Research
Stable Spike: Dual Consistency Optimization via Bitwise AND Operations for Spiking Neural Networks
Stable Spike: Dual Consistency Optimization via Bitwise AND Operations for Spiking Neural Networks Authors: Yongqi Ding, Kunshan Yang, Linze Li, Yiyang Zhang, Mengmeng Jing, Lin Zuo (corresponding aut

- arXiv
- 2603.11676
- Published
- 2026-03-12
- Authors
- Yongqi Ding, Kunshan Yang, Linze Li, Yiyang Zhang, Mengmeng Jing, Lin Zuo
AI summary
Stable Spike: Dual Consistency Optimization via Bitwise AND Operations for Spiking Neural NetworksAuthors: Yongqi Ding, Kunshan Yang, Linze Li, Yiyang Zhang, Mengmeng Jing, Lin Zuo (corresponding author, linzuo@uestc.edu.cn) Affiliation: School of Information and Software Engineering, University of Electronic Science and Technology of China arXiv: 2603.11676v1 [cs.NE], 12 Mar 2026 | License: CC BY-NC-ND 4.0 Funding acknowledged: National Natural Science Foundation of China, Grant No. 62276054
Overview
Research area: Spiking neural networks (SNNs), neuromorphic computing, and neuromorphic object recognition.
Technical level: Advanced. The paper assumes familiarity with spiking neuron dynamics (leaky integrate-and-fire models), timestep-based inference, and training objectives such as cross-entropy and KL divergence.
Scope: The paper proposes "Stable Spike," a training-time method that decouples a consistent spike skeleton across timesteps using bitwise AND operations and injects amplitude-aware spike noise, improving SNN accuracy without modifying neurons or architecture.
What This Paper Is About
SNNs transmit sparse binary spikes over multiple discrete timesteps, and because neuron membrane potentials carry over between timesteps, the spike maps and predictions produced at different timesteps vary substantially. This inherent inconsistency (including many redundant, variable "noise" spikes) degrades representation quality, especially when only the earliest timesteps are available for low-latency inference. The paper's goal is to directly enforce consistency across timesteps while preserving feature diversity, so SNNs become more accurate at ultra-low latency without changing the spiking neuron model or the network architecture.
Key Contributions
- Stable spike decoupling via bitwise AND. The authors extract a "stable spike feature skeleton" across timesteps by applying the hardware-friendly AND operation to spike maps of adjacent timesteps, then use it as an anchor to guide the variable, unstable spike maps toward consistency.
- Amplitude-aware spike noise for perturbation consistency. They inject discrete binary spike noise into the stable spike firing rate, with the noise probability at each position set to that position's own stable firing rate, then align the perturbed prediction's probability distribution with the clean prediction.
- Plug-and-play, neuron- and architecture-agnostic design. The method requires no modification to spiking neurons or model architecture, and the paper demonstrates it can be combined with other SNN methods.
- Extensive validation across architectures and datasets. Experiments cover VGG-9, ResNet-18, QKFormer, VGGSNN, and ResNet-34 on CIFAR10-DVS, DVS-Gesture, N-Caltech101, CIFAR10/100, and ImageNet, including ultra-low-latency regimes.
Main Findings
-
Consistent gains from both loss components across architectures. On CIFAR10-DVS, adding the spike consistency loss alone raised VGG-9 from 72.9 to 75.2 (+2.4), ResNet-18 from 66.1 to 69.7 (+3.6), and QKFormer from 81.2 to 82.5 (+1.3); adding both losses gave 77.1 (+4.2), 70.3 (+4.2), and 82.9 (+1.7) respectively. On DVS-Gesture, both losses gave VGG-9 94.44 (+7.29), ResNet-18 85.42 (+3.83), and QKFormer 95.49 (+1.74) from baselines of 87.15, 81.59, and 93.75.
-
Largest gains appear at ultra-low latency. On DVS-Gesture the improvement at T=2 timesteps was 8.33% (83.68 to 92.01), decreasing to +7.29 at T=4 (87.15 to 94.44), +5.91 at T=6 (88.19 to 94.10), +4.52 at T=8 (90.97 to 95.49), and +3.48 at T=10 (92.01 to 95.49). On CIFAR10-DVS the gains were +2.0 at T=2, +4.2 at T=4, +3.8 at T=6, +2.6 at T=8, and +2.7 at T=10.
-
AND outperforms other bitwise operations. On CIFAR10-DVS, AND scored 77.1 versus OR at 68.9 and XOR at 74.5; on DVS-Gesture, AND scored 94.44 versus OR at 88.54 and XOR at 89.58. The authors attribute OR's degradation to it simultaneously retrieving consistent and inconsistent spike patterns.
-
Amplitude-aware noise beats fixed-probability and Gaussian noise. On DVS-Gesture, fixed noise probabilities gave p=0.4 → 87.15, p=0.5 → 88.89, p=0.6 → 86.81, while Gaussian noise gave std=0.1 → 88.19, std=0.5 → 91.67, std=1.0 → 89.93, all below the proposed method's 94.44. On CIFAR10-DVS these variants ranged from 74.5 to 75.5 versus 77.1 for the proposed method. The paper notes that a fixed probability of p=0.75 caused the model to fail to converge.
-
The consistency function is not restricted to MSE. On CIFAR10-DVS, spike consistency with MSE reached 77.1, KL 76.2, and cosine 75.6 (versus vanilla 72.9); the noise loss with MSE reached 76.1, KL 77.1, and cosine 76.0. On DVS-Gesture the corresponding numbers were 94.44 (MSE), 94.10 (KL), 93.40 (cosine) for the spike loss, versus vanilla 87.15.
-
Results on neuromorphic benchmarks. With VGGSNN and standard data augmentation, CIFAR10-DVS reached 83.7 at 4 timesteps (compared with MPS's 83.2 on VGGSNN at 4 timesteps) and VGG-9 without augmentation reached 77.1 at 4 timesteps (MPS reported 76.77 with VGG-9 at 5 timesteps). On N-Caltech101, combining with Knowledge-Transfer gave 94.25 at 10 timesteps; VGG-9 reached 83.92 at 4 timesteps. On DVS-Gesture, QKFormer reached 98.61 at 16 timesteps and 95.49 at 4 timesteps.
-
Static dataset results. On ImageNet with ResNet-34 at 4 timesteps the method reached 70.59%, above Strong2Weak (70.53), STAA-SNN (70.40), FSTA-SNN (70.23), RateBP (70.01), MPS (69.03), IMP+LTS (68.90), Shortcut (68.14), TAB (67.78), EnOF (67.40), Weak2Strong (69.87), and SSCL (66.78). CIFAR10 and CIFAR100 reached 96.73% and 82.29% respectively.
-
Lower firing rate and power in most layers. On CIFAR10-DVS with VGG-9, the proposed method reduced firing rates in layers 2 through 7 (for example, layer 1 went from 9.24 to 10.14, layer 2 from 4.26 to 3.61, and layer 5 from 2.17 to 1.70), with total power dropping from 189.83 to 181.02 (×10⁶ pJ). The first layer's firing rate rose and the final listed layer rose from 7.63 to 9.81.
-
Smoother loss landscape. Visualizations show the vanilla SNN's loss landscape has multiple local minima and saddle points, while the proposed method's landscape is smoother and more centralized with a clear global minimum, despite spike noise being applied during training.
-
Robustness to balance coefficients. Sweeping β and γ over {0.1, 0.25, 0.5, 0.75, 1.0, 1.25, 1.5, 1.75, 2.0} on DVS-Gesture produced a lowest accuracy of 92.01% and a highest of 95.14%, both above the vanilla SNN's 87.15%.
-
Compatibility with other SNN methods. Adding the method on CIFAR10-DVS improved CLIF from 74.3 to 76.2 (+1.9), TAB from 73.1 to 75.4 (+2.3), and SLT from 74.1 to 75.9 (+1.8). On DVS-Gesture it improved CLIF from 89.58 to 95.83 (+6.25), TAB from 87.50 to 92.36 (+4.86), and SLT from 88.19 to 90.97 (+2.78).
Methodology in Plain English
The method builds on the observation that while different timesteps in a trained SNN produce different-looking spike maps, they still tend to capture object-relevant features; the variability mostly comes from redundant noise spikes. The authors exploit this by taking the spike maps of two adjacent timesteps and applying a bitwise AND, which keeps only positions where both timesteps fired a 1. Repeating this over T timesteps yields T−1 "stable spikes," and averaging them gives a stable spike firing rate that acts as a consistent feature skeleton.
The original T-timestep spike firing rate is then pushed to match this skeleton using a mean squared error loss (KL divergence and cosine similarity were tested as alternatives, and the paper notes the method works with them too). This is the first half of the "dual consistency."
For the second half, the authors address the fact that neural networks benefit from feature diversity but SNNs cannot simply absorb continuous Gaussian noise: binary spikes must stay binary to avoid a training-inference precision mismatch, and discrete firing rates are sensitive to noise amplitude. Their solution, amplitude-aware spike noise, samples a binary 1 or 0 at each position from a Bernoulli distribution whose probability equals that position's own stable firing rate value. High-firing-rate elements are therefore perturbed more often, and low-firing-rate elements are perturbed less, so key semantics are preserved. The perturbed firing rate is then forward-propagated through the classifier to produce a noise prediction, and a KL divergence between the temperature-softened distributions of the clean and noisy predictions encourages the SNN to make perturbation-consistent predictions. Temperature α is set to 2.
Both losses are combined with the standard cross-entropy loss as L_total = L_CE + β·L_spike + γ·L_noise, with β and γ defaulting to 1.0. Stable spikes and consistency guidance are computed only on backbone features, so extra forward propagation is needed only past the classifier. The final output is the original SNN output averaged over T timesteps; the noise prediction has no temporal dimension and is used directly. Default experiments use T=4 timesteps, with the LIF neuron's membrane time constant τ at its default 2.0. The accompanying supplementary material includes an interpretation of the AND operation through mutual information, experimental setup details, additional results, an overhead analysis stating training overhead is negligible and inference is entirely unaffected, and further visualizations.
Why This Matters
Impact on research. The work directly targets temporal inconsistency, which prior work had addressed only indirectly by altering neuron dynamics or distilling logits between adjacent timesteps — approaches that are hard to deploy on neuromorphic chips where neuron models are predetermined. Because Stable Spike leaves neurons and architecture untouched, it functions as a drop-in training enhancement that composes with other SNN advances, as demonstrated with CLIF, TAB, and SLT. It also frames the AND operation as a mutual-information-style consensus extractor between timesteps, which may inspire other consistency mechanisms.
Real-world applications:
- Event-camera-based gesture recognition, where low-latency classification is required and the method delivered 92.01% at just 2 timesteps on DVS-Gesture.
- Neuromorphic object recognition on edge devices such as drones or robots, where the paper reports reduced power consumption on CIFAR10-DVS with VGG-9 (181.02 vs 189.83 ×10⁶ pJ).
- Always-on vision sensors that must run within tight power budgets, since spike rates dropped across most layers.
- General energy-constrained classification, given demonstrated gains on static ImageNet, CIFAR10, and CIFAR100 in addition to event data.
Industry relevance. The method requires no custom neuron hardware, adds no inference cost according to the paper, and lowers measured power on most layers — all favorable properties for neuromorphic chip deployment, where the paper notes neuron models are typically fixed in advance.
Future Directions
- Whether the method's benefit is sustained at very large scale and on longer-duration neuromorphic datasets, since the paper's ImageNet evaluation uses 4 timesteps and ResNet-34.
- How far the consistency objective can be generalized: the paper shows KL divergence and cosine similarity also yield gains, but does not identify which function is optimal or under what conditions.
- How to avoid manual tuning of the balance coefficients β and γ, given the observed spread from 92.01% to 95.14% across the tested range on DVS-Gesture.
- Whether the amplitude-aware noise formulation can be adapted to other spiking neuron models and SNN training pipelines beyond those tested (CLIF, TAB, SLT, and the backbone architectures evaluated).
- The paper identifies no explicit limitation, and the truncated content does not report an error analysis or failure cases.
Target Audience
Researchers and engineers working on spiking neural networks, neuromorphic computing, and event-based vision, particularly those concerned with low-latency or energy-constrained inference. It is also relevant to practitioners seeking training-time techniques that can be added to existing SNN pipelines without hardware or architectural changes, and to readers interested in consistency regularization and data-augmentation-style perturbation methods adapted to binary, discrete representations. Readers without background in spiking neuron dynamics will need to consult the cited literature first, since the paper assumes that context.
Authors’ abstract
Although the temporal spike dynamics of spiking neural networks (SNNs) enable low-power temporal pattern capture capabilities, they also incur inherent inconsistencies that severely compromise representation. In this paper, we perform dual consistency optimization via Stable Spike to mitigate this problem, thereby improving the recognition performance of SNNs. With the hardware-friendly ``AND" bit operation, we efficiently decouple the stable spike skeleton from the multi-timestep spike maps, thereby capturing critical semantics while reducing inconsistencies from variable noise spikes. Enforcing the unstable spike maps to converge to the stable spike skeleton significantly improves the inherent consistency across timesteps. Furthermore, we inject amplitude-aware spike noise into the stable spike skeleton to diversify the representations while preserving consistent semantics. The SNN is encouraged to produce perturbation-consistent predictions, thereby contributing to generalization. Extensive experiments across multiple architectures and datasets validate the effectiveness and versatility of our method. In particular, our method significantly advances neuromorphic object recognition under ultra-low latency, improving accuracy by up to 8.33\%. This will help unlock the full power consumption and speed potential of SNNs.