Skip to content
AI.info

Research

Synergy Between the Strong and the Weak: Spiking Neural Networks are Inherently Self-Distillers

Overview Research area: Spiking neural networks (SNNs), knowledge distillation, and energy-efficient / neuromorphic machine learning. Technical level: Intermediate. The paper assumes familiarity with

Synergy Between the Strong and the Weak: Spiking Neural Networks are Inherently Self-Distillers
arXiv
2510.07924
Published
2025-10-09
Authors
Yongqi Ding, Lin Zuo, Mengmeng Jing, Kunshan Yang, Pei He, Tonglan Xie

AI summary

Overview

Research area: Spiking neural networks (SNNs), knowledge distillation, and energy-efficient / neuromorphic machine learning.

Technical level: Intermediate. The paper assumes familiarity with spiking neurons (LIF dynamics, surrogate gradients), logit-based knowledge distillation, and standard image classification benchmarks, though the core idea is describable without heavy mathematics.

Scope: The paper proposes and evaluates two label-free self-distillation schemes — Strong2Weak and Weak2Strong — that treat each timestep of an SNN as a separate submodel and distill between the most and least confident submodels, improving accuracy, low-latency inference, and adversarial robustness without adding teachers or auxiliary modules.

What This Paper Is About

Spiking neural networks save energy by passing binary spikes instead of running dense multiply-accumulate operations, but they still trail conventional artificial neural networks in accuracy. Prior work closes that gap using knowledge distillation with a separate, larger teacher network — which costs extra storage, extra pre-training, and often task-specific customization. This paper asks whether an SNN can distill itself: because its output changes across timesteps, an SNN can be logically split into one submodel per timestep, and the strongest and weakest of those submodels can teach each other at no additional cost.

Key Contributions

  1. Temporal deconstruction of the SNN. The authors show that an SNN running for T timesteps can be logically decomposed into T submodels that share architecture and parameters but produce different outputs, creating the multiplicity of outputs that distillation requires — with no extra modules.

  2. Confidence-based identification of strong and weak submodels. The output confidence (maximum softmax probability) of each submodel is used to rank them, requiring no labels, no step-by-step comparison with ground truth, and minimal computation.

  3. Two self-distillation schemes: Strong2Weak and Weak2Strong. In Strong2Weak the highest-confidence submodel teaches the lowest-confidence one; in Weak2Strong the weakest submodel teaches the strongest, transferring "dark knowledge" and acting as a regularizer.

  4. Flexible implementations and broad empirical validation. Distillation losses (KL divergence, MSE, logit standardization) and teacher/student configurations (one-to-one, ensemble teacher, ensemble student, simultaneous, cascade) can be swapped, and gains are demonstrated on static and neuromorphic datasets across VGG, MS-ResNet, Transformer (SDT), and QKFormer architectures.

Main Findings

  • Consistent accuracy gains over vanilla SNNs. On CIFAR10 with VGG-9, vanilla scores 94.21%; Strong2Weak reaches 94.79% (+0.58) and Weak2Strong 94.70% (+0.49). With MS-ResNet18, vanilla is 94.88%, Strong2Weak 95.15% (+0.27) and Weak2Strong 95.13% (+0.25).

  • Larger gains on neuromorphic data. On CIFAR10-DVS the gains are far bigger: VGG-9 vanilla 73.97%, Strong2Weak 78.93% (+4.96), Weak2Strong 79.33% (+5.36); MS-ResNet18 vanilla 66.40%, Strong2Weak 70.50% (+4.10), Weak2Strong 71.57% (+5.17). The paper attributes this to the richer temporal features of event data. The paper's text states CIFAR10-DVS accuracy improved by up to 5.26%, while its ablation table lists deltas as high as +5.36.

  • Naive accuracy/loss-based distillation is weaker. In Table 1, a "high-accuracy teacher" gives 78.43% (±0.33) on CIFAR10-DVS and 90.62% (±0.49) on DVS-Gesture, and a "low-loss teacher" gives 78.93% (±0.33) and 90.39% (±0.71), versus Strong2Weak at 78.93% (±0.12) / 91.43% (±0.43) and Weak2Strong at 79.33% (±0.29) / 91.20% (±0.33).

  • Competitive against published methods. On ImageNet at 4 timesteps, Strong2Weak reaches 70.53% with SEW-ResNet34, versus 69.60% for TKS (a teacher-free SNN distillation method), a 0.93-point margin; Weak2Strong reaches 69.87%. On CIFAR10/CIFAR100 the paper reports 96.66% and 82.02% with ResNet-19.

  • Strong results on neuromorphic benchmarks. Weak2Strong reaches 86.70% on CIFAR10-DVS with VGGSNN at T=10, which the paper states is 1.40% above TKS (85.30%); Strong2Weak reaches 85.60%. On DVS-Gesture, cascade distillation produces 93.41% (Strong2Weak) and 92.13% (Weak2Strong).

  • Better behavior at low inference timesteps. Using a 5-timestep pre-trained model on CIFAR10-DVS, vanilla SNN collapses to 10.00% at T=1, while Strong2Weak reaches 71.50% and Weak2Strong 73.40%; at T=5 the values are 78.60% and 79.70% versus 74.10% for vanilla.

  • Improved adversarial robustness. On CIFAR100 with VGG-11 at T=8, combined with RAT the clean accuracy rises to 70.68% (Strong2Weak) and 70.09% (Weak2Strong) versus 69.99% for RAT alone. Weak2Strong improves FGSM robust accuracy to 21.23% from RAT's 19.00%, a 2.23-point gain; PGD robustness goes from 9.11% to 10.70%.

  • Compatibility with early-exit dynamic inference. On CIFAR10-DVS with an exit threshold of 0.8, Strong2Weak achieves 77.80% with an average of 2.30 timesteps and Weak2Strong 78.80% with 2.27 timesteps, roughly half the timesteps of the full 5-step model.

  • Different confidence metrics perform similarly. Replacing confidence with entropy, margin, or diversity changes results only slightly (e.g., Strong2Weak on CIFAR10-DVS: confidence 78.93, entropy 78.80, margin 78.83, diversity 78.27; the corresponding second group: confidence 79.33, entropy 78.57, margin 78.57, diversity 78.73), supporting the use of the simple confidence measure.

  • Simultaneous Strong2Weak + Weak2Strong does not clearly help. The paper reports that using both at once did not significantly improve performance because excessive similarity between submodels reduces overall diversity, a trade-off left for future work.

Methodology in Plain English

An SNN's neurons keep a membrane potential that leaks over time; a spike is emitted only when that potential crosses a threshold. Because the membrane potential and the incoming current differ at every timestep, the same network produces different outputs at t=1, t=2, … t=T. The authors exploit this by treating each timestep's output as if it came from a separate model.

At each training iteration they compute the softmax output of every timestep submodel and take the maximum probability as that submodel's "confidence." The highest-confidence submodel is labeled strong and the lowest-confidence one weak. Two distillation directions are then possible. In Strong2Weak, the strong submodel's softened output distribution (temperature α = 2) is used as the teacher for the weak submodel via a KL divergence loss. In Weak2Strong the roles are reversed: the weak output teaches the strong one. The distillation loss is added to the ordinary cross-entropy loss with coefficient λ = 1 for both schemes, and training uses a rectangular surrogate gradient (width parameter a = 1.0) since spikes are non-differentiable.

The schemes are described as implementations of flexible components: the loss can be KL divergence, MSE, or logit standardization, and the teacher or student can be a single submodel, an ensemble of T−1 submodels, both directions at once, or a cascade ordered by confidence. Confidence is averaged over a batch rather than computed per sample.

Why This Matters

Research impact. The paper reframes self-distillation for SNNs as something the architecture already provides, rather than something bolted on. It removes the teacher model, its storage and pre-training cost, and the need for labels during the distillation step — all of which are barriers in prior SNN distillation work such as TKS and TSSD. It also links distillation to adversarial robustness through the stability of submodel outputs, a connection the authors say they intend to explore further.

Real-world applications (as motivated by the paper):

  • Power-constrained edge devices, where the paper notes ANN multiply-accumulate costs limit deployment, while an SNN on the Tianmouc chip needs only 0.7 mW for typical vision tasks.
  • Neuromorphic chips paired with dynamic vision sensors, where temporal event streams give the largest gains in these experiments.
  • Latency-critical inference, since a single trained model can be run at 1–5 timesteps without retraining, or combined with early-exit thresholds.
  • Streaming or online learning settings where labels are unavailable, because strong/weak identification relies only on output confidence.

Industry relevance. Neuromorphic hardware and energy-efficient AI are active commercial and defense-adjacent directions; a training method that improves accuracy and adversarial robustness with zero additional parameters or teacher models is directly compatible with existing inference hardware and pipelines, and it applies across VGG, ResNet, and Transformer variants.

Future Directions

  • Balance diversity and similarity between submodels. The paper states that running Strong2Weak and Weak2Strong simultaneously did not clearly help because excessive similarity reduces diversity, and explicitly leaves this ensemble-learning problem for future work.
  • Tune the distillation losses and coefficients. The stated limitation is that no deliberate tuning of loss functions or coefficients was done, so performance could rise further.
  • Explore the robustness connection more deeply. The authors say they will further study how distillation promotes robustness.
  • Extend to more temporal tasks. The paper suggests the approach could generalize to other temporally rich tasks, and notes that combination with CILF neurons on Tiny-ImageNet is another promising direction.

Target Audience

Researchers and graduate students working on spiking neural networks, neuromorphic computing, or knowledge distillation who want efficient training methods without extra teacher models; practitioners building low-power or low-latency vision systems on neuromorphic hardware; and readers interested in how architectural properties of SNNs can be repurposed as a training signal. The experimental setup, alternative implementation analyses, and overall output visualizations are deferred to appendices that are not fully included in the provided paper content.

Authors’ abstract

Brain-inspired spiking neural networks (SNNs) promise to be a low-power alternative to computationally intensive artificial neural networks (ANNs), although performance gaps persist. Recent studies have improved the performance of SNNs through knowledge distillation, but rely on large teacher models or introduce additional training overhead. In this paper, we show that SNNs can be naturally deconstructed into multiple submodels for efficient self-distillation. We treat each timestep instance of the SNN as a submodel and evaluate its output confidence, thus efficiently identifying the strong and the weak. Based on this strong and weak relationship, we propose two efficient self-distillation schemes: (1) \textbf{Strong2Weak}: During training, the stronger "teacher" guides the weaker "student", effectively improving overall performance. (2) \textbf{Weak2Strong}: The weak serve as the "teacher", distilling the strong in reverse with underlying dark knowledge, again yielding significant performance gains. For both distillation schemes, we offer flexible implementations such as ensemble, simultaneous, and cascade distillation. Experiments show that our method effectively improves the discriminability and overall performance of the SNN, while its adversarial robustness is also enhanced, benefiting from the stability brought by self-distillation. This ingeniously exploits the temporal properties of SNNs and provides insight into how to efficiently train high-performance SNNs.

Read the original paper