Skip to content
AI.info

Research

Compression and Inference of Spiking Neural Networks on Resource-Constrained Hardware

Compression and Inference of Spiking Neural Networks on Resource-Constrained Hardware Overview Research area: Neuromorphic computing / embedded machine learning — specifically Spiking Neural Networks

Compression and Inference of Spiking Neural Networks on Resource-Constrained Hardware
arXiv
2511.12136
Published
2025-11-15
Authors
Karol C. Jurzec, Tomasz Szydlo, Maciej Wielgosz

AI summary

Compression and Inference of Spiking Neural Networks on Resource-Constrained Hardware

Overview

Research area: Neuromorphic computing / embedded machine learning — specifically Spiking Neural Networks (SNNs) and their deployment on microcontrollers.

Technical level: Intermediate. The paper assumes familiarity with neural networks, spiking neuron models (Leaky Integrate-and-Fire), and basic embedded systems concepts, but the methodology is described accessibly.

Scope (1 sentence): This paper presents a custom C-based inference runtime plus spike-driven pruning that executes SNNTorch-trained spiking neural networks on desktop CPUs and an Arduino Portenta H7 microcontroller, achieving order-of-magnitude speedups and large memory reductions compared with a Python baseline.

What This Paper Is About

Spiking neural networks communicate through discrete spikes over time rather than continuous activations, which makes them attractive for low-power, event-driven computing, but the Python frameworks used to train them (SNNTorch, Brian2) carry interpreter overhead and memory costs that prevent deployment on embedded devices. The paper's goal is to build a lightweight C runtime that loads SNN models exported from SNNTorch as JSON and runs them efficiently on resource-constrained hardware, while exploiting the sparse firing behaviour of SNNs to prune inactive neurons and filters without losing accuracy.

Key Contributions

  1. A modular C-based SNN inference runtime consisting of three components: a Network Loader that parses JSON network descriptions (layers, weights, connectivity) into C data structures, a Dataset Loader that converts event-based input into discrete time-step frames, and an Execution Pipeline that runs the layers forward over simulation time. It supports convolution, fully-connected, and Leaky Integrate-and-Fire layers, uses Structure-of-Arrays memory layout for cache efficiency, and pre-allocates all buffers up front.
  2. A spike-driven pruning method that records per-neuron spike counts over a representative data subset, groups spiking neurons by their source convolutional filter, and removes filters whose neuron groups fire zero or very few times. Pruning is propagated forward so that a removed convolutional output channel also removes the corresponding input channel of the next convolutional layer, and is applied iteratively and conservatively.
  3. A functional-equivalence validation showing the C runtime reproduces SNNTorch's predicted labels on all tested samples for both N-MNIST and ST-MNIST, and that pruning only zero-activity neurons leaves accuracy unchanged.
  4. Cross-platform benchmarking on a desktop Intel Core i3-10100F CPU and an Arduino Portenta H7 (STM32H747, Cortex-M7 at 480 MHz with a lower-power Cortex-M4), reporting inference latency, memory footprint, and the feasibility of real-time inference. Code is released at https://github.com/karol-jurzec/snn-generator/.

Main Findings

  • 11× desktop speedup for the C runtime: On N-MNIST (averaged over approximately 300 time steps), a single-sample inference took 2.393 s in Python (SNNTorch) versus 0.224 s in C without pruning — a 10.68× speedup — with identical accuracy of 84.20% and F1 of 0.843.
  • Pruning roughly doubles that gain: With pruning, N-MNIST inference dropped to 0.113 s, a 21.18× speedup over Python and roughly 2× over the unpruned C version, while accuracy rose slightly to 84.60% and F1 to 0.848. This corresponded to removing approximately 20% of neurons and their associated synapses.
  • Larger benefit on ST-MNIST: ST-MNIST showed a larger fraction of inactive units, producing over a 7× reduction in inference time relative to the unpruned baseline. The paper reports a 0% accuracy drop in all ST-MNIST configurations shown (Fig. 6).
  • Memory footprint far below Python: The SNNTorch environment requires hundreds of megabytes to load the interpreter, libraries, and model, which is infeasible on microcontrollers. The compiled C model and buffers for N-MNIST occupy only a few hundred kilobytes — approximately 50 KB for weights and 200 KB for neuron states and buffers.
  • Pruning reduces peak memory: Peak memory fell from approximately 6 MB to 5.9 MB on the desktop PC and from 0.58 MB to 0.46 MB on the Arduino. The paper frames the embedded reduction as valuable headroom on a device with only 1 MB of available SRAM (note: the platform description elsewhere states 512 KB SRAM for the Cortex-M7, and the paper refers both to 1024 KB and 1 MB of SRAM in later passages).
  • Microcontroller deployment works: The Portenta H7 produced correct classifications for N-MNIST and ST-MNIST samples. Unoptimized N-MNIST inference on the Cortex-M7 took on the order of a few seconds per sample, and the authors anticipate on the order of a few hundred milliseconds per inference with the fully optimized model, extrapolating from the 10× desktop speedup. Microcontroller tests ran single-threaded due to the single-core nature of the target configuration.
  • Accuracy was preserved by conservative pruning: Because only neurons with zero activity were removed in one round after initial training, classification accuracy was unchanged on the test sets, and minor discrepancies in spike timing did not affect final predicted labels.

Methodology in Plain English

The researchers trained SNN models in SNNTorch (on top of PyTorch) using backpropagation-through-time with surrogate gradients, then exported the trained architecture and weights to JSON files of a few megabytes using a custom exporter. A hand-written C program on the target device parses those JSON files, reconstructs the network in memory, converts event-based datasets into a fixed number of time-step frames (for example, 10 frames per N-MNIST sample), and runs the forward pass step by step. For classification, the final layer has one spiking neuron per class and the prediction is the neuron with the highest spike count over the inference window.

Because SNNs fire sparsely, the authors recorded how many spikes each neuron emitted on a representative subset of the data. Since each spiking neuron receives input from a specific convolutional filter, they grouped neurons by source filter and deleted filter groups whose neurons never fired — a safe criterion that guarantees no change in output for the analysed data. Deleting a filter also shrinks the corresponding input channel of the next convolutional layer. Models were trained for one epoch, which the authors describe as sufficient for these relatively simple tasks. Evaluation compared Python CPU-mode SNNTorch against the C runtime on the same inputs, measured end-to-end inference time for one sample averaged over 500 runs, and estimated memory by summing the sizes of all C data structures and monitoring heap and stack usage on the microcontroller.

Two datasets were used: N-MNIST, a spiking version of MNIST captured with a Dynamic Vision Sensor, which preserves MNIST's 60,000 training and 10,000 test instances as asynchronous address-event streams; and ST-MNIST, a spiking tactile dataset in which human participants wrote digits on a 10×10 array of artificial tactile sensors, providing roughly 30,000 event sequences of digits 0–9. The N-MNIST SNN had two 5×5 convolutional layers, each followed by a LIF spiking layer (leak factor β = 0.5) and 2×2 max pooling, then a fully-connected layer and a 10-neuron output LIF layer, with on the order of tens of thousands of parameters. The ST-MNIST SNN was simpler, with two fully-connected LIF layers (hidden size 128) and a 10-neuron output layer.

Why This Matters

The paper addresses the deployment gap between SNN research frameworks and practical embedded hardware. It shows that SNNs — usually pitched as requiring specialised neuromorphic chips — can run efficiently on conventional microcontrollers when paired with low-level software and model-level compression, and it treats sparsity as a structural optimisation signal rather than only an energy argument.

Real-world applications:

  • Event-based vision on battery-powered cameras and IoT sensors, where Dynamic Vision Sensor streams must be classified on-device.
  • Tactile and touch sensing in robotics or prosthetics, using the ST-MNIST style of spike-stream input from sensor arrays.
  • Always-on keyword, gesture, or anomaly detection on wearables and mobile devices with strict memory and latency budgets.
  • Edge AI in industrial or remote settings where connectivity is limited and per-device power draw matters.

Industry relevance: For embedded ML engineers, the contribution is a concrete pattern for exporting trained models into a compact C runtime with static memory allocation, avoiding interpreter and garbage-collection overhead. For neuromorphic hardware vendors and chip designers, it suggests a software path that lets spiking models reach existing Cortex-M class silicon today rather than waiting for dedicated neuromorphic accelerators.

Future Directions

  1. Measure power consumption — explicitly listed as future work; the paper only reports latency and memory, so the energy-efficiency argument for SNNs remains unquantified here.
  2. Validate real-time performance under streaming input conditions, rather than single-sample benchmark runs.
  3. Explore more aggressive pruning thresholds with retraining or fine-tuning to recover accuracy, since the paper only performed one conservative round removing zero-activity neurons.
  4. Exploit multi-core and RTOS parallelism by partitioning workloads between the Cortex-M7 and Cortex-M4 cores, which was not exercised because microcontroller tests ran single-threaded.

Target Audience

Embedded and edge machine learning engineers who need to run neural models within tight memory and latency budgets; neuromorphic computing researchers interested in bridging training frameworks and deployment hardware; and students or practitioners with a background in neural networks and C programming who want a concrete example of model export, pruning, and microcontroller inference. Readers looking for state-of-the-art SNN accuracy results or energy measurements will not find them here, since the paper's focus is deployment efficiency and functional equivalence.

Authors’ abstract

Spiking neural networks (SNNs) communicate via discrete spikes in time rather than continuous activations. Their event-driven nature offers advantages for temporal processing and energy efficiency on resource-constrained hardware, but training and deployment remain challenging. We present a lightweight C-based runtime for SNN inference on edge devices and optimizations that reduce latency and memory without sacrificing accuracy. Trained models exported from SNNTorch are translated to a compact C representation; static, cache-friendly data layouts and preallocation avoid interpreter and allocation overheads. We further exploit sparse spiking activity to prune inactive neurons and synapses, shrinking computation in upstream convolutional layers. Experiments on N-MNIST and ST-MNIST show functional parity with the Python baseline while achieving ~10 speedups on desktop CPU and additional gains with pruning, together with large memory reductions that enable microcontroller deployment (Arduino Portenta H7). Results indicate that SNNs can be executed efficiently on conventional embedded platforms when paired with an optimized runtime and spike-driven model compression. Code: https://github.com/karol-jurzec/snn-generator/

Read the original paper