Research
Convolutional Spiking-based GRU Cell for Spatio-temporal Data
Overview Research area: Spiking Neural Networks (SNNs) and recurrent architectures for temporal and spatio-temporal data, specifically a spiking variant of the Gated Recurrent Unit (GRU). Technical le

- arXiv
- 2510.25696
- Published
- 2025-10-29
- Authors
- Yesmine Abdennadher, Eleonora Cicciarella, Michele Rossi
AI summary
Overview
- Research area: Spiking Neural Networks (SNNs) and recurrent architectures for temporal and spatio-temporal data, specifically a spiking variant of the Gated Recurrent Unit (GRU).
- Technical level: Intermediate to Advanced. Readers need familiarity with GRU gating, LIF/Cuba-LIF neuron models, surrogate gradients, and event-based vision benchmarks to follow the equations, though the intuition behind the four modifications is accessible.
- Scope: The paper proposes a Convolutional Spiking GRU (CS-GRU) cell built from four modifications to the existing SpikGRU design, and evaluates it on two temporal speech datasets (N-TIDIGITS, SHD) and three spatio-temporal benchmarks (MNIST, DVSGesture, CIFAR10-DVS), reporting accuracy and spiking-activity-based energy efficiency. (arXiv:2510.25696v1 [cs.LG], 29 Oct 2025; authors from the Department of Information Engineering, University of Padova.)
What This Paper Is About
Recurrent models like GRUs are good at tracking long-term sequence dynamics, but they tend to lose fine-grained local structure in the input, and spiking neuron models such as Cuba-LIF are limited in handling complex sequential inputs — local information gets diluted over long sequences. The paper addresses this by modifying the existing SpikGRU cell so that it preserves local spatio-temporal dependencies through convolution, dynamically controls how past information is retained, uses the current variable when computing its update gate, and swaps the Heaviside-based surrogate gradient for an arctangent one. The goal is a single cell that works well on both purely temporal and event-driven spatio-temporal classification tasks while firing fewer spikes.
Key Contributions
- A dynamic current gate (mod1): The constant decay parameter α in the SpikGRU current equation is replaced by a gate r_t computed from input spikes with its own weights W_r, U_r and bias b_r, giving time-dependent control over retention of past information.
- Current-aware update gate (mod2): The update gate z_t is recomputed from the current variable i_t rather than only from the membrane-potential-related terms, making the gate context-sensitive to rapidly changing or bursty spiking activity.
- Convolutional operations in the cell (mod3): Fully-connected matrix products such as W s_t are replaced with convolutional filters W * s_t so the cell can learn local spatial patterns that SpikGRU may miss.
- Arctangent surrogate gradient (mod4): The Heaviside activation's non-differentiability is handled with a scaled arctangent surrogate gradient instead of a triangular surrogate, described as always nonzero for all membrane potentials because of its polynomially decreasing slope.
- Empirical validation: The concurrent combination of all four modifications (CS-GRU) is tested against SpikGRU, a vanilla GRU, Cuba-LIF, and each individual/partial modification combination across five benchmarks.
Main Findings
- Temporal data, combined modifications win: On a 1×128 network, SpikGRU-mod1-2-3-4 (the full CS-GRU) reaches 90.66% on N-TIDIGITS and 91.96% on SHD — the highest accuracy of any configuration in the table.
- Baselines on temporal data: Vanilla GRU scores 86.19% (N-TIDIGITS) and 84.49% (SHD); Cuba-LIF scores 81.25% and 73.01%; SpikGRU scores 86.23% and 87.89%.
- Individual modifications vary widely: SpikGRU-mod2 reaches 86.99% / 87.63%; mod3 alone reaches 82.12% on N-TIDIGITS but 90.06% on SHD; mod1 alone drops to 82.61% / 84.71%; mod4 alone gives 86.28% / 85.77%.
- Some combinations collapse: SpikGRU-mod1-3 falls to 37.61% on N-TIDIGITS, mod2-3 to 69.55% / 40.50%, mod3-4 to 55.79% / 8.43%, mod1-2-3 to 73.9% / 89.48%, mod1-3-4 to 24.60% / 86.79%, and mod2-3-4 to 17.30% / 86.7%.
- Spatio-temporal data: CS-GRU scores 99.31% on MNIST (vs. SpikGRU's 98.10%), 82% on DVSGesture (vs. 66%), and 51.40% on CIFAR10-DVS (vs. 26.53%) — improvements that grow as the data become more event-driven and spatio-temporally complex.
- Overall headline: The abstract reports an average improvement of 4.35% over SpikGRU, accuracies above 90% on the sequential datasets, and up to 99.31% on MNIST.
- Energy efficiency: On the DVS Gesture dataset, CS-GRU fires 0.13 spikes per neuron per timestep versus SpikGRU's 0.42, a relative reduction of 69.05% in average spiking activity rate, used as a proxy for energy consumption since inactive neurons consume no power on neuromorphic hardware.
Methodology in Plain English
The team started from the existing SpikGRU cell — itself a combination of the Cuba-LIF spiking neuron model with GRU-style gating — and made four targeted changes.
First, they removed the fixed decay constant that SpikGRU uses to carry over past input current, and replaced it with a learned gate that looks at the input spikes, so the model can decide dynamically how much history to keep. Second, they changed the update gate to be computed from the input current rather than from the previous membrane state alone, letting the cell respond more sharply to bursts of spiking activity. Third, they swapped every fully-connected matrix multiplication inside the cell for a convolution, so the cell processes input shaped as channels × height × width and learns local spatial patterns rather than flattening everything into a vector. Fourth, because spikes are produced by a step function that has no usable gradient, they used an arctangent surrogate gradient during backpropagation instead of the triangular approximation used in SpikGRU.
Training used a single recurrent layer of 128 units followed by a readout layer with self-recurrence, a max-over-time loss on the readout activations, backpropagation through time, and the Adam optimizer with a learning rate of 0.001, 100 epochs and batch size 128. This setup covers N-TIDIGITS (2,464 training / 2,486 test samples, 64 channels, 250 timesteps, trimmed to 1.25 seconds) and SHD (8,156 training / 2,264 test samples, 700 channels, 20 classes, 250 timesteps). For the spatio-temporal experiments, MNIST images were stretched over 10 timesteps via rate coding, DVS128Gesture streams were split into 10 timesteps, and CIFAR10-DVS samples were organized into 10 timesteps, with training extended to 300 epochs. Inputs were reshaped to C×H×W — DVS Gesture to 1×8×8 and SHD to 7×10×10 — with 2×2 max-pooling applied to DVS Gesture and a 3×63 convolution used to reduce SHD's spatial dimensions. The spiking threshold was set to v_th = 1.
Why This Matters
The work argues that spiking networks are attractive for low-resource settings such as edge and wearable devices because they are event-driven, and that existing spiking GRU designs lose local detail that matters when spatial relations drive the prediction. CS-GRU is positioned as a drop-in replacement for conventional neuron models like LIF and Cuba-LIF inside larger architectures.
- Gesture recognition: The DVSGesture results (82% for CS-GRU vs. 66% for SpikGRU) target event-based hand and arm gesture streams from dynamic vision sensors, relevant to human-computer interaction and control.
- Audio and speech processing: The N-TIDIGITS and SHD experiments address spoken digit recognition from spiking cochlea sensors and audio-to-spike mappings, applicable to always-on keyword or digit spotting.
- Event-based vision and video: CIFAR10-DVS and the general framing around video analysis point to neuromorphic camera pipelines where frames are sparse events rather than dense images.
- Sensor and wearable data: The stated motivation includes biological signals and power-constrained devices, where the measured 69.05% reduction in spiking activity relative to SpikGRU translates into fewer redundant computations, longer battery life and less heat generation on neuromorphic hardware.
Industry relevance: The combination of competitive accuracy on standard benchmarks with a directly reported energy proxy is aimed at deployers of neuromorphic and ultra-low-power inference hardware, as well as at toolchains that need efficient recurrent building blocks for streaming sensor input.
Future Directions
- Scaling CS-GRU into larger networks: The authors explicitly suggest embedding the cell in ensembles, ResNets and VGG-style spiking architectures, and replacing conventional LIF or Cuba-LIF neurons within them.
- Broader task coverage: Stated open directions include speech recognition, video processing and general sensor data processing beyond the five benchmarks evaluated here.
- Explaining the unstable combinations: Several partial combinations of mod1–mod4 degrade sharply (for example 8.43% and 17.30% on SHD and N-TIDIGITS respectively), which raises the question of why all four modifications are needed together rather than any subset.
- Validation on physical neuromorphic hardware: Energy efficiency is reported only through the spiking activity rate proxy on the DVS Gesture dataset; the paper does not report measured power on neuromorphic chips, nor efficiency on the other datasets.
Target Audience
Researchers and engineers working on spiking neural networks, neuromorphic computing, and efficient sequential models will get the most from this paper, particularly those already familiar with GRU gating and LIF-style neuron dynamics. It is also relevant to practitioners building event-based vision or audio pipelines who need a recurrent cell that handles local spatial structure, and to readers interested in the trade-off between classification accuracy and spiking activity as a proxy for energy use.
Authors’ abstract
Spike-based temporal messaging enables SNNs to efficiently process both purely temporal and spatio-temporal time-series or event-driven data. Combining SNNs with Gated Recurrent Units (GRUs), a variant of recurrent neural networks, gives rise to a robust framework for sequential data processing; however, traditional RNNs often lose local details when handling long sequences. Previous approaches, such as SpikGRU, fail to capture fine-grained local dependencies in event-based spatio-temporal data. In this paper, we introduce the Convolutional Spiking GRU (CS-GRU) cell, which leverages convolutional operations to preserve local structure and dependencies while integrating the temporal precision of spiking neurons with the efficient gating mechanisms of GRUs. This versatile architecture excels on both temporal datasets (NTIDIGITS, SHD) and spatio-temporal benchmarks (MNIST, DVSGesture, CIFAR10DVS). Our experiments show that CS-GRU outperforms state-of-the-art GRU variants by an average of 4.35%, achieving over 90% accuracy on sequential tasks and up to 99.31% on MNIST. It is worth noting that our solution achieves 69% higher efficiency compared to SpikGRU. The code is available at: https://github.com/YesmineAbdennadher/CS-GRU.