Research
Energy-based Autoregressive Generation for Neural Population Dynamics
Energy-based Autoregressive Generation for Neural Population Dynamics Authors: Ningling Ge, Sicheng Dai, Yu Zhu, Shan Yu Affiliations: Institute of Automation, Chinese Academy of Sciences; School of A

- arXiv
- 2511.17606
- Published
- 2025-11-18
- Authors
- Ningling Ge, Sicheng Dai, Yu Zhu, Shan Yu
AI summary
Energy-based Autoregressive Generation for Neural Population DynamicsAuthors: Ningling Ge, Sicheng Dai, Yu Zhu, Shan Yu Affiliations: Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology; Beijing Academy of Artificial Intelligence arXiv: 2511.17606v1 [cs.LG], 18 Nov 2025
Overview
- Research area: Computational neuroscience and generative machine learning — modeling the spiking activity of neural populations, evaluated against the Neural Latents Benchmark.
- Technical level: Advanced. The paper assumes familiarity with energy-based models, strictly proper scoring rules, masked autoregressive transformers, latent diffusion, and variational autoencoders.
- Scope: The paper introduces an Energy-based Autoregressive Generation (EAG) framework that learns to generate realistic neural spike trains in a learned latent space, benchmarked on one synthetic Lorenz dataset and two real neural datasets (MC_Maze and Area2_bump).
What This Paper Is About
Modeling how populations of neurons fire is central to neuroscience, but existing generative approaches face a trade-off: VAE-based methods are fast but capture complex population and single-neuron statistics poorly, while diffusion-based methods model variability well but require costly iterative sampling. The paper's goal is to build a generative model of neural spiking that is both high-fidelity and computationally cheap by using energy-based learning in a latent space instead of iterative denoising. The authors also test whether the generated data is useful for practical tasks, specifically generalizing to unseen behavioral conditions and improving brain-computer interface decoding.
Key Contributions
- A new EAG framework that is described as resolving the trade-off between computational efficiency and high-quality neural population modeling through energy-based learning in latent space.
- State-of-the-art generation quality with substantial efficiency gains: the paper reports SOTA results across four spiking-statistic metrics on MC_Maze and Area2_bump, alongside large reductions in sampling latency relative to diffusion-based LDNS.
- Demonstration that conditional generation generalizes to unseen behavioral contexts — reach direction labels and per-timepoint hand velocity labels never seen in training — while preserving trial-to-trial variability.
- Demonstration that EAG-generated synthetic neural data improves downstream BCI decoding accuracy across multiple decoder models, with the largest reported gain reaching 12.1% on MC_Maze using the Neural Data Transformer (NDT).
The paper also states that it introduces energy-based models to neural computational modeling "for the first time." Code is released at https://github.com/NinglingGe/Energy-based-Autoregressive-Generation-for-Neural-Population-Dynamics.
Main Findings
- Best overall generation quality on real neural data: On MC_Maze, EAG reports D_KL of 0.0014 ± 2.0e-4 for the population spike count histogram, pairwise correlation RMSE of 0.0024 ± 1.0e-5, mean ISI RMSE of 0.024 ± 0.001, and std ISI RMSE of 0.018 ± 0.0024 — the best values among all listed methods. On Area2_Bump, EAG reports 0.0018 ± 1.6e-4, 0.0075 ± 9.1e-5, 0.035 ± 0.004, and 0.025 ± 0.003 respectively.
- Significance testing: The paper states EAG outperforms all baselines (Wilcoxon, p < 0.001) except for pairwise correlation versus LDNS, which is not significant (p = 0.09).
- Large latency reduction: EAG-32 generates 2008 trials in 10.29 s, while LDNS-1000 requires 330.64 s, which the paper reports as a 96.9% speed-up. Against the minimal-step LDNS-200, EAG-32 still achieves an 84.4% latency reduction. The Results introduction separately describes a "30× higher sampling efficiency" than diffusion models.
- Quality at low step counts: Even minimal-step EAG-16 is reported to outperform LDNS-1000. EAG-32 delivers a 49.0% gain over LDNS-200 and a 32.4% gain over LDNS-1000 on RMSE mean ISI.
- The energy loss matters: A plain autoregressive Transformer trained with MSE loss performs significantly worse than EAG (Supp. Table S1), and EAG still wins after baselines are augmented with spike-history inputs.
- Strict propriety matters: Varying the energy-score exponent α shows models with α = 2.0 degrade sharply (D_KL 0.0541 ± 7.8e-4, mean ISI RMSE 0.051 ± 0.004), while α in [1, 2) performs consistently well. α = 1.0 is the default. The authors note α < 1 can cause early-training instability from unbounded gradients, so they focus on α in [1, 2).
- Generalization to unseen labels: Conditioned on novel reach directions (~60°, unseen in training), EAG-sampled rates align closely with real rates and show natural trial-to-trial variability rather than copying training data. Under velocity conditioning on an entirely unseen trajectory, decoding from EAG-sampled rates reaches R² = 0.89 versus R² = 0.65 for LDNS.
- BCI decoding improvement on MC_Maze: With one-fold EAG augmentation, co-smoothing bps improved for GRU (+1.1%), SLDS (+6.9%), LFADS (+7.1%), NDT (+12.1%), AutoLFADS (+0.9%), and NDT(ray) (+8.4%). Gains persist but shrink after hyperparameter optimization, which the authors present as evidence the benefit is not a training artifact.
- Scaling augmentation helps most on small data: On the small Area2_Bump dataset, multi-scale augmentation produced a 54.7% co-smoothing improvement at 2× with NDT, and a 51.9% improvement at 4×; PSTH R² gains reached 31.5% (2×) and 32.4% (4×).
- Runtime and memory scale well: The paper reports that EAG's runtime and memory usage remain largely unaffected by increases in neuron count or time length.
Methodology in Plain English
EAG uses a two-stage pipeline:
-
Learn a compact representation of spikes. Following the LDNS approach, an autoencoder maps high-dimensional spike trains into a low-dimensional latent space, using a Poisson observation model with temporal smoothness constraints. The authors state they use identical architecture and training configuration as LDNS for this stage so the comparison is fair. Their own decoder is deeper than LDNS's, with more S4 blocks, to match the energy transformer's capacity.
-
Generate new latents with an energy-based autoregressive transformer. Rather than denoising step by step, the model is trained to output a distribution of latents at each masked time point. It does this with the energy score, a strictly proper scoring rule that needs no explicit likelihood calculation. The training loss compares two independent samples from the model against the real latent: it pulls both samples toward the data while also rewarding separation between them, which is what preserves realistic trial-to-trial variability.
The generative architecture borrows from masked autoencoding: during training, random masking ratios are drawn uniformly from [0.7, 1.0], the encoder sees only unmasked positions, and the decoder fills in mask tokens. During inference, the masking ratio decreases from 1.0 to 0 along a cosine schedule, so generation is progressive rather than one-shot. Stochasticity comes from a noise vector sampled uniformly from [-0.5, 0.5] and injected into an MLP generator through adaptive layer normalization (shift, scale, and gate terms), which is what makes the output stochastic rather than deterministic.
For conditional generation, behavioral variables (initial reach angle, encoded as cosine/sine, or hand velocity along two channels) are projected and concatenated with the latents. In 10% of training trials the condition is replaced with a learnable null token, enabling classifier-free guidance at inference via a weighted combination of conditional and unconditional outputs.
Evaluation setup: the synthetic dataset comes from a 3D Lorenz system producing 128-dimensional neural spiking over 256 time steps. MC_Maze is a delayed center-out reaching task recorded from premotor and primary motor cortex; Area2_Bump is a small dataset of roughly 300 trials from somatosensory cortex during a bump-perturbed reaching task. Baselines are TNDM, pi-VAE, and AutoLFADS (VAE-based) plus LDNS (diffusion-based). Metrics follow the LDNS study: population spike count distribution (D_KL), pairwise spike-count correlation RMSE, mean ISI RMSE, and ISI standard deviation RMSE. All experiments ran on NVIDIA A40 GPUs with 40 GB of memory, using two GPUs for Lorenz and MC_Maze and one for Area2_Bump.
Why This Matters
Impact on research: Encoding models of neural activity have been underexplored relative to decoders. EAG argues that high-fidelity, cheap neural generation can fill that gap, providing a tool for studying low-dimensional manifold dynamics and trial-to-trial variability that deterministic predictive models cannot capture. The computational savings matter methodologically: iterative diffusion sampling is expensive enough to limit experimental scale, and the paper reports a 96.9% latency reduction relative to LDNS-1000.
Real-world applications named or motivated in the paper:
- Motor brain-computer interfaces for paralysis, where synthetic spike data augment decoder training when recorded data is scarce.
- Therapeutic interventions for neurological disorders, with Parkinson's disease cited as an example, where better models of population dynamics could inform intervention design.
- Neural data augmentation for low-data regimes, demonstrated concretely on Area2_Bump, where the paper reports gains up to 54.7% co-smoothing improvement with 2× augmentation.
- Closed-loop neuroengineering experiments, where conditionally generated activity from hypothetical movements can substitute for or supplement real recordings, provided trial-to-trial variability is preserved.
Industry relevance: For BCI companies and neurotechnology groups, the appeal is a generator that can synthesize realistic neural data at a fraction of the compute — relevant when data collection is expensive, when only a few hundred trials per subject are available, and when latency budgets matter. Releasing code alongside the paper lowers the barrier to trying the approach. The paper does not report commercial partnerships or deployment results; those are not reported.
Future Directions
- Pushing augmentation scale further. The authors hypothesize that for small datasets, data scarcity is the more pressing bottleneck and that larger scale augmentation brings larger gains. They tested 1×, 2×, and 4×; whether gains continue beyond 4× is not reported.
- Broadening the α range and understanding early-training instability. The paper restricts experiments to α in [1, 2) because α < 1 produced unstable gradients early in training. A principled fix could extend the usable range of this strictly proper scoring rule.
- Extending conditional generation to more behavioral and cognitive variables. The paper demonstrates reach direction and hand velocity. Whether the same framework generalizes to cognitive states, sensory variables, or other task structures is not established.
- Testing generality across species, brain areas, and recording modalities. Evaluation covers macaque motor and somatosensory cortex plus a synthetic Lorenz system. Generalization to other cortical regions, non-motor tasks, or human recordings is not reported.
Target Audience
This paper is best suited to machine learning researchers working on generative models for scientific data, computational neuroscientists studying population dynamics and latent manifold structure, and BCI engineers who need synthetic neural data for decoder training. Readers need comfort with energy-based models, proper scoring rules, and transformer architectures to follow the methods section, though the experimental results and application claims are accessible to a broader neuroscience audience.
Authors’ abstract
Understanding brain function represents a fundamental goal in neuroscience, with critical implications for therapeutic interventions and neural engineering applications. Computational modeling provides a quantitative framework for accelerating this understanding, but faces a fundamental trade-off between computational efficiency and high-fidelity modeling. To address this limitation, we introduce a novel Energy-based Autoregressive Generation (EAG) framework that employs an energy-based transformer learning temporal dynamics in latent space through strictly proper scoring rules, enabling efficient generation with realistic population and single-neuron spiking statistics. Evaluation on synthetic Lorenz datasets and two Neural Latents Benchmark datasets (MC_Maze and Area2_bump) demonstrates that EAG achieves state-of-the-art generation quality with substantial computational efficiency improvements, particularly over diffusion-based methods. Beyond optimal performance, conditional generation applications show two capabilities: generalizing to unseen behavioral contexts and improving motor brain-computer interface decoding accuracy using synthetic neural data. These results demonstrate the effectiveness of energy-based modeling for neural population dynamics with applications in neuroscience research and neural engineering. Code is available at https://github.com/NinglingGe/Energy-based-Autoregressive-Generation-for-Neural-Population-Dynamics.