Research
ASGMamba: Adaptive Spectral Gating Mamba for Multivariate Time Series Forecasting
Overview Research area: Long-term multivariate time series forecasting (LTSF) with State Space Models (SSMs), frequency-domain filtering, and efficient sequence modeling for high-performance computing

- arXiv
- 2602.01668
- Published
- 2026-02-02
- Authors
- Qianyang Li, Xingjun Zhang, Shaoxun Wang, Jia Wei, Yueqi Xing
AI summary
Overview
Research area: Long-term multivariate time series forecasting (LTSF) with State Space Models (SSMs), frequency-domain filtering, and efficient sequence modeling for high-performance computing environments.
Technical level: Advanced. The paper assumes familiarity with state space models, the Mamba selective scan, FFT-based spectral analysis, patching strategies, and standard LTSF benchmark protocols.
Scope: The paper proposes ASGMamba, a linear-complexity forecasting architecture that injects a local spectral gating prior into a Mamba backbone and combines it with hierarchical multi-scale patching and learnable variable (node) embeddings, evaluated on nine public benchmarks.
What This Paper Is About
Transformer-based forecasters model long-range dependencies well but scale quadratically with sequence length, incurring a heavy memory footprint that can saturate GPU bandwidth on very long inputs. Linear SSMs such as Mamba avoid this cost, but the paper argues that their limited latent state gets saturated by high-frequency noise when applied to noisy sensor data, because distinguishing signal from noise in the time domain is expensive. ASGMamba's goal is to let a linear-complexity Mamba model spend its state capacity on genuine temporal structure by filtering each local patch based on its own spectral energy, while also restoring variable-specific semantics that channel-independent designs discard.
Key Contributions
- Spectral-conditioned state evolution framework: ASGMamba conditions the SSM input on local spectral energy density, so high-frequency noise does not contaminate the latent state, which the authors describe as increasing the model's effective capacity for trend modeling.
- Adaptive Spectral Gating (ASG) module: A learnable, input-dependent filter that modulates input fidelity according to frequency properties, acting as a computation-efficient prior that separates signal from noise within a strict O(L) budget.
- Hierarchical multi-scale system design with Node Embeddings: A multi-branch architecture (patch sizes 8, 16, 32) augmented with learnable Node Embeddings, allowing the shared backbone to capture multi-granularity temporal patterns and the distinct physical semantics of each variable at once.
- Efficiency-accuracy trade-off demonstration: Evaluations on nine real-world benchmarks, reported as competitive or superior to state-of-the-art Transformer and SSM baselines, with reduced memory footprint and inference latency in long-sequence scenarios.
Main Findings
- Spectral gating as noise suppression: Patches dominated by high-frequency energy produce gate values approaching 0, while trend-rich patches are preserved with gates approaching 1; attenuating the gated input reduces the magnitude of the input projection and prevents the recurrent state from updating on spurious fluctuations.
- Strict linear complexity: Because FFT is applied to fixed-size local patches, the gating cost is N × O(P log P), which the paper derives as approximately (L/S) · P log P; since P and S are small constants, complexity remains O(L), in contrast to the O(L²) of attention and the O(L log L) of global Fourier methods such as FEDformer.
- Reported accuracy on benchmarks (from the visible portion of Table 3, input length L = 96): ASGMamba records the lowest average MSE among the listed models on Weather (0.244 MSE, 0.270 MAE) and Exchange (0.351 MSE, 0.393 MAE). Other listed models record lower average MSE on some datasets: TimeMixer on Solar (0.216 vs. ASGMamba's 0.231), S-Mamba on Electricity (0.170 vs. 0.172) and on Traffic (0.414 vs. 0.478). The table is truncated in the supplied content, so a full dataset-by-dataset verdict is not available here.
- Long-horizon efficiency claim: The abstract states that ASGMamba significantly reduces memory usage on long-horizon tasks while keeping O(L) complexity, positioning it as scalable for high-throughput forecasting in resource-limited environments. Specific memory or latency figures are not reported in the supplied content.
- Multi-scale fusion is learnable: Rather than fixing a single temporal resolution, the model uses a learnable convex combination (softmax over three weights) to prioritize whichever patch scale best matches the data's frequency characteristics.
- Ablation and sensitivity studies are described but not shown: The experiments section states that ablative analysis of each component and sensitivity analysis of critical hyperparameters were conducted, but their results do not appear in the supplied content.
Methodology in Plain English
The model begins by normalizing each input series with Reversible Instance Normalization (RevIN) and then treats each variable independently (a Channel-Independent strategy), reshaping the batch to (B · N) × L × 1 — this reduces parameter complexity from O(N²) to O(1) but removes cross-variable context. To recover some of that lost identity, the authors add two learnable embeddings to the patched input: a positional embedding for temporal order and a Node Embedding that acts as a static semantic descriptor for each variable, so the shared backbone can adapt to, for example, the different periodicities of voltage versus load.
Inputs are cut into overlapping patches at three sizes (8, 16, 32) with 50% overlap; overlap is used because hard patch boundaries create spectral leakage when an FFT is applied. Each patch then passes through the Adaptive Spectral Gating module: a real-valued FFT is computed on the patch, its power spectrum is aggregated into three coarse bands (low, mid, high relative to the Nyquist frequency), and a small two-layer MLP with a bottleneck of D/4 and a Sigmoid output turns that 3-dimensional spectral descriptor into a gate. The gate multiplies the patch embedding (after LayerNorm), so noise-dominated patches are suppressed before entering the Mamba encoder; the Mamba block output is added back via a residual connection. The three scale branches each produce a forecast, and a learnable softmax-weighted sum fuses them before the inverse RevIN is applied.
Training uses the Adam optimizer, a maximum of 10 epochs, L2 weight decay, dropout, and early stopping with patience 5, with all results averaged over five independent runs on a single NVIDIA RTX 4090 (24GB) GPU.
Why This Matters
Impact on research. The paper frames noise handling in SSMs as a state efficiency problem rather than an accuracy problem: rather than forcing a recurrent system to learn global spectral filtering from scratch in the time domain, it supplies a cheap, local spectral prior. If the approach generalizes, it suggests that lightweight, patch-local frequency cues are a broadly useful interface between classical signal processing and modern linear recurrent models, and it offers an alternative to global FFT designs that break streaming prediction.
Real-world applications (as named in the paper):
- Real-time energy grid management and electricity consumption forecasting.
- Large-scale traffic flow simulation and traffic volume forecasting.
- High-frequency financial market decision-making.
- Deployment of forecasting on high-performance computing and resource-constrained supercomputing infrastructures handling large sensor streams.
Industry relevance. The motivation is explicitly operational: in HPC and grid settings, a model is judged not only on statistical accuracy but on computational efficiency, memory footprint, and latency. A backbone with O(L) scaling and no attention map means ultra-long sequences (the paper cites L > 10³) can be processed without the quadratic memory matrix that bottlenecks GPU bandwidth, which matters for throughput-oriented forecasting services.
Future Directions
- Complete and quantify the efficiency story: The paper claims reduced memory footprint and inference latency on long-horizon tasks, but specific memory, throughput, or latency numbers are not reported in the supplied content; a systematic efficiency benchmark against the same baselines would substantiate the core selling point.
- Full ablation and sensitivity results: The experimental plan promises ablations for each component and sensitivity analysis over critical hyperparameters, but those results are absent from the supplied content; isolating the contribution of the ASG gate, the Node Embeddings, and the multi-scale fusion would clarify where the gains come from.
- Reconciling the channel-independence trade-off: Node Embeddings are a lightweight substitute for cross-channel modeling; whether they can be extended toward richer variable interactions without reintroducing quadratic channel mixing is an open question the paper raises but does not resolve.
- Beyond the tested regimes: The evaluation fixes a look-back window of L = 96 and horizons of 96, 192, 336, and 720 across nine datasets. Behavior at substantially longer look-backs, on streaming/online settings, and on domains outside energy, traffic, weather, economics, and solar is not reported.
Target Audience
Researchers and practitioners working on efficient time series forecasting, particularly those interested in state space models, Mamba-style architectures, and spectral or frequency-domain methods. It is also relevant to engineers deploying forecasting models under memory and latency constraints in energy, traffic, weather, finance, and HPC operations, and to readers following the trade-off between attention-based and linear-complexity sequence models. A background in deep learning and basic signal processing is assumed.
Authors’ abstract
Long-term multivariate time series forecasting (LTSF) plays a crucial role in various high-performance computing applications, including real-time energy grid management and large-scale traffic flow simulation. However, existing solutions face a dilemma: Transformer-based models suffer from quadratic complexity, limiting their scalability on long sequences, while linear State Space Models (SSMs) often struggle to distinguish valuable signals from high-frequency noise, leading to wasted state capacity. To bridge this gap, we propose ASGMamba, an efficient forecasting framework designed for resource-constrained supercomputing environments. ASGMamba integrates a lightweight Adaptive Spectral Gating (ASG) mechanism that dynamically filters noise based on local spectral energy, enabling the Mamba backbone to focus its state evolution on robust temporal dynamics. Furthermore, we introduce a hierarchical multi-scale architecture with variable-specific Node Embeddings to capture diverse physical characteristics. Extensive experiments on nine benchmarks demonstrate that ASGMamba achieves state-of-the-art accuracy. While keeping strictly $$\mathcal{O}(L)$$ complexity we significantly reduce the memory usage on long-horizon tasks, thus establishing ASGMamba as a scalable solution for high-throughput forecasting in resource limited environments.The code is available at https://github.com/hit636/ASGMamba