Research
Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modelling and State Tracking
Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modelling and State Tracking Overview Research area: Efficient sequence modelling for language, spanning state-space models

- arXiv
- 2602.10743
- Published
- 2026-02-11
- Authors
- Vaisakh Shaj, Cameron Barker, Aidan Scannell, Andras Szecsenyi, Elliot J. Crowley, Amos Storkey
AI summary
Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modelling and State TrackingOverview
Research area: Efficient sequence modelling for language, spanning state-space models (SSMs), gated linear attention (GLA), Bayesian filtering / Kalman filtering, and parallel scan algorithms.
Technical level: Advanced. The paper derives its central result through Gaussian state-space models, information-form parameterisations, Möbius (fractional-linear) transforms and associative prefix scans, and it assumes familiarity with Mamba, linear attention and the Kalman filter.
Scope: The paper proposes Kalman Linear Attention (KLA), a drop-in sequence-mixing layer that performs exact Kalman filtering in a temporally parallel form, and evaluates it against modern SSMs and GLAs on state-tracking, associative recall and zero-shot commonsense reasoning.
What This Paper Is About
State-space language models such as Mamba and gated linear attention offer linear-complexity, parallelisable alternatives to transformers, but their hidden-state updates are linear or affine, which limits their expressivity and their ability to track state reliably. Classical Kalman filters give principled state and uncertainty estimates but are viewed as inherently sequential, so they have not been used as scalable sequence mixers. The paper's goal is to close this gap by showing that the Kalman filter's updates can be rewritten so that they compose associatively and can be run as a parallel scan, yielding a layer that is strictly more expressive than GLA-style linear updates at the same computational cost and that carries an explicit belief-state uncertainty.
Key Contributions
-
Associative reparameterisation of Kalman filtering (C1). The authors reparameterise the diagonal linear-Gaussian filter in information form and show that the precision recursion is a Möbius (fractional-linear) map that composes associatively, which enables parallel prefix scans. They argue no augmentation is needed relative to prior work: Kalman updates in information form are already Möbius maps, so precision composes by 2×2 matrix multiplication.
-
Nonlinear gating from uncertainty (C2). The precision-ratio gates are history-dependent and nonlinear, going beyond linear/affine SSM and gated linear-attention updates while preserving linear-time scan structure. The same factor reappears as the forget gate in the mean recursion, tying the precision and mean tracks together.
-
The Kalman Linear Attention layer (C3). KLA is introduced as a drop-in sequence mixer for modern language-modelling pipelines that produces explicit belief-state uncertainty. It follows Mamba's fused-MLP block design with the Kalman filter as the mixer.
-
Scaling and empirical validation (C4). The scan-based implementation scales efficiently with sequence length and matches or outperforms modern SSMs and GLAs on state-tracking, associative recall and zero-shot commonsense reasoning. The authors state it is among the first stacked Bayesian filters pretrained at the billion-token scale.
Main Findings
-
Expressivity on permutation composition: KLA solves the A5 (alternating group on 5 elements) permutation-composition task at constant depth (1–2 blocks), where linear SSMs and attention require depth growing with sequence length. Its fractional-linear (Möbius) updates are positioned between a fully nonlinear RNN and linear SSMs/transformers, while remaining parallel-trainable.
-
Efficiency profile: Per Table 1, KLA has 𝒪(T) parallel training cost (versus 𝒪(T²) for softmax attention), 𝒪(1) inference cost, supports parallel training, and is the only one of the three primitives listed that carries sequence uncertainty.
-
Nonlinearity relative to GLA: KLA's per-token recurrent update is nonlinear (a Möbius/precision recursion) yet remains temporally parallel, and the authors describe it as strictly more expressive than GLA-style linear updates at the same computational cost.
-
Parallel depth: The information-form derivation yields 𝒪(T) work and 𝒪(log T) parallel depth (Corollaries 1.1 and 2.1), the same parallel depth the paper attributes to models such as Mamba.
-
Online-learning view: The moment-form posterior mean collapses the predict/update pair into a single gated recurrence (Corollary 2.2), which is the exact minimiser of a precision-weighted least-squares objective with a proximal term and a fit term.
-
Downstream scaling: In Figure 1(b), replacing a single final attention layer with KLA (GPT+KLA) gives the strongest complement to attention, evaluated by zero-shot accuracy averaged over eight commonsense benchmarks at 45M and 180M parameters. The exact accuracy values are not reported in the provided content.
-
Ablation of OU dynamics: Ablating the Ornstein-Uhlenbeck prior dynamics and discretisation shows that OU discretisation improves accuracy and learning stability, especially for deeper models.
-
Special case: Under deterministic (p_t = 0) and linear time-invariant settings, the KLA updates reduce to convolutions computable in 𝒪(log T) time via FFT (Theorem 3, in the Appendix).
-
Positioning against prior parallel filtering: The authors state that Särkkä and García-Fernández (2020) required a 5-tuple augmented representation in the linear case to achieve associativity, whereas KLA needs only 2×2 matrix multiplication. Closest layers, MesaNet and Gated KalmaNet, solve a similar regularised least-squares objective but are described as assuming a static, deterministic latent state (no transition dynamics) and requiring expensive iterative solvers (conjugate gradient and Chebyshev respectively).
Methodology in Plain English
The work recasts sequence mixing as Bayesian state estimation rather than deterministic state updating. Instead of treating each token as a control signal that pushes a hidden state forward, KLA treats each token as noisy evidence about an unobserved latent state.
Two sources of uncertainty are modelled: process noise, capturing uncertainty in how the state evolves, and observation noise, capturing uncertainty in the information each token provides. Crucially, this uncertainty is not just an output of the model — it directly controls how new information is gated into the state.
For the dynamics prior, the authors use a continuous-time Ornstein-Uhlenbeck process, described as the canonical mean-reverting diffusion and the continuous-time analogue of a stable AR(1), then discretise it exactly. This yields a Gaussian transition in which the discretised decay acts as a forget factor and the process-noise variance is coupled by construction to that decay, so the same parameters that control decay also determine how uncertainty accumulates between observations.
Each token supplies a (noisy) observation of the latent state through an observation operator, with a value precision representing confidence in the token evidence; keys and values parameterise this likelihood, while the query parameterises a readout applied after inference. In the deterministic-readout limit (zero readout noise), the output is simply the query-conditioned projection of the posterior mean.
The central technical move is to work in information form — precisions and natural parameters — because incorporating token evidence then amounts to adding canonical parameters, and the predict transformation takes a structured form. That structure turns out to be a Möbius transform for the precision and an affine transform for the information mean, both of which compose associatively across time and can therefore be computed with parallel prefix scans. The queries are then read out from the resulting belief state. The paper fixes the state-expansion factor to N = 1 in the main exposition and presents the diagonal (per-channel) filter, with the state-expanded (N > 1) form given in a table and algorithm; the covariance/precision remains diagonal throughout, and the N × D matrix arises solely from state expansion.
Why This Matters
Impact on research. The paper argues that exact Bayesian filtering is itself scan-parallelisable, with no approximations, steady-state assumptions or test-time solvers. If that holds, it expands the class of probabilistic primitives available to sequence modellers and provides a probabilistic bridge between attention-like mixers and classical filtering. It also claims a concrete expressivity separation: a linear-cost mixer that solves A5 permutation composition at constant depth where linear SSMs and attention require depth growing with sequence length. The paper states it is among the first stacked Bayesian-filtering primitives trained at the billion-token scale, and unlike recent mixers it admits a strictly more expressive Möbius update, which the authors say translates into measurable gains on state tracking and commonsense reasoning.
Potential real-world applications (the paper's stated motivations are long contexts, on-device deployment and energy efficiency; it does not enumerate applications itself, so these follow from those motivations):
- Long-context language modelling, where quadratic attention is costly.
- On-device or resource-constrained deployment, given 𝒪(1) inference cost and linear or sublinear memory.
- State tracking and sequence-manipulation tasks where reliable tracking of composed state matters.
- Uncertainty-aware systems, since KLA exposes an explicit belief-state uncertainty that SSMs and attention do not.
Industry relevance. The layer is presented as a drop-in replacement for standard SSM or attention layers within existing pipelines, and the paper's GPT+KLA experiments replace a single attention layer in a GPT-style stack. That makes it plausible as a component-level substitution rather than a full architectural rewrite. The explicit uncertainty signal is also relevant to applications where knowing when a model is uncertain matters, and the energy-efficiency framing connects to the cost pressures driving sub-quadratic mixer research.
Future Directions
-
Extending beyond the diagonal filter. KLA's covariance/precision remains diagonal throughout, and the main exposition fixes the state-expansion factor to N = 1. Whether richer covariance structures can retain the Möbius/associative property and the scan-parallelism is left open by this framing.
-
Better use of the uncertainty signal. The paper notes that uncertainty directly controls gating, but treats the readout in the deterministic limit where output precision goes to infinity. Reintroducing a finite output precision, or otherwise exploiting the uncertainty estimate downstream, is a natural extension.
-
Positioning against iterative-solver mixers. MesaNet and Gated KalmaNet solve a similar regularised least-squares objective but with different assumptions and solvers. An open question is whether KLA's closed-form transition-and-process-noise formulation can absorb the settings those layers target.
-
Scaling and architecture placement. The scaling study spans 45M and 180M parameters with a single replaced attention layer, and the authors describe billion-token-scale pretraining as a first. Where KLA layers are best placed in a deeper stack, and how that interacts with the observed stability benefit of OU discretisation at depth, are open questions.
Target Audience
Researchers and engineers working on efficient sequence architectures, particularly those familiar with Mamba, gated linear attention and linear-attention variants. It is also relevant to readers from probabilistic modelling, control and robotics backgrounds interested in whether Kalman filtering can be scaled into a competitive language-modelling primitive, and to those working on state tracking and expressivity limits of sub-quadratic mixers. The mathematical density and reliance on information-form Gaussian derivations make it a poor fit for beginners, despite the plain-language analogies the authors use (for example, comparing latent beliefs to image sensors feeding proprioceptive outputs).
Authors’ abstract
State-space language models such as Mamba and gated linear attention (GLA) offer linear-complexity, parallelisable alternatives to transformers, but their linear state updates limit expressivity and robust state tracking. We close this gap from a probabilistic angle, casting sequence mixing as exact Bayesian filtering with the Kalman filter as the core primitive. Classical Kalman filters give principled state and uncertainty estimates but are viewed as inherently sequential; we show that reparameterising them in information form turns their updates into an associative scan - so the per-token recurrent update is non-linear (a Möbius/precision recursion) yet remains temporally parallel. The resulting Kalman Linear Attention (KLA) layer is a drop-in sequence mixer that performs time-parallel probabilistic inference, carries an explicit belief-state uncertainty, and is strictly more expressive than GLA-style linear updates at the same computational cost. This expressivity translates directly into stronger state tracking: KLA solves permutation-composition ($A_5$) tasks that linear SSMs and attention cannot, while staying scan-parallel. As a drop-in primitive it also matches or improves on modern SSMs and GLAs across synthetic token-manipulation and zero-shot commonsense benchmarks, and is among the first stacked Bayesian-filtering primitives trained at the billion-token scale.