Research
HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models
HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models Overview Research area: Computer Vision — autoregressive video generation and video world models, specifically memory mechanisms for

- arXiv
- 2610.05739
- Published
- 2026-10-05
- Authors
- Zhuokun Chen, Feng Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
AI summary
HLA-WM: Hybrid Linear Attention for Long-Horizon Video World ModelsOverview
Research area: Computer Vision — autoregressive video generation and video world models, specifically memory mechanisms for long-horizon rollout (Gated DeltaNet / recurrent linear attention).
Technical level: Advanced. The paper assumes familiarity with attention, KV caching, linear-attention recurrences, and camera-based video generation pipelines.
Scope: The paper diagnoses long-range forgetting in Gated DeltaNet (GDN)-based video world models and proposes HLA-WM, a training-free hybrid memory framework that retrieves geometry-relevant historical chunks and recomposes them into a query-specific recurrent state.
What This Paper Is About
Video world models must remember scenes they observed earlier so that when the camera returns to a previously seen location, the model regenerates the same structures rather than a merely similar-looking view. Recurrent linear attention (GDN) keeps memory compact by squeezing all history into a single fixed-size state, but the authors show this causes distant scene information to be progressively attenuated by every intervening update. HLA-WM's goal is to let a pretrained recurrent video world model recover distant scene history without a full token-level KV cache and without any additional training.
Key Contributions
-
Empirical characterization of long-range forgetting in GDN. The authors formalize the chunk-level recurrence $S_i = S_{i-1}A_i + B_i$, define a cumulative retention measure $W_{i\rightarrow j} = \text{AvgNorm}(A_{i+1}A_{i+2}\cdots A_j)$, and show on SANA-WM-Bench that information from Chunk 3 retains a value of only 0.0416 by Chunk 35, even though the camera has returned to the same region.
-
HLA-WM, a training-free hybrid linear-attention framework. It caches compact chunk-wise affine summaries $(A_i, B_i)$ of the GDN transition, retrieves scene-relevant historical chunks using camera geometry, and recomposes them into a query-specific recurrent state. All pretrained parameters remain unchanged.
-
Geometry-guided retrieval and selective state recomposition. A symmetric frustum-overlap score $R(i,q)$ built from poses, intrinsics, and a scene-level median-depth estimate addresses memory without a learned router, and an associative composition rule recombines the selected summaries without replaying historical visual tokens.
-
Broad empirical validation. On the 60-second SANA-WM-Bench, HLA-WM improves all six aggregate revisit-consistency and camera-control metrics of the base autoregressive generator, and generalizes to MBench-A across all four subsets and all evaluated inference modes over 547 samples, at 12× lower historical-state memory than full KV caching.
Main Findings
-
Distant memory is attenuated by unrelated updates. On a representative trajectory, the camera leaves the region observed at Chunk 3, loops, and returns near Chunk 35. Despite high overlap, Chunk 3's retained influence drops to 0.0416 at Chunk 35, and the generated scene deviates substantially from the earlier observation. Most intermediate chunks observe different viewpoints and are only weakly related, yet their transition matrices still transform the earlier memory.
-
Stage-1 gains on SANA-WM-Bench (Hard split). PSNR rises from 9.37 to 10.11 (+0.74 dB), SSIM from 0.1774 to 0.1997 (+0.0223), LPIPS drops from 0.6044 to 0.5848 (−0.0196), and RotErr falls from 20.3759 to 15.0887 (26.0% reduction). TransErr improves from 1.9952 to 1.8388 and CamMC from 2.1507 to 1.9482.
-
Stage-1 gains on SANA-WM-Bench (Simple split). PSNR rises from 9.18 to 9.91 (+0.73 dB), SSIM from 0.1729 to 0.1999, LPIPS from 0.6327 to 0.6049, and RotErr from 19.7040 to 13.5864 (31.0% reduction).
-
Headline numbers. The abstract reports a 0.74 dB PSNR gain and a 28.5% reduction in rotation error on the 60-second benchmark, with the same six-metric improvement pattern.
-
Gains persist after downstream refinement. Under full-sequence (bidirectional) refinement, PSNR gains are 0.52 dB / 0.46 dB and RotErr reductions are 22.0% / 20.1% on the Simple/Hard splits. Under causal AR refinement, all three revisit-consistency metrics improve on both splits, and all three camera-control metrics improve on Hard, though TransErr and CamMC rise slightly on Simple.
-
Generalization to MBench-A (547 samples). Stage-1 PSNR rises from 10.17 to 11.01 (+0.84 dB), SSIM from 0.2545 to 0.2848 (+0.0303), and LPIPS falls from 0.6251 to 0.5823 (−0.0428). With causal AR refinement, PSNR improves by 0.37 dB (12.36 → 12.73) and with bidirectional refinement by 0.58 dB (12.94 → 13.52). All three revisit-consistency metrics improve across all four subsets and all evaluated inference modes.
-
Memory and throughput. At a 60-second context, HLA-WM requires 2.15 GiB of historical-state memory, 12× less than full KV caching, while incurring at most a 1.6% reduction in inference throughput. Reported FPS drops from 22.403 to 22.053 at Stage 1, from 8.280 to 8.211 under AR refinement, and stays at 7.700 vs 7.701 under bidirectional refinement.
-
Component ablation (Hard split, Stage 1, K=1). The full HLA-WM reaches PSNR 10.11 / SSIM 0.1997 / LPIPS 0.5848 / RotErr 15.0887. Removing the sink chunk drops PSNR to 9.71 and worsens RotErr to 15.9351. Retaining all unselected transition terms ($A$) drops PSNR to 9.61, retaining all historical write terms ($B$) drops it to 9.64, and removing turn-aware recent context drops it to 9.96 — each also degrading camera control.
-
Both memory pathways matter. Under Top-1 retrieval on the Simple split at Stage 1, applying recomposition only to GDN reaches PSNR 9.45, applying selected history only to the softmax KV cache reaches 9.50, and applying it to both (full HLA-WM) reaches 9.91.
Methodology in Plain English
The authors first break a GDN rollout into temporal chunks and show that each chunk's effect on memory can be written as a single affine transform: a matrix $A_i$ describing how the chunk reshapes existing memory, and a matrix $B_i$ describing what the chunk writes. These pairs are small, are computed once per chunk, and combine associatively — so any subset of them can be stitched back together in chronological order without re-running the original video frames through the model.
At generation time, for each query chunk the system asks which historical chunks actually observed the same place. It answers this using only the prescribed camera trajectory: sample a 7×7 grid of image-plane rays per camera frame, place proxy 3D points at three depths around the scene's median depth (0.5×, 1.0×, and 1.5× $d_\text{med}$), project those points into the other chunk's views, and measure the fraction that land with positive depth. Averaging the two directions gives a symmetric frustum-overlap score. The top-$K$ chunks by that score are retrieved, plus the first chunk as a permanent "sink" for global context and the most recent chunk for local continuity. When the camera's translation direction changes by more than 25° inside a chunk, the recent window temporarily expands from one to three chunks.
The retrieved chunks are reordered chronologically, their cached summaries are composed into a fresh recurrent state, and that state replaces the single accumulated GDN memory for the current step. Only the main GDN memory is replaced — the camera-control GDN state and local convolutional states keep their original streaming updates. The softmax-attention pathway receives the same selected chunk set. Everything runs on top of chunk-wise Triton GDN kernels, with the $t=0$ forward pass caching each completed chunk's summary for later retrieval. No parameters are trained or fine-tuned.
Why This Matters
This work reframes long-horizon memory in world models as a retrieval problem rather than a compression problem. It shows that the mismatch between temporal memory updating and spatial relevance is the root cause of forgetting, and that a pretrained model can be repaired at inference time with camera geometry alone — an unusually cheap intervention for a problem that normally invites retraining or larger caches. It also provides a quantitative diagnostic (the cumulative retention measure $W_{i\rightarrow j}$) that other recurrent architectures can reuse.
Real-world applications:
- Interactive game and simulation environments, where players or agents roam and return to previously visited areas that must stay spatially consistent for 60 seconds or more of continuous generation.
- Robotics and embodied AI, where world models let agents anticipate the consequences of actions and need stable memory of earlier observations.
- Autonomous driving and navigation simulation, where camera trajectories revisit intersections, landmarks, and road geometry during long rollouts.
- Virtual production and cinematic previsualization, where a camera path loops back on a set or location and footage must remain visually coherent.
Industry relevance: The 12× memory reduction and at-most-1.6% throughput cost map directly onto the economics of serving long video generation — long-context inference is limited by KV-cache memory, and HLA-WM offers a middle regime between constant-memory recurrence and full-history caching using metadata (camera poses, intrinsics, a depth scalar) that already exists in most camera-controlled pipelines.
Future Directions
- Integrate selective retrieval into training. The authors note that HLA-WM's training-free nature limits the achievable gains and suggest training with selective state retrieval to better preserve and exploit long-range memory.
- Richer retrieval signals. Current retrieval uses only camera geometry and field-of-view overlap, without explicitly modeling occlusion, visibility, or semantic relevance; learned or richer signals could improve robustness in complex scenes.
- Better handling of mode-dependent refinement trade-offs. Under causal AR refinement on Simple trajectories, RotErr improves but TransErr and CamMC slightly increase, and on Hard trajectories AR refinement produces a smaller PSNR gain (14.12 → 14.47) than bidirectional refinement — these interactions are not fully resolved.
- Retrieval budget and historical-selection strategies. The main results use Top-1 retrieval; the paper states that ablations on retrieval budget and historical selection strategies are provided in the Appendix, leaving room to characterize the accuracy/efficiency frontier more fully.
Target Audience
Researchers and engineers working on video generation, video world models, and efficient long-context inference — particularly those using linear-attention or recurrent architectures such as Gated DeltaNet, and those building camera-controlled or action-conditioned simulation systems. It is also relevant to practitioners interested in training-free inference-time methods and in memory-efficiency engineering for long video rollouts. Readers without a background in attention mechanisms and recurrent state formulations will find the method section demanding.
Authors’ abstract
Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-size states with substantially lower memory cost. However, we identify severe long-range forgetting in Gated DeltaNet (GDN), where information from distant but relevant scenes is progressively attenuated by subsequent state updates. To address this limitation, we propose HLA-WM, a training-free hybrid linear-attention framework that combines coarse-grained geometry-guided retrieval with fine-grained recurrent linear-state computation. HLA-WM exploits the affine structure of GDN to cache compact chunk-wise transition summaries, retrieve scene-relevant historical chunks using camera geometry, and recompose them into query-specific recurrent states. On the $60$-second SANA-WM-Bench, HLA-WM improves all six aggregate revisit-consistency and camera-control metrics of the base autoregressive generator without additional training, including a $0.74$ dB PSNR gain and a $28.5\%$ reduction in rotation error. The improvements persist after downstream refinement and generalize to MBench-A, where HLA-WM consistently improves all three revisit-consistency metrics across all four subsets and all evaluated inference modes over $547$ samples. At a $60$-second context, HLA-WM reduces historical-state memory by $12\times$ relative to full KV caching while incurring at most a $1.6\%$ reduction in inference throughput. These results demonstrate that selectively addressable recurrent memory can improve long-range scene recall while preserving the efficiency advantages of GDN. Project page: https://caesarhhh.github.io/hla-wm/