Skip to content
AI.info

Research

LOCI: Spatial Linear Memory for Streaming World Models

Overview Research area: Computer vision and generative video — specifically camera-controllable, streaming "world models" that must remember and reproduce previously observed places when the camera re

LOCI: Spatial Linear Memory for Streaming World Models
arXiv
2609.40222
Published
2026-09-30
Authors
Ji Xia, Tingting Liao, Xuezhi Liang, Hao Li, Guangyi Liu

AI summary

Overview

Research area: Computer vision and generative video — specifically camera-controllable, streaming "world models" that must remember and reproduce previously observed places when the camera returns to them.

Technical level: Advanced. The paper builds on transformer key–value caching, Kimi Delta Attention (a linear-attention/delta-rule recurrent mechanism), and projective camera positional encoding (PRoPE); some familiarity with attention and video diffusion is assumed.

Scope: The paper introduces LOCI, a hybrid memory architecture that combines a fixed-size, camera-conditioned recurrent state with an explicitly stored key–value cache of retained observations, and evaluates it on the public MIND memory benchmark, WBench, and held-out recorded Unreal Engine trajectories against several external world models and an identically trained full-softmax baseline.

What This Paper Is About

Video world models can generate long, interactive camera paths, but they often fail when the camera turns back toward something it saw earlier: the scene structure and appearance get replaced or distorted. Fixing this requires two things at once — storing past observations, and retrieving the right one for the current viewpoint. Existing approaches force a trade-off, since key–value caches preserve detail but grow without bound, while recurrent memory stays compact but compresses history into a fixed-size state where individual past observations are no longer directly accessible. LOCI is designed to keep both representations and let them interact.

Key Contributions

  1. A hybrid spatial memory that couples the two memory types. Projectively conditioned recurrent integration is combined with direct historical key–value access, so recurrent context informs the queries used by later historical-attention blocks while observation-level detail is still preserved in the cache. Channel-wise retention is applied once per chunk, so explicit forgetting follows elapsed video time rather than token order; in an extended comparison on MIND, this lowers local error at revisits more than 20 s apart relative to per-token retention (Appendix F).

  2. A bounded streaming mode. The recurrent state, carried over the full history, is paired with a fixed-capacity bank of retained observations for streaming. Under an identical bounded KV budget, the recurrent path improves fidelity over a same-recipe full-softmax model on the MIND memory test (all 50 segments; full-segment PSNR +0.89 dB, 95% interval [+0.64, +1.15], with lower LPIPS and MSE), while generation runs for 300 seconds at constant memory.

  3. Zero-shot results on the public MIND benchmark. With full-history access and no MIND training data, LOCI attains lower MSE and higher PSNR and SSIM than the values reported for GIM-World, which is trained on MIND. At short-horizon revisits, the hybrid is also more consistent locally than full softmax.

Main Findings

  • Revisit fidelity versus an identical full-softmax model. With the same 5B backbone (Wan2.2-TI2V-5B), data, recipe and 5,000 updates, LOCI improves reference PSNR at revisits by 0.62 dB on held-out Unreal Engine trajectories, and by 0.99 dB on set A when both models access the same bounded set of retained observations. On the MIND memory test under a bounded budget, PSNR rises by 0.89 dB [+0.64, +1.15] and the hybrid is better in 44 of 50 segments — so the gain does not come from storing more history.

  • Full MIND memory test results. Evaluated on the entire prediction segment with full history, LOCI scores MSE 0.0455, PSNR 14.36, SSIM 0.464, LPIPS 0.643 across 50 segments, versus reported values for GIM-World (1.3B, trained on MIND) of MSE 0.0614, PSNR 13.40, SSIM 0.414, LPIPS 0.630. Among the streaming world models the authors ran with the memory segment as context, LOCI is best on all four metrics among those scored on all segments — 1.82 dB PSNR above the strongest, HY-WorldPlay (12.54 dB), which is an 8B model. In bounded sparse mode, LOCI scores MSE 0.0622, PSNR 13.02, SSIM 0.409, LPIPS 0.679.

  • Comparison against external world models. On the 36 segments CaR completes, LOCI has MSE 0.0460, PSNR 14.39, SSIM 0.478 and LPIPS 0.630, versus CaR's MSE 0.0538, PSNR 13.85, SSIM 0.458 and LPIPS 0.622 — so CaR has lower LPIPS on that subset. On held-out recorded trajectories, LOCI has significantly lower LPIPS than six of the seven external models the authors run and significantly higher PSNR than five of them.

  • Full softmax runs out of GPU memory where LOCI does not. With full history, LOCI completes the seven segments (100–138 s of prediction) on which full softmax exceeds GPU memory. On the 32 segments both models complete, LOCI's advantage is PSNR +0.35 dB [+0.08, +0.62] and LPIPS −0.029 [−0.040, −0.018].

  • Long-range recall under a bounded budget. Both models were given the first 68–245 s of four held-out recorded trajectories as ground-truth history under the same bounded access and asked to generate the remaining 45–60 s, which revisit places first seen at least 60 s earlier. LOCI reproduces these returns more faithfully on all four trajectories (revisit PSNR +0.68 to +1.93 dB, lower LPIPS on each).

  • Local consistency at short-horizon revisits. Scoring the worst 5% of patches by DINOv2 feature distance between generated and ground-truth revisit frames (revisits 8–20 s after the first visit) on the MIND memory test, local error is significantly lower for LOCI than full softmax (−0.073) and lower in 12 of 13 segments with such revisits, whereas whole-frame LPIPS does not separate the models. Resetting the recurrent state at every chunk during inference raises this local error (+0.021 [+0.013, +0.028]) and blurs revisit frames (Laplacian-variance sharpness 0.12 with reset vs. 0.18 without), each in 15 of 16 segments — evidence that the recurrent state contributes to short-horizon fidelity.

  • The recurrent state is camera-decodable. Camera yaw relative to the first frame remains linearly decodable from the accumulated KDA state: held-out R² = 0.946 [0.929, 0.960] with a median error of 3.6°, and coordinate-invariant summaries of the state retain R² = 0.941. An untrained branch with the same conditioning decodes about as well, so decodability reflects the encoding rather than learned behaviour. Among models trained on the uniform mixture, PRoPE raises this R² by +0.646 [+0.544, +0.749] over a recurrent memory without camera encoding.

  • Historical attention concentrates on co-visible content. Co-visible tokens make up only about 0.1% of the history, yet LOCI's 15 softmax layers place 11.1% of their history attention on them, versus 10.45% for the same 15 layers of full softmax — an enrichment over uniform attention of 115× versus 108× (+6.64), higher in all 8 clips. Under bounded access the pattern holds (88× vs. 81×, +6.82, in 8 of 8). Zeroing only the cross-chunk readout at inference lowers downstream co-visible enrichment (+0.56 for on minus off; +0.86 right after a hybrid block), and a norm-matched random readout lowers it further. On its own rollouts, LOCI reconstructs the co-visible region more closely (masked LPIPS 0.597 vs. 0.654, lower at 131 of 154 instants).

  • Cost and latency. With full history, replacing half of the historical-KV layers with recurrent state extends reachable length from 157 s to 225 s and lowers peak memory at equal length by about 30%; LOCI is also faster (7.2 vs. 8.1 s per video second over the first 40 s). Under bounded sparse access it streams the full 300 s at a constant 23.6 GiB and constant time per chunk — 15% less memory than full softmax under the same access (27.6 GiB), at the same speed of 5.4 s per video second on one H200 with an output-preserving optimized implementation.

  • Design details that mattered. Converting the trained full-softmax model into the same hybrid layout with the ARL² recipe does not reach the directly trained hybrid, and neither converted model exceeds its own teacher. The recurrent branch adds 90.7M parameters (1.7% of 5.38B). Bounded mode keeps at most 34 latent frames: the conditioning frame as a sink, a bank of at most 20 older views, at most 8 recent completed frames, and the current chunk of 5.

  • Limitations reported by the authors. Held-out trajectories are rendered in Unreal Engine; real-world captures are not evaluated. Absolute fidelity is low for all models (benchmark-average PSNR below 15 dB), so gains should be read as relative. The study covers one 5B backbone with a short fine-tuning budget of 5,000 updates. Both models keep an explicit-history camera-attention branch in all 30 blocks, so the recurrent path halves main-attention KV rather than all stored history. The recurrent state is built from noised chunk features in training but committed from generated chunks at inference, and speed parity in bounded sparse mode relies on the optimized implementation.

Methodology in Plain English

The authors start from a 30-block pretrained video transformer (Wan2.2-TI2V-5B) and make it generate video chunk by chunk, where each chunk is five latent frames, with camera pose and intrinsics supplied per frame. They then split the blocks in half. Fifteen blocks keep ordinary softmax attention with an explicit historical key–value cache — the "what exactly did we see" store. The other fifteen keep only a short local window of the current chunk plus a recurrent linear-attention memory built on Kimi Delta Attention, which is a fixed-size state that is written token by token but never grows with video length.

The key move is that this recurrent state is conditioned on camera geometry: using PRoPE, each token's queries, keys and values are transformed by that token's camera projection, so what the memory writes and reads depends on viewpoint rather than only on how recently the frame appeared. A second move is to apply channel-wise retention (forgetting) only once per chunk, at the first token, computed from the chunk's mean representation, while the token-level delta corrections that revise individual key–value associations still happen per token. This makes forgetting track elapsed video time instead of raster position within a frame.

Information flows through network depth: the recurrent readout is scaled by a learned per-token, per-head gate and added to the local attention output, and the resulting token features are what later historical-attention blocks use to form their queries. So compressed history steers which retained observations the explicit cache attends to, with no separate router or auxiliary objective.

For streaming, they add a bounded mode: a fixed bank holding the conditioning frame, up to 20 diverse older views chosen by a field-of-view coverage criterion, up to 8 recent frames, and the current chunk. Discarded observations are not archived, and the recurrent state is unaffected by bank selection, so neither memory store grows.

Training uses chunk-wise diffusion forcing over 81 latent frames (one conditioning frame plus sixteen future chunks), a single full-window forward scanning noised chunks from zero state, and 5,000 updates. Training data is roughly 98 hours of Unreal Engine and CARLA scenes the authors rendered plus real walking videos from Sekai; no MIND data is used. At sampling, each denoising iteration reads the committed state without modifying it, and a separate forward after the last denoising step commits the state for the next chunk. Evaluation follows MIND's official protocol on the entire prediction segment, defines a revisit instant as a pose within 0.15 m and 32° of an earlier one reached after at least 8 s and after looking away by 60°, and reports MSE, PSNR, SSIM and LPIPS with 95% bootstrap intervals over clips using 10,000 resamples.

Why This Matters

The paper attacks a specific and practically important failure mode of generative world models: they look plausible moment to moment but forget the world they generated. LOCI shows that a fixed-size recurrent memory can be made spatially meaningful by conditioning it on camera geometry, and that it can be combined with explicit caching so the two mechanisms reinforce each other rather than being alternatives. That is a design pattern — geometry-aware compression plus retained detail, interacting through the feature stream — that could transfer beyond this particular backbone or task.

For research, the paper contributes a controlled comparison: the hybrid is measured against a full-softmax model trained with the same backbone, data, recipe and number of updates, isolating architecture from scale or data advantages. It also contributes negative and boundary results, such as the conversion recipe failing to reach direct training, and quantified evidence (via linear decoding of yaw from the recurrent state) that the conditioning actually encodes viewpoint.

Real-world applications that this kind of capability would support:

  • Interactive game and simulation environments where a player or camera agent wanders and later returns to a location that must look the same.
  • Virtual and augmented reality exploration, where head movement frequently revisits earlier viewpoints and inconsistency is immediately visible.
  • Embodied AI and robotics, where an agent navigating a previously mapped space needs a consistent visual model of places it has left.
  • Virtual production and previsualization, where long camera takes in a virtual set must stay stable as the camera revisits areas.

For industry, the cost profile is as relevant as the quality numbers: bounded sparse mode streams 300 seconds at constant 23.6 GiB and 5.4 s per video second, 15% less memory than full softmax under the same access and at the same speed. Constant-memory streaming at fixed latency is what determines whether long interactive generation is deployable on a given GPU budget.

Future Directions

  • Evaluation on real-world captures. The authors note that real-world footage is not evaluated; extending the revisit protocol to captured video rather than Unreal Engine renders is the clearest next step.
  • Scaling the backbone and training budget. The study uses one 5B backbone and 5,000 updates; whether the reported gains persist or grow with larger models, longer training, and more diverse data is open.
  • Reducing the remaining explicit-history footprint. Because the camera-attention branch keeps explicit history in all 30 blocks, the recurrent path only halves main-attention KV; making the camera branch recurrence-compatible would test how far fixed-size memory can be pushed.
  • Training–inference mismatch of the recurrent state. The state is built from noised chunk features during training but committed from generated chunks at inference, and speed parity depends on an optimized implementation — closing that gap and making the optimized path standard are practical open items.
  • Retention schedule design. Per-chunk retention helps at long revisit intervals while per-token retention matched it at short-range revisits, leaving room for a schedule that adapts to both.

Target Audience

This paper is most useful to researchers and engineers working on video generation, world models, and long-horizon interactive content, especially those already familiar with attention caching and linear-attention or state-space sequence models. It will also interest practitioners building camera-controllable generative systems who need to understand the memory-versus-cost trade-off, and readers tracking how geometric conditioning (camera poses) can be injected into recurrent memory rather than only into softmax attention. Given the technical density of the preliminaries and appendices, some background in transformer attention and diffusion video models is assumed.

Authors’ abstract

When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.

Read the original paper