Skip to content
AI.info

Research

Memorizon: Training World Models Beyond Their Context Window

Memorizon: Training World Models Beyond Their Context Window Overview Research area: Computer vision, specifically generative video world models — the class of models that simulate a 3D environment by

Memorizon: Training World Models Beyond Their Context Window
arXiv
2610.00544
Published
2026-09-30
Authors
Tingting Liao, Xuezhi Liang, Hao Li, Guangyi Liu

AI summary

Memorizon: Training World Models Beyond Their Context Window

Overview

Research area: Computer vision, specifically generative video world models — the class of models that simulate a 3D environment by generating video along a camera trajectory.

Technical level: Advanced. The paper assumes familiarity with diffusion transformers, causal attention, rotary positional embeddings (RoPE), attention sinks, and knowledge distillation.

Scope (one sentence): The paper proposes a training recipe that lets a streaming video world model be supervised on revisits separated by spans far longer than its attention context (up to 400 s tested), by decoupling the span a sample covers from the sequence the transformer actually attends over.

What This Paper Is About

Streaming world models should render a place the same way each time the camera returns to it, but training samples are limited to short clips (roughly tens of seconds) because attention cost is quadratic in sequence length, while the models are deployed for minutes. If the camera takes longer to leave a place and come back than a training clip lasts, no sample contains both visits, so no loss term ever relates them — on this corpus the shortest genuine return spans 13.25 s, and stricter definitions push it past 28 s. The paper's goal is to make long-span supervision affordable: samples can cover arbitrarily long spans, but the model only attends to a small bounded set of positions.

Key Contributions

  1. Memorizon, a training recipe with bounded cost. A training sample covers a span of any length, but only its last k chunks are scored. Each scored chunk retrieves its own top-K latents by camera co-visibility, and the union of those requests forms a shared memory bank. The bank is bounded by k·K, so the sequence length is independent of the span; at the shortest span the recipe reduces exactly to ordinary training.

  2. A demonstration that per-chunk retrieval and span length do different jobs. Per-chunk retrieval under one shared positional index lets a model read frames far older than any it was trained on; a span long enough to reach the first visit turns returns into supervised signal, adding gains on the returns that test memory hardest (mid-path returns and returns more than 10 s apart), with no further gain once the span covers the first visits.

  3. Evidence that the bank saturates and step cost barely grows. For K=6 the bank holds 28 entries at a 100 s span and 31.5 at 400 s, while the candidate pool grows tenfold from 50 to 400 s. Going from a 100 s to a 400 s span adds 12% to step time.

  4. A causal check that the model reads what it retrieves. Replacing the bank's contents at inference (random frames, an empty bank, or frames from another episode) degrades memory performance, and filling the bank from another episode lowers revisit correlation by 83%.

Main Findings

  • Retrieval helps on every split, even from a 10 s model. Adding per-chunk top-K retrieval to a sliding-window baseline raises Revisit on all splits and multiplies Gain several times over, already for a model trained on 10 s clips. The rise is smallest on seen scenes and largest on web photographs, where only retrieved frames tell the model what the place looked like.

  • More attention is no substitute for retrieval. Widening the window to the whole clip without retrieval helps on unseen scenes, is level on seen scenes, and lowers Revisit on web photographs — on 16 of the 20 photographs and in every one of the five rollout seeds.

  • Who retrieves matters. Sharing one retrieval across the whole scored block (the Context-as-Memory arrangement) lowers Revisit and Gain on every split at the same span, slightly on rendered scenes and by about a third on web photographs.

  • Longer spans help up to about 200 s, then stop. Revisit and Gain rise on every split from 10 s to 100 s to 200 s, and fall back at 400 s. Spans of 100, 200 and 400 s reach the first visit of 64%, 84% and 93% of returning latents respectively; once the span covers the first visits, more length only adds candidates no chunk chooses.

  • A bank index that is shared beats temporal ordering for features and quality. Numbering the bank entries in temporal order raises Revisit in pixels on seen and unseen scenes (0.600 vs 0.559 on unseen scenes at 100 s) and is level on web photographs, but DINO Revisit is higher only on seen scenes and imaging quality is lower on every split.

  • Image quality is a real cost. Imaging quality falls as the span grows to 200 s, most on unseen scenes and web photographs, and recovers at 400 s where Revisit falls back. The cause is not a weaker generator — given a ground-truth history, every model from 10 s to 200 s reaches the quality of the rendered videos; the loss comes from conditioning on the model's own output, most of it through the bank.

  • Memorizon and CaR are the only systems with clear memory among the compared open world models. Memorizon is highest in five of the eight memory columns of Table 1 and CaR in the other three: 0.365 Revisit at the starting pose and 0.447 mid-path for Memorizon, versus 0.350 and 0.450 for CaR and 0.254 and 0.283 for Matrix-Game 3.0, the strongest other baseline. Gain (which subtracts what any two frames of one video share at the same time gap) is at most 0.089 for the other five baselines, and −0.006 for LingBot-World 2.0 at the starting pose, against 0.230 and 0.329 for Memorizon and 0.220 and 0.312 for CaR.

  • The margin is not just camera control. The other baselines follow the prescribed path with correlations of 0.70 to 0.88 against 0.95 to 0.97 for Memorizon; and of the 173 returns in this window, 125 (72%) are fold-backs, where a consistent scale error largely cancels.

  • VBench is not led. HY-WorldPlay 1.5 scores higher on subject and background consistency (0.839 and 0.908 against 0.778 and 0.857), and four of the six baselines score higher on imaging quality, up to 0.761 against 0.705. The paper notes both consistency scores compare neighbouring frames, so a smoothly drifting rollout scores well on them and poorly on Revisit.

  • Retrieval criterion graded against ground truth. Overlap alone reaches 47.4% of achievable gain; the mixed score reaches 57.2% at λ=0.2, level with the best value of 57.3% at λ=0.1; pose distance alone reaches 55.9%, mostly through its orientation term, since ranking by camera centres alone gives only 19.8%.

  • Long-horizon returns. Over 400 s rollouts on web photographs, the 10 s top-K model decays from 0.390 ± 0.023 (returns under half a minute apart) to 0.303 ± 0.025 at two to four minutes and 0.145 ± 0.044 beyond four minutes. The 200 s model is ahead in every bin, including 0.148 ± 0.040 at one to two minutes and 0.103 ± 0.061 beyond four. The 400 s model is level with the 10 s model under a minute and at two to four minutes.

Methodology in Plain English

The key idea is to stop tying the span a training sample covers to the sequence the transformer attends over. A training sample is built from a first frame, a history of chunks, a recent block of r chunks, and a scored block of the last k chunks — but the history is never tokenized. Instead, each scored chunk asks for its own top-K frames from the history, ranked by camera co-visibility, and the union of those requests becomes a shared bank. Because the bank is capped at k·K entries, the input sequence has a fixed size no matter how long the span is. Attention is bidirectional inside a chunk and causal across chunks with a sliding window of r+1 chunks; the first frame acts as an attention sink; and a per-chunk mask lets each chunk read only its own retrieved entries. Every scored chunk therefore attends to at most 1 + K + (r+1)c = 15 latents (with K=6, r=1, c=4), whatever the span.

Retrieval is scored by frustum overlap minus a weighted relative-pose distance, so a 15° turn weighs about as much as moving one unit (4 m). A shared RoPE index is given to all bank entries, so positions never depend on history length and the model is not told how old a retrieved frame is. To reduce exposure bias — training conditions on clean context, inference on the model's own output — the trained model is distilled into a four-step generator using Self Forcing and distribution matching distillation.

Training starts from Wan2.2-TI2V-5B with a PRoPE camera branch, as a chunked causal diffusion transformer using diffusion forcing: chunks of 4 latents, k=10 scored chunks, r=1 recent chunk, spans drawn from U[10, mmax], K=6, 6,000 steps, batch size 32 across 32 H200 GPUs. Evaluation averages five rollouts per model with 20 denoising steps and classifier-free guidance at scale 4, over 20 60 s clips from ten training scenes, 16 clips from four held-out scenes, and 20 web photographs. A return is defined as two latents whose cameras lie within 0.5 units and 15° of each other, at least 8 s apart with the camera away in between.

Why This Matters

Impact on research. The paper reframes long-horizon memory in world models as a data and supervision problem rather than purely an architectural one, and shows that the span a sample covers and the sequence a transformer attends to can be decoupled. It also gives a clean ablation structure — retrieval, bank formation, and span length each do a distinct job — and reports honest costs (image quality, VBench columns not led) that later work can target.

Real-world applications:

  • Interactive 3D scene generation and exploration, where a user walks away from a location and returns expecting it to look the same.
  • Game and simulation content production, where consistent environments must be generated over minutes of play rather than seconds of clip.
  • Robotics and embodied simulation, where an agent revisits landmarks and needs a stable internal model of the environment.
  • Virtual tours, film previsualization, and digital twins of real spaces, where footage or renders must remain coherent across long traversals.

Industry relevance. The recipe is a change in the training sampler and memory bank rather than a new architecture, and it was applied on top of an existing open model (Wan2.2-TI2V-5B), which lowers the barrier to adoption. The cost profile — one-time cost to add the bank, then about 12% extra step time going from a 100 s to a 400 s span — is the kind of scaling that matters for teams already paying for long-context training.

Future Directions

  • Moving content. The paper states it mainly focuses on static scenes where nothing moves but the camera, so a return is always to a place that should look the same. Scenes with moving objects or places that change between visits are explicitly left out.
  • Reducing the image-quality cost of long spans. The paper traces the quality loss to conditioning on generated frames, mostly through the bank, and notes the 400 s run reverses both trends — suggesting a model that relies less on its bank. A training objective that narrows this gap is an open target.
  • Better retrieval signals. The authors state that their criterion accounts for neither occlusion nor motion along the viewing axis.
  • Span as a coverage setting rather than a quantity to maximize. Since gains plateau once the span covers the first visits of returns, and returns more than four minutes apart are few (30 per seed), the natural next step is choosing span per corpus by measured return coverage rather than by maximizing length.

Target Audience

Researchers and engineers working on generative video world models, video diffusion transformers, and long-horizon memory in sequence models; practitioners interested in training-time retrieval rather than architectural memory; and readers tracking the gap between the clip lengths models are trained on and the minute-scale rollouts they are deployed for. Readers without a background in diffusion transformers, causal attention and RoPE will find the method sections demanding, though the ablation and comparison results are readable on their own.

Authors’ abstract

Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervision expensive. Memorizon breaks this coupling: long spans are needed for supervision, but not for attention, since the two visits can share a forward pass without including every intervening frame. A training sample covers a span of any length but is scored only on its last $k$ chunks. Instead of tokenizing the history before them, each scored chunk retrieves its own top-$K$ latents by camera co-visibility, and the union of these requests forms a shared bank. The bank is bounded by $kK$, so the sequence stays bounded however long the span; at the shortest span the recipe is exactly conventional training. Adding the bank raises the cost of a step once; beyond that, a longer span costs little, and going from 100 to 400 s adds 12% to the step time. Against a sliding-window baseline, retrieval raises revisit consistency on every split, and a span long enough to reach the first visit of each return adds a further 24% to 30%, at some cost in image quality; beyond that span, more length no longer helps. Filling the bank from another episode lowers revisit correlation by 83%, so the model uses what it retrieves. Project page: https://tingtingliao.github.io/memorizon

Read the original paper