Research
LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation
Overview Research area: Computer vision and generative modeling — specifically autoregressive video diffusion transformers, long-video generation, and memory retrieval under bounded key-value (KV) cac

- arXiv
- 2608.28460
- Published
- 2026-08-28
- Authors
- Yixuan Ding, Jiahao Kong, Wei Huang, Ruijie Quan, Yi Yang
AI summary
Overview
Research area: Computer vision and generative modeling — specifically autoregressive video diffusion transformers, long-video generation, and memory retrieval under bounded key-value (KV) caches.
Technical level: Advanced. The paper assumes familiarity with diffusion transformers (DiTs), attention, KV caching, and autoregressive chunk-based generation.
Scope (one sentence): The paper introduces LayerRecall, a small trainable memory router that decides which historical key/value states to retrieve and which network layers should receive them, plus a training scheme (Cross-Horizon Prediction Matching) that supervises this router without long-horizon video labels.
What This Paper Is About
Autoregressive video diffusion models generate long videos chunk by chunk, but they can only "see" a bounded recent window of frames. When a person, object, attribute, or scene detail leaves that window and later reappears, recency-based caching has already thrown away the evidence needed to reproduce it, causing identity, attribute, count, and scene drift. The paper's goal is to let the model reuse distant history effectively — not merely store it — by deciding what to retrieve from past chunks and where in the network to inject what was retrieved.
Key Contributions
- Layer-wise temporal preference analysis. The authors profile how attention mass is distributed across current, recent, and distant context in each DiT layer, showing that video DiT layers have distinct and backbone-specific preferences rather than a uniform response, and that these patterns are relatively stable across 100 prompts.
- LayerRecall, a state-conditioned memory router. A small module that builds a retrieval query from the current generation state, scores resident historical chunk summaries by cosine similarity, and injects the selected full K/V payload only into a fixed set of memory-sensitive layers, leaving all other layers on the backbone's original sink-plus-sliding-window attention.
- Cross-Horizon Prediction Matching (CHPM). A training objective that uses a frozen long-context teacher as a behavioral reference and supervises the bounded-memory student in prediction space, requiring neither high-quality long-horizon target videos nor explicit memory-allocation labels, and without fine-tuning the generative backbone.
- Comprehensive evaluation. Experiments covering multi-shot consistency (100 evaluation prompts), temporal stability, cross-backbone portability, and inference efficiency, with ablations isolating layer selection and the CHPM objective.
Main Findings
- Best overall memory-oriented scores. LayerRecall ranks first on both MemoBench and MovieBench overall and remains among the top methods on most diagnostic subdimensions, across videos generated from 100 evaluation prompts.
- No loss on general quality. LayerRecall matches its backbone LongLive-2.0 on VBench-Long, reporting 0.978 on the VBench-Long average, indicating memory gains do not sacrifice consistency or motion quality.
- Layer selection matters. In a matched ablation with a fixed checkpoint, retrieval mechanism, memory budget, and the same 100 prompts, the profiled layer policy
[4, 9, 10, 12, 13, 15, 16, 17, 18, 26]raises the MemoBench overall score from 0.538 (ten randomly sampled layers[0, 3, 4, 7, 8, 17, 20, 23, 24, 28]) to 0.570, improving object reappearance, state, layout, and camera consistency. Identity is the only reported dimension where the random policy scores higher (0.507 versus 0.495). - All-layer routing hurts temporal stability. Compared with injecting retrieved memory into every layer, LayerRecall improves VBench motion smoothness from 0.9910 to 0.9916, Helios smoothness from 0.9917 to 0.9923, adjacent-frame CLIP consistency from 0.9397 to 0.9574, and adjacent-frame DINO consistency from 0.8795 to 0.9212. The high-frequency frame-change energy ratio drops from 0.60 to 0.38.
- CHPM supervision is effective. Using the same ten memory-sensitive layers, a randomly initialized router scores 0.519 MemoBench overall versus 0.548 for the CHPM-trained router, with gains in every reported subdimension (ID 0.484 to 0.530, state 0.487 to 0.525, camera 0.508 to 0.573).
- Cross-backbone portability without training. Reusing the trained LayerRecall parameters on LongLive and Self-Forcing (1.3B Wan2.1 family rather than the 5B Wan2.2 of LongLive-2.0) improves MemoBench overall on both backbones in most dimensions. Replacing only the insertion-layer policy with each target backbone's own profiled policy gives the best overall score on both models, though individual dimensions trade off. The authors restrict this claim to the tested models.
- Negligible inference cost. End-to-end generation time on H100 hardware changes from 305.9 to 309.4 seconds per video, and FPS from 5.22 to 5.16.
- Memory-guided self-correction. In qualitative examples, an initially mismatched local attribute (for instance a patterned inner garment) returns to its historical appearance later within the same shot, while identity, action, and scene structure continue without a global reset. The authors note LayerRecall does not explicitly detect errors; the correction emerges from state-conditioned retrieval and restricted injection.
- Very small trainable footprint. The LayerRecall overlay has 11 trainable tensors and 1,648,416 parameters, approximately 0.033% of the 5B backbone, and only these parameters are optimized.
Methodology in Plain English
The authors start by measuring where in the network a pretrained autoregressive video DiT actually pays attention to old versus recent context. They find that different layers care about history to very different degrees, so they fix a backbone-specific list of "memory-sensitive" layers for injection and leave every other layer untouched.
Retrieval works like a lightweight search index. For every past chunk, each layer stores a small pooled summary of its normalized pre-RoPE keys alongside a pointer to the full cached keys and values. At the current chunk, a query is formed from the current hidden state — a shared learnable global query plus a current-conditioned, gated residual whose final projection is zero-initialized, so the router begins as a plain global query and learns to specialize. Each layer ranks its resident historical summaries by cosine similarity and picks the top entries without replacement. The forward pass uses the hard-selected chunk's real K/V, while a temperature-scaled soft mixture supplies gradients through a straight-through estimator, so routing remains differentiable despite discrete selection.
Training uses two frozen copies of the same backbone: a teacher that attends to the full expanded context (384 latent frames, 48 chunks) and a student restricted to a 32-frame attention-visible budget with an 80-frame physical cache (10 chunks). Both see the same noisy latent, prompt, and diffusion timestep at supervised anchors, and the router is trained to make the student's denoised prediction match the teacher's detached prediction, plus a regularization term on the router's parameter magnitudes. The student rolls out its own trajectory — non-anchor chunks are denoised without gradients and written back as detached context — so the router experiences the self-generated history it will face at inference, and no gradients span the whole video.
Training used 1,600 multi-shot prompts of 48 blocks each on 16 NVIDIA H100 GPUs with sequence parallelism 2 and data parallelism 8 for one epoch. The loss weights (λ_pred and λ_reg) are stated to be given in the supplementary material, not in the main text, and no human study or user preference evaluation is reported.
Why This Matters
Research impact. The paper reframes long-video memory as a two-part allocation problem — across historical content and across network depth — rather than a single retrieval problem. It offers a training recipe that avoids scarce long-horizon supervision and avoids backbone fine-tuning, and it provides layer-preference analysis for two additional backbones, which is useful to anyone studying how attention is distributed in video DiTs.
Real-world applications.
- Long-form narrative or episodic content generation where characters, costumes, props, and locations must stay consistent across many shots.
- Streaming or interactive video synthesis, where bounded caches and negligible added latency (roughly 3.5 seconds per video in the reported H100 setting) are practical requirements.
- Advertising and product video, where a specific object or brand element must survive shot changes and reappear faithfully.
- Virtual production and previsualization, where continuity of identity and set detail across an extended sequence reduces manual correction work.
Industry relevance. The method is a small, portable add-on: 1,648,416 parameters, inference overhead described as negligible, and demonstrated reuse on two backbones without retraining. Those properties suit deployment on top of existing pretrained video generators rather than requiring new foundation models.
Future Directions
- Adaptive layer activation. The authors explicitly list exploring adaptive layer activation as future work; the current insertion policy is fixed per backbone rather than varying with content.
- Compressed memory. Also listed as future work, pointing toward smaller or more structured historical stores than keeping full K/V payloads around in an 80-frame physical cache.
- Broader backbone validation. Portability was tested on LongLive and Self-Forcing (1.3B Wan2.1 family) with zero training; the authors limit their claim to these evaluated models, leaving open whether the routing policy generalizes further.
- Label-free alternatives to the teacher reference. CHPM still requires a privileged long-context teacher at training time; whether equivalent supervision can be obtained without running an expanded-context trajectory is not addressed.
Target Audience
Researchers and engineers working on video diffusion, autoregressive or streaming generation, and efficient attention/KV-cache management will get the most from this paper. It is also relevant to practitioners who need long-horizon consistency on top of an existing video backbone, and to readers interested in layer-wise analysis of transformer attention. Because it assumes comfort with DiTs, attention, and KV caching, it is not an introductory read.
Authors’ abstract
Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent context. While recency-based caching preserves local continuity, it evicts historical cues needed when subjects, objects, scenes, or attributes reappear. Existing memory mechanisms expose models to nonlocal history, but access alone does not ensure effective use. Our analysis reveals that video DiT layers exhibit distinct preferences for current, recent, and distant context, suggesting that long-range memory requires deciding both what to retrieve and where to use it. We introduce LayerRecall, a current-conditioned, layer-selective memory router that retrieves relevant historical K/V states and injects them only into backbone-specific memory-sensitive layers while preserving local attention elsewhere. To reduce reliance on scarce high-quality long-horizon videos and explicit memory-allocation labels, we further propose Cross-Horizon Prediction Matching (CHPM), which uses a privileged long-context reference to supervise the bounded-memory router in prediction space. Across 100 multi-shot evaluation prompts, LayerRecall achieves the best overall results on MemoBench and MovieBench while matching its backbone on VBench-Long, demonstrating stronger long-range recovery without sacrificing local continuity. Qualitative analyses further reveal memory-guided self-correction, whereby initially mismatched local attributes return to their historical appearance without resetting ongoing motion or scene structure. Additional analyses show cross-backbone portability and negligible inference overhead.