Skip to content
AI.info

Research

Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

Overview Research area: computer vision, specifically long-horizon video generation and video world models, combined with multimodal large language models (MLLMs) used as memory controllers. Technical

Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory
arXiv
2610.02521
Published
2026-10-01
Authors
Ying Yang, Guiyu Zhang, Lianghua Huang, Chang Nie, Chenyang Si, Haofan Wang, Shaoshuai Shi, Li Jiang

AI summary

Overview

Research area: computer vision, specifically long-horizon video generation and video world models, combined with multimodal large language models (MLLMs) used as memory controllers. Technical level: Advanced. This paper proposes Spatial Memory Intelligence (SMI), a framework in which a fine-tuned MLLM manages the long-term spatial memory of an action-conditioned video world model through four "atomic operations" — spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering.

What This Paper Is About

Video world models generate future frames conditioned on past observations and user actions, so their memory of previously seen regions keeps growing in size and disorder. The paper's core problem is that existing memory strategies — compression, sparse attention, and geometry- or semantic-based retrieval — either discard fine-grained spatial evidence or mis-select memories in cluttered, occluded, or revisited scenes. SMI's goal is to delegate "what to keep, what to recall, and what to trust" to an understanding model that reasons about both the geometry and the semantics of the observation history.

Key Contributions

  1. The first MLLM-driven memory manager for world models. SMI is presented as the first unified framework that systematically employs an understanding model for spatial-memory management in long-horizon video world models, defined through four coordinated atomic operations (spatial clustering, within-cluster sparsification, action-aware retrieval, reliability-aware filtering).
  2. A structured memory representation. Spatial clustering groups temporally separated but spatially proximate chunks and assigns each cluster a prototype selected by largest average camera-projection coverage; new chunks are matched to a candidate prototype or start a new cluster.
  3. Operation-oriented supervision and training. The authors collected roughly one thousand inference trajectories, filtered low-quality samples, manually annotated more than one hundred examples, used those as demonstrations for multiple state-of-the-art teacher MLLMs to annotate the remainder, then had human annotators review, discard, and correct labels. The MLLM is then fine-tuned with standard cross-entropy (an SFT objective).
  4. Broad empirical validation. Experiments across two world-model backbones (HY1.5-8B and Wan2.2-5B from WorldPlay-1.5), five memory-management baselines, VBench, a three-agent GPT-5.6-sol blind evaluation, reconstruction consistency, and a user study.

Main Findings

  • Memory sparsification: SMI reaches 83.6842% sparsity on HY1.5 and 82.2073% on Wan2.2, versus 94.3820% for FramePack, 79.2135% for Deep Forcing, 49.4253% for MemFlow, and 0.0000% for the Base, MoC, and VMem.
  • HY1.5 quantitative gains: Aesthetic Quality rises from the Base's 0.445389 to 0.469357, GPT-evaluated Camera-Constrained Scene Consistency from 51.6053 to 67.0200, Overall Visual Quality from 53.1538 to 65.3677, PSNR from 12.478528 to 13.309393, and LPIPS improves from 0.603425 to 0.554519. Motion Smoothness also increases (0.991386 to 0.991999) and Background Consistency rises from 0.937983 to 0.942890.
  • Wan2.2 results are mixed but favorable on the main axes: SMI obtains the best Motion Smoothness (0.935492), Camera-Constrained Consistency (46.9309), Overall Visual Quality (51.4909), PSNR (11.564875), LPIPS (0.587550), and Aesthetic Quality (0.445095) among the compared methods, but its Background Consistency (0.917329) is below the Base (0.930115) and several compression baselines.
  • Latency cost: The understanding model adds overhead. SMI runs at ×1.38 the Base latency on HY1.5 and ×1.31 on Wan2.2, whereas Deep Forcing is faster than the Base (×0.81 and ×0.75 respectively). The authors expect this overhead to shrink as generation and understanding become unified through shared representations.
  • Component ablation (HY1.5): Removing reliability-aware filtering gives PSNR 13.232422 and sparsity 81.5789%; removing spatial clustering gives PSNR 11.050455; removing spatial clustering and sparsification together gives 12.985179 PSNR with 35.7368% sparsity; disabling all three gives 0.937700 Background Consistency, 0.446470 Aesthetic Quality, 12.812125 PSNR, 0.581249 LPIPS, 0.0000% sparsity, and ×1.27 latency. The paper states that without spatial clustering, sparsification can discard important spatial evidence and consistency declines.
  • Fine-tuning matters greatly: Without operation-oriented fine-tuning, Qwen3.5-4B yields 42.8380 Camera-Constrained Consistency, 44.7440 Overall Visual Quality, 9.396166 PSNR, 0.714904 LPIPS, and 86.3444% sparsity, versus 67.0200, 65.3677, 13.309393, 0.554519, and 83.6842% for the fine-tuned SMI.
  • Reported operational settings: Seven historical chunks are geometrically shortlisted before MLLM retrieval; up to P_r memories can be returned (including an empty set); both backbones retain the two most recent chunks as context; reliability filtering scores chunks 1–10 and discards anything below 6.
  • Not reported in the supplied content: the numeric size N_D of the training dataset, the sparsification trigger threshold T_s, details of the user study (Appendix D), the full baseline configuration (Appendix E), and the full dataset-composition statistics (Appendix F).

Methodology in Plain English

The world model generates video chunks one at a time, conditioned on a few recent chunks and the current action. SMI sits alongside it as a memory controller. When a new chunk is generated, the controller first decides whether it depicts a region already in memory: camera position narrows the candidates to the nearest cluster prototypes, and the MLLM decides on a match. Matched chunks join an existing cluster; unmatched ones create a new cluster. Inside a cluster, chunks are ordered by generation time, and once a cluster reaches its trigger threshold the MLLM compares them and deletes intermediate chunks that add no new spatial structure, object state, or occlusion information. For the next generation step, camera geometry shortlists the seven most spatially overlapping memories, and the MLLM — seeing the recent context and the current action — chooses which of these to actually reuse, possibly none. Before any new chunk is stored, the MLLM rates its visual reliability from 1 to 10; chunks below 6 are thrown away so that visual drift is not recycled into later frames. The MLLM itself is Qwen3.5-4B, fine-tuned with AdamW at a learning rate of 2×10⁻⁵ and global batch size 64 on data built from an annotation pipeline that moves from real inference trajectories to human seed labels, to teacher-MLLM labels, to human correction.

Why This Matters

The paper reframes memory management in world models from a fixed heuristic (camera geometry or embedding similarity) into a learned reasoning task performed by a multimodal understanding model, which is a conceptual shift for how long-horizon generative systems can be governed.

Real-world applications:

  • Interactive entertainment and games: persistent, explorable 3D worlds where revisiting a location must show the same objects and layout.
  • Embodied simulation and robotics: training environments where an agent accumulates a long observation history and must not corrupt its own world state with erroneous predictions.
  • Virtual production and digital twins: long camera takes through a simulated space where consistency across revisits determines usability.
  • Streaming generative content pipelines: reducing storage and context cost while keeping generated footage spatially coherent.

Industry relevance: the work sits directly at the intersection of two active industrial efforts — video world models for interactive simulation and unified understanding-and-generation architectures. Any team building long-horizon video generation faces the storage and consistency problems SMI targets, and the framework is designed to plug into different backbones rather than being tied to one generator.

Future Directions

  • End-to-end joint training: the framework is not trained jointly with the world model, so it does not improve the backbone's intrinsic generation capability and remains bounded by the base generator's visual quality, controllability, and stability. The authors propose learning generation, understanding, and memory management jointly in a shared representation space to cut overhead and let the two capabilities reinforce each other.
  • Reinforcement learning for memory control: applying RL to long-horizon world models with unified understanding and generation, to strengthen the understanding model's long-term memory management.
  • Better and larger datasets: constructing higher-quality, more comprehensive datasets so understanding models can learn broader memory-management capabilities.
  • Long-horizon streaming understanding: the current task setting does not explicitly address it, and may remain insufficient for memory tasks requiring complex long-term dependencies or logical reasoning.

Target Audience

Researchers and engineers working on video generation, world models, and interactive or embodied simulation; MLLM researchers interested in understanding models acting as controllers rather than as passive perceivers; and practitioners who need long-horizon generation to stay spatially consistent while keeping memory budgets bounded. Readers should be comfortable with autoregressive video generation, diffusion transformer backbones, camera geometry, PSNR/LPIPS, and supervised fine-tuning of multimodal models.

Authors’ abstract

Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observations conditioned on user actions and historical memory. However, as memory sequences grow longer and their structures become increasingly complex, managing long-range spatial context becomes increasingly challenging, calling for a more intelligent and systematic memory-management strategy. Building on the advancing spatial reasoning capabilities of multimodal large language models (MLLMs) and the broader vision of unified models, we propose Spatial Memory Intelligence (SMI), the first framework to systematically employ an understanding model for spatial-memory management in long-video world models. SMI introduces four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering. Extensive experiments across multiple baselines, benchmarks, and world-model backbones demonstrate the effectiveness and generalizability of SMI, achieving comprehensive improvements in memory sparsity, generation stability, and spatial consistency.

Read the original paper