Skip to content
AI.info

Research

Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models

Overview Research area: Machine learning for world models, specifically episodic memory retrieval for video diffusion world models, drawing on retrieval-augmented generation and information-theoretic

Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models
arXiv
2609.34677
Published
2026-09-28
Authors
Beomsu Kim, Chieh-Hsin Lai, Bac Nguyen, Amir Bar, Jong Chul Ye, Yuki Mitsufuji

AI summary

Overview

Research area: Machine learning for world models, specifically episodic memory retrieval for video diffusion world models, drawing on retrieval-augmented generation and information-theoretic relevance.

Technical level: Advanced. The paper formalizes recall as a latent-variable problem, proves a predictive relevance bound, and derives a training objective from a future-aware posterior.

Scope: The paper proposes Future-Aware Recall (FAR), a framework that learns which past observations to recall and which retrieval cues (time, pose, vision, audio) to trust, and evaluates it on LoopNav, SoundSpaces, and AI2-THOR.

Paper details: "Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models," by Beomsu Kim, Chieh-Hsin Lai, Bac Nguyen, Amir Bar, Jong Chul Ye, and Yuki Mitsufuji (KAIST, Sony Group Corporation, Imperial College London), arXiv:2609.34677v1 [cs.LG], 28 Sep 2026, CC BY 4.0.

What This Paper Is About

World models predict future observations from the current query and actions, but the information needed for a prediction can lie far in the past, outside the current view. Episodic memory keeps past observations available for later recall, but as memory grows, the system must decide which stored memories are useful for the current prediction and which of several available retrieval cues can actually find them. Existing world models use fixed relevance rules such as temporal recency, pose-based field-of-view overlap, or visual embedding similarity, and the paper argues that similarity under a particular cue does not necessarily reflect predictive usefulness, and that cue reliability varies across environments and queries.

Key Contributions

  1. A predictive definition of memory relevance. FAR defines the value of a recalled context by its predictive utility, the conditional log-likelihood of the realized future given that context, approximated in the video diffusion instantiation by negative diffusion prediction loss. Proposition 1 shows that this quantity lower-bounds the mutual information between the future and the recalled context given the prediction query, so maximizing it is a principled relevance objective.

  2. Future-aware supervision of a future-blind retriever. During training, the observed future is combined with the retriever's own score via Bayes' rule to form a future-aware posterior over contexts, which serves as a stop-gradient teacher. The retriever is trained by KL divergence to this posterior, so it learns predictive relevance while remaining unable to see the future at inference.

  3. Adaptive multi-cue fusion. FAR learns a separate relevance function per cue (time, pose, vision, audio), standardizes cue scores across episodic memory to handle different numerical scales, and combines them with query-dependent weights produced by a learned gate. This lets the model decide per query which cues to trust rather than committing to one cue or a fixed combination.

  4. A video diffusion world-model instantiation and evaluation across three settings. Memory-side encoders for high-dimensional cues are contrastively pretrained and frozen so historical keys can be cached, lightweight query-side adapters and fusion components stay trainable, and recall is temporally chunked so Top-K slots are not spent on redundant nearby frames. FAR is evaluated on LoopNav, SoundSpaces, and AI2-THOR against temporal, WorldMem, and LongLive-RAG retrieval baselines.

Main Findings

  • Learned recall beats hand-designed recall with the same cues (LoopNav). At loop closure, the learned visual retriever outperforms LongLive-RAG with 17% lower DreamSim, and the learned metadata retriever outperforms WorldMem with 19% lower DreamSim, despite using the same respective cue types. The temporal baseline performs worst. The paper attributes the gap to cue similarity not implying predictive utility: WorldMem retrieves a frame with high field-of-view overlap but a heavily occluded view of the goal, whereas the learned metadata retriever selects a clearer predictive view.

  • Multi-cue fusion tracks the better single cue (LoopNav). The Multi-Cue variant fusing metadata and visual cues closely follows the better of the metadata-only and visual-only variants across the horizon, which the authors read as adaptive fusion emphasizing whichever cue is more informative for the current prediction.

  • Audio helps when geometry and pose are ambiguous (SoundSpaces). With dense periodic 360° scans, metadata alone is a strong signal and multi-cue fusion of metadata and audio yields modest gains. Under the sparser endpoint-scan corpus, the advantage of Multi-Cue grows with return-phase length, particularly in LPIPS and DreamSim. In the corridor example, temporal retrieval picks the most recent observation, WorldMem picks the same context due to high field-of-view overlap with the goal pose, metadata-only FAR picks an observation farther along but obstructed by corridor geometry, and adding audio retrieves a context with a clearer view of the region around the goal.

  • State-consistent recall under world changes (AI2-THOR). With object manipulation making memories potentially stale, Multi-Cue FAR fusing metadata and vision substantially outperforms temporal and geometry-based recall across state-change magnitudes for both surface changes and container reveals. The advantage is especially pronounced for container reveals, where the current observation contains little information about hidden contents. Performance decreases for all methods when current and stale states differ only locally, which the authors connect to limited sensitivity of the diffusion prediction objective to small visual changes.

  • Counterfactual, history-dependent futures (AI2-THOR). Holding the current observation and action fixed while varying history, FAR renders the tomato inside the refrigerator only when it was previously placed there. Temporal and WorldMem fail to preserve the interaction-dependent state, with WorldMem additionally hallucinating the tomato at its stale table location.

  • Off-scene dynamics prediction (AI2-THOR two-agent corridor). Each recalled context consists of two observations separated by 1.5 seconds. Temporal achieves 41.9% trajectory-level accuracy and WorldMem 32.4%, while FAR with metadata alone reaches 67.1% and Multi-Cue (metadata plus an A₂ observation cue) reaches 94.6%. WorldMem retrieves a geometrically relevant but dynamically uninformative memory.

  • Ablation on LoopNav (Table 1, average loop closure errors).

    Retrieval PSNR ↑ LPIPS ↓ DreamSim ↓
    Temporal (NWM) 13.229 ± 0.114 0.582 ± 0.007 0.200 ± 0.004
    LongLive-RAG 14.847 ± 0.121 0.463 ± 0.006 0.133 ± 0.003
    WorldMem 15.372 ± 0.112 0.448 ± 0.006 0.130 ± 0.003
    FAR with Vision Cue, Enc. Pretrained 16.248 ± 0.112 0.390 ± 0.006 0.100 ± 0.003
    + MLP Adapter 16.480 ± 0.108 0.374 ± 0.006 0.091 ± 0.002
    + Meta, λ = 0.5 16.912 ± 0.117 0.353 ± 0.005 0.088 ± 0.002
    + Meta, λ learned 17.267 ± 0.116 0.337 ± 0.005 0.082 ± 0.002

    A frozen pretrained vision encoder already substantially outperforms baselines across all three metrics. Adapting the query representation with an MLP improves 16.25 to 16.48 PSNR, with gains in LPIPS and DreamSim. Fixed metadata fusion at λ = 0.5 raises PSNR to 16.91, and learning the fusion weights achieves the best performance across all metrics, indicating metadata and visual cues are complementary and adaptive weighting is preferable to uniform fusion.

  • Where errors concentrate (LoopNav). Frame-wise PSNR, LPIPS, and DreamSim between generated and ground-truth return trajectories are highest near the middle of the return phase, where relevant observations are typically farthest in time and memory is most sparse. Curves report the mean across trajectories with shaded regions indicating 95% bootstrap confidence intervals.

Methodology in Plain English

The setup is simple to state. An agent explores an environment, then returns; the world model must generate the return-phase video using exploration-phase observations as its memory pool. Each stored observation is tagged with retrieval cues, such as time, pose, a visual representation, or an audio representation. For each prediction step, an external retriever picks a compact Top-K set of past observations, and the world model conditions on that recalled set to predict the next observation.

The hard part is that at inference time the future is unknown, so the retriever cannot know which memories will actually help. The researchers turn this into a training trick. Because the future is observed during training, they can score every candidate memory by how much it improves the likelihood of what actually happened. They formalize this as predictive utility, the log probability the world model assigns to the realized future given a recalled context, and show that this quantity lower-bounds the mutual information between future and context given the query. For the diffusion world model, exact likelihood is too expensive, so they substitute negative diffusion prediction loss, averaged over four diffusion timesteps, as a tractable proxy; memories that yield lower prediction loss receive more credit.

They then combine two signals with Bayes' rule: the retriever's own future-blind relevance score, and the future-aware predictive credit. The resulting posterior is treated as a fixed teacher, and the retriever is trained by KL divergence toward it with a stop-gradient, so no gradient flows through the teacher. The effect is to distill information about the future into a retriever that will not have it at inference.

For multiple cues, each cue produces its own observation-level relevance score, standardized across the memory to align scales. A learned gate produces query-dependent weights over cues, these are summed into a fused score, and the Top-K highest-scoring observations form the recalled context. The paper adapts a candidate-level latent-variable approximation from EMDR² to handle the otherwise intractable sum over possible context sets. To avoid filling Top-K slots with redundant nearby frames, memory is partitioned into temporal chunks and only the highest-scoring representative from each chunk is eligible. High-dimensional cue encoders on the memory side are contrastively pretrained and frozen so historical keys can be cached; query-side adapters and the cue-fusion components remain trainable.

Why This Matters

Impact on research. The paper separates two things that prior external-memory world models conflate: whether a memory looks similar under a cue, and whether it actually helps predict the future. It imports the latent-variable retrieval-training principle from retrieval-augmented generation (specifically EMDR²) into episodic recall for world models, and it makes cue selection a learned, query-dependent decision rather than a fixed heuristic. Table 2 in Appendix A positions this as the first entry in its comparison to combine a world model, an external retriever, prediction-supervised training, and adaptive cue fusion over time, pose, vision, and audio. The framework is complementary to persistent-state memory (recurrent latent state or persistent 3D representation) and to internal long-context mechanisms using attention, routing, compression, linear attention, or recurrent memory, so it can be layered onto existing systems rather than replacing them.

Real-world applications.

  • Embodied navigation and household robots. The AI2-THOR results target agents that move objects between surfaces and containers and later revisit those locations; recalling the current state rather than a stale one is exactly what a service robot needs.
  • Audio-visual simulation and content generation. The SoundSpaces corridor example shows spatial audio disambiguating memories when pose and visual geometry are ambiguous, relevant to generating or re-rendering navigable 3D scenes for games and film.
  • Long-horizon video generation. The LoopNav setting is a video diffusion world model generating a return trajectory from exploration footage; learned recall reduces loop-closure error against the temporal, WorldMem, and LongLive-RAG baselines.
  • Multi-agent and off-scene reasoning. The two-agent corridor task predicts where an intermittently visible patroller will be when the observer looks down the corridor, which maps onto tracking entities that leave and re-enter the field of view.

Industry relevance. The author list includes Sony Group Corporation alongside KAIST and Imperial College London, and work was done while Beomsu Kim was an intern at Sony. The techniques apply to any system that must maintain long-horizon consistency while generating video or simulating embodied interaction under a bounded context window, which is a central cost and quality problem for video generation, gaming, simulation, and robotics.

Future Directions

  • Memory lifecycle beyond recall. The paper explicitly leaves memory writing, compression, forgetting, and higher-order interactions among recalled memories to future work; it only studies recall from an external episodic memory.
  • Better utility functions. The authors note their diffusion-loss utility may underweight semantically important local changes, which is consistent with the observed drop in AI2-THOR accuracy when current and stale states differ only locally, and suggest alternative task-aware utility functions.
  • Reducing the cost of future-aware training. FAR requires informative retrieval cues and additional training-time computation for predictive-utility evaluation, both flagged as limitations.
  • Richer evaluation and integration. Experiments use controlled simulations; the authors call for evaluation in richer real-world settings and integration with learned memory formation and persistent-state representations.

Target Audience

Researchers and advanced practitioners working on world models, long-horizon video diffusion, and embodied AI. It will be most useful to readers already comfortable with latent-variable models, information-theoretic objectives, and diffusion training, and to those building retrieval-augmented generation systems who want to see the discrete-retrieval supervision idea applied to an agent's own episodic history rather than a document corpus. Practitioners in robotics, simulation, gaming, and video generation who need long-horizon consistency under bounded context will find the empirical comparisons and the ablation directly actionable, while readers looking for an introductory treatment should expect to work through the formalization in Section 3.

Authors’ abstract

World models predict future observations from current experience and actions, yet prediction can depend on observations seen far in the past. Episodic memory preserves past observations for later recall; however, as memory accumulates, it raises a fundamental question: which memories are useful for the current prediction, and which available retrieval cues should be trusted to find them? This is challenging because fixed criteria based on recency, pose overlap, or visual similarity can be unreliable across environments and queries. We propose Future-Aware Recall (FAR), a framework that learns episodic recall from future-aware predictive supervision and adaptive multi-cue scoring. During training, FAR measures predictive utility by the conditional log-likelihood of the realized future given recalled context, approximated by negative diffusion prediction loss, and uses it to train a retriever that remains future-blind at inference. The retriever learns cue-specific relevance and automatically determines which available retrieval cues, such as time, pose, vision, and audio, to trust for each query when selecting memories. Across three complementary settings, FAR outperforms hand-designed recall even with the same retrieval cues, automatically adapts which available cues to trust, and recalls the right history as the world changes. Together, these results establish FAR as a flexible, principled approach to episodic memory access in world models.

Read the original paper