Research
What Should World Models Forget? Stratified Retention for Continual Adaptation
Overview Research area: Continual learning and world models — specifically, how a world model deployed in a changing environment should decide what knowledge to keep and what knowledge to revise. Tech

- arXiv
- 2610.03713
- Published
- 2026-10-02
- Authors
- Nishit Anand, Ramani Duraiswami, Dinesh Manocha
AI summary
Overview
Research area: Continual learning and world models — specifically, how a world model deployed in a changing environment should decide what knowledge to keep and what knowledge to revise.
Technical level: Intermediate. The paper is a position paper with a conceptual framing and a proposed evaluation protocol; it contains no new experiments and only light formal notation.
Scope: A one-sentence scope: The paper argues that world models violate the stationary-target assumption underlying standard continual-learning metrics, proposes stratifying world-model knowledge by how long it stays true, and proposes a two-part evaluation protocol called differential retention.
What This Paper Is About
Continual learning normally treats any drop in performance on previously seen data as a failure, because in supervised classification a correct label stays correct forever. The authors argue that world models do not satisfy this assumption: their prediction target is the environment itself, which changes, so some knowledge that was accurate when learned later becomes false and must be discarded. The goal is to give the field a way to tell apart a world model that has correctly revised outdated knowledge from one that has genuinely broken, and to stop ranking a frozen, never-updating model as the best possible continual learner.
Key Contributions
-
Obsolescence as a distinct failure mode. The paper argues that obsolescence — retaining knowledge that has become false — is a failure mode of continual world models distinct from catastrophic forgetting, and that the presence of hard invariants makes it unavoidable (Section 2).
-
Stratification by invariance timescale. It proposes organizing world-model knowledge by τ, the timescale over which a piece of knowledge remains true, with protection increasing in τ (Section 3).
-
Diagnosis of what current evaluation cannot measure. It identifies that standard forgetting metrics cannot distinguish warranted from unwarranted degradation, producing a perverse optimum in which a frozen model scores best (Section 4).
-
The differential retention protocol. It proposes reporting invariant regression testing across the adaptation stream jointly with revision latency, without aggregating the two into a single number (Section 5).
Main Findings
-
World models break the stationary-target assumption. A world model approximates p(s_{t+1} | s_t, a_t). When a road is reconstructed, an object is relocated, or a scene is rearranged, knowledge accurate at acquisition becomes false, and a model that keeps asserting it is stale rather than faithful.
-
Three phenomena are reported as one. The paper separates catastrophic forgetting (loss of knowledge that remains true — a defect), deliberate unlearning (removal of true but unwanted knowledge — well studied), and obsolescence (retention of knowledge that has become false — revision is mandatory).
-
Existing benchmarks avoid the problem structurally. Continual-Dreamer is evaluated on sequences of Minigrid and Minihack tasks, and WMAR is evaluated on Procgen and Atari sequences, where ground truth within a task is fixed throughout. DRAGO studies task sequences that share the same dynamics and differ only in reward — exactly the setting where knowledge cannot go out of date.
-
Where environments do change, world models cope poorly. PlaNet and DreamerV2 are reported to adapt poorly to local changes in the environment. In DreamerV3 with an unbounded replay buffer, the world model retains essentially all measurable knowledge of earlier tasks while the policy degrades regardless, suggesting the failure mechanism need not be loss of stored information.
-
Consolidation methods protect on the wrong basis. Elastic weight consolidation, replay, and distillation-based consolidation decide what to protect by importance to past performance, which is unrelated to how long knowledge stays true. The result fails in both directions: physical and geometric invariants are under-protected, while a replay buffer that stores an old room layout keeps reasserting it (over-protection of lower strata).
-
No formal result guarantees invariant preservation. LeJEPA identifies the isotropic Gaussian as the embedding distribution minimizing downstream prediction risk and enforces it with a regularizer that provably rules out representational collapse, and LeVJEPA extends this to video. The authors state that no comparable proof exists that a representation will preserve a physical or geometric invariant once it starts updating.
-
Standard metrics reward the frozen model. There is a perverse optimum: the best possible forgetting score belongs to a model that never updates at all. In a stationary benchmark this is harmless; in a changing environment the frozen model is precisely the failure mode continual adaptation is meant to remedy.
-
Refreshing labels is insufficient for world models. In classification and textual factuality the metric can be repaired by refreshing labels, but matching a refreshed target says nothing about whether invariants survived the update. A model that accommodates a rearranged scene by degrading its representation of object permanence would score well under any updated-target metric.
-
Brittleness under distribution shift is documented. Recoloring the background in Push-T reduces planning success from 50.8% to 6.0% for both LeWM and PLDM. The authors note that if such a model adapted online and recovered, current evaluation could not determine whether recovery came from correct revision of an appearance prior or from degradation of the dynamics model.
-
Scale of current systems. Dreamer 4 trains agents entirely within a learned model of Minecraft from offline data alone; V-JEPA 2 pretrains on over one million hours of video and transfers to zero-shot robot manipulation; LeVJEPA matches V-JEPA 2 at between 5.6x and 20.8x less total pretraining compute. In each case the model is trained once and deployed without further modification.
Methodology in Plain English
This is a position paper, so there are no experiments. The authors proceed by argument and by re-framing existing tools.
First, they contrast world models with supervised classification to show that the "any drop is forgetting" convention rests on a stationary target that world models do not have.
Second, they propose a table of strata ordered by τ (invariance timescale), a property of the environment rather than of the model or its update schedule. The top stratum is invariants with τ = ∞ (gravity, object permanence, impenetrability, geometric consistency, action–effect structure), which must never be revised because any change is a defect. Below that sits long-horizon environment regularity (room layout, road network, typical behavior of other agents), revised slowly on accumulated evidence; short-horizon instance configuration (where objects are, which door is open, current lighting), revised promptly; and immediate transient state (who is moving where, momentary occlusion), for which no retention is expected.
Third, they argue that existing consolidation methods select what to protect by past performance importance rather than by τ, and that this is why protection is misallocated in both directions.
Fourth, they propose the differential retention protocol, which consists of two separately reported quantities. Invariant regression testing fixes a probe set drawn from the τ = ∞ stratum, held out from adaptation data, and re-runs it after every adaptation round; a single evaluation at the end cannot show which update damaged an invariant. Revision latency measures how many observations the model needs before its predictions reflect a fact that has changed, and separately counts facts it never revises at all. The authors emphasize that neither quantity is informative alone, and that the two must not be aggregated.
They note that existing concept-isolated physical probes are directly reusable: WorldBench disentangles individual physical concepts to support per-concept diagnosis, and PhyGround covers thirteen physical laws spanning solid-body mechanics, fluid dynamics and optics. Both are published as evaluations of frozen checkpoints; using them as a regression suite across an adaptation stream requires no new data, only a change in when the probes are run.
Why This Matters
Impact on research. The paper challenges a convention inherited from supervised classification that has shaped how continual-learning results are reported and ranked. If its argument holds, a substantial body of continual world-model benchmarking has been measuring the wrong thing, and the metric that currently rewards a frozen model needs replacing. It also draws a line between the authors' stratification (by the rate at which ground truth moves) and multi-timescale memory systems such as the Continuum Memory System built on Titans, which stratify by update frequency — a distinction the authors argue matters because only the former tells you whether a change was correct.
Real-world applications:
- Mobile robots and autonomous vehicles operating in environments where roads are reconstructed, objects are relocated, or scenes are rearranged.
- Robot manipulation policies transferred from large-scale video pretraining, where an appearance prior may need revision but the dynamics model must not degrade.
- Planning and policy training in imagination for systems where collecting real experience is costly or unsafe.
- Simulation-based training pipelines that must stay accurate as the simulated environment is deliberately changed.
Industry relevance. The companies acknowledged as partial funders — Adobe, Amazon, NVIDIA and Sesame — span content generation, cloud AI infrastructure, accelerated computing, and interactive agents. Any product that ships a learned model of an environment and expects it to keep working after deployment faces the same question the paper poses: which parts of what the model knows are allowed to change, and how would you know if the wrong part changed?
Future Directions
-
Where do the strata boundaries fall? Table 1 is a starting proposal; its boundaries are likely domain-dependent. It remains unclear whether τ can be estimated from observation or must be supplied as domain knowledge.
-
Probes for non-physical invariants. Probes of sufficient rigor exist today only for physical invariants, so realizing the protocol in full requires comparable probes for geometric consistency and causal structure. The authors state that invariant regression testing is implementable today only for the physical sub-case.
-
Are invariants localized? It is not yet known whether invariants are localized in a learned world model. If they are entangled with other knowledge, stratified retention may need architectural support rather than a purely evaluation-side fix.
-
Formal guarantees for retention. Existing formal results describe what a representation must not do (such as ruling out representational collapse); none describes what a representation must retain. A proof that a representation preserves a physical or geometric invariant once it begins updating remains open.
Target Audience
Researchers and practitioners working on world models, model-based reinforcement learning, and continual or lifelong learning — particularly those who build or evaluate systems that must keep operating after deployment, such as robotics and embodied-agent teams. The paper is also aimed at benchmark designers and anyone who reports forgetting metrics, since the central claim is about what current metrics can and cannot measure. Readers from the concept drift and temporal-factuality communities will find the framing familiar but applied to a new setting. Because it is a position paper with no experiments, it is accessible to graduate students and to engineers who need the argument without the mathematics.
Authors’ abstract
Continual learning treats degradation on previously seen data as evidence of failure, a convention inherited from settings with a stationary prediction target, where a correct label remains correct indefinitely. World models do not satisfy this condition. Their prediction target is the environment, which changes, so knowledge that was accurate when acquired may later become false, and discarding it is required behavior rather than a defect. Non-stationary ground truth is well studied in the concept drift literature and in the temporal factuality of language models, but has not been formulated for world models, which are distinctive in that they also encode knowledge that must never be revised. We argue that continual world models require retention stratified by invariance timescale, separating invariants such as physics and object permanence, which must never be revised, from instance-level facts that should be revised as soon as the environment changes. Standard forgetting metrics cannot distinguish a world model that has correctly revised outdated knowledge from one that has suffered catastrophic forgetting, and consequently rank a frozen model highest, while existing physical-reasoning benchmarks evaluate only frozen checkpoints. We propose differential retention, which reports invariant regression testing across the adaptation stream jointly with revision latency, without aggregation.