Research
Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks Overview Research area: Artificial intelligence — specifically LLM-based interactive agents, program/world-model learning
- arXiv
- 2610.11794
- Published
- 2026-10-08
- Authors
- Haoyu Zhao, Zhengxu Yu, Zhiyuan He, Meng Fang, Rasul Tutunov, Haitham Bou-Ammar, Weilin Luo, Jun Wang
AI summary
Memento 3: Model-Based Recursive Self-Improvement through Reflective RulebooksOverview
- Research area: Artificial intelligence — specifically LLM-based interactive agents, program/world-model learning, and recursive self-improvement (RSI).
- Technical level: Advanced. The paper formalizes its method with a partially observable Markov decision process (POMDP), a Bayes-adaptive POMDP belief update, and shortest-path planning over a learned deterministic model, alongside an LLM-driven code-generation loop.
- Scope: The paper proposes a method by which a frozen large language model learns explicit, revisable world models — stored as a natural-language "rulebook" plus compiled executable code — and demonstrates it on ARC-AGI-3 and an Atari Pong case study.
What This Paper Is About
An agent dropped into an unfamiliar environment must figure out both what actions do and what counts as success, but any finite history of interactions is consistent with many different world models that disagree about states the agent has never seen. Memento 3 addresses this by having a frozen LLM maintain an explicit, inspectable hypothesis about the environment's rules in a natural-language rulebook, compile that rulebook into executable code, and repeatedly revise both when predictions turn out wrong. The goal is a model-based route to recursive self-improvement: the agent's own verified world model guides further exploration and planning, improving its behavior without ever changing the underlying LLM's parameters.
Key Contributions
- Code as Model. Memento 3 extends the Memento series from policy memory and procedural memory to semantic world-model memory. A continual loop of observation, reflection, rule revision, compilation, and verification updates a rulebook (persistent semantic memory) and its executable realization (used for prediction and planning), while the LLM's parameters stay fixed (Eq. 5: φ_{t+1} = φ_t = φ).
- A population extension for ambiguous generalization. Instead of committing to one world model, the agent maintains N world models in parallel as rulebook–executable pairs, each sharing the same interaction history. One member is sampled per planning episode, and the resulting transitions are used to check all members; N = 1 recovers the single-model setting.
- Model-based RSI as an explicit definition and objective. The paper defines model-based recursive self-improvement as autonomously revising a persistent world-model hypothesis h_t to improve later prediction and decision-making, and formalizes the objective as maximizing the probability p_B of reaching a goal within an action budget B, with conditional action cost c_B as a secondary objective.
- Empirical evaluation on ARC-AGI-3 plus a control case study. The paper reports that the single-model agent clears every level of all 25 public games and that a learned Atari Pong controller wins 21-0 in three evaluated episodes without further LLM calls.
Main Findings
- ARC-AGI-3 results: On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count.
- Comparison to a frontier model harness: Against Claude Opus 5 with the ARC Prize Standard harness, the reported agent achieves a 59.3-point higher mean RHAE.
- Learned feedback control: In an Atari Pong case study, the learned model supports a feedback controller that wins 21-0 in each of three evaluated episodes with different openings, acting without further LLM calls during execution.
- Verification, not just model confidence, gates acceptance: A revised rulebook–executable pair is adopted only when the LLM judges the code faithful to the rulebook (θ ∈ [[ψ]]) and cell-exact replay reproduces every recorded next observation. The paper states this acceptance is not based on language-model confidence alone, citing prior work showing that ranking by language-model scores can exclude correct candidates.
- A minimum-description-length bias is used for model selection: Among admissible pairs, the method prefers shorter rulebooks and code, framed as an approximation to an argmin over description length L(ψ) + L(θ | ψ), which the paper connects to a maximum a posteriori interpretation under a prior μ₀(ψ,θ) ∝ 2^(−L(ψ)−L(θ|ψ)).
- Planning is reduced to shortest paths under a certainty-equivalent belief: The single-model agent approximates the joint belief with a point mass δ over one accepted hypothesis and its replayed internal state, then plans by computing shortest-path costs V_θ(x) to the inferred goal set Ĝ_θ.
- The paper does not report (in the available content) specific values of N used in the population extension, per-game ARC-AGI-3 scores, the ablations isolating the rulebook and population components described as "targeted comparisons" in the contributions list, action-budget B values, LLM call counts for ARC-AGI-3, or compute/time costs. The available content ends mid-section.
Methodology in Plain English
The agent interacts with an environment it cannot see the source code for. It observes the screen and the actions available, but gets no task description and no access to the underlying state or transition rules.
Each interaction round runs five stages:
- Observation — the agent takes an action and appends the resulting observation to its interaction history.
- Reflection — the frozen LLM looks at the updated history and the current model and chooses one of three options: retain the model, revise the rulebook, or repair the code. A prediction mismatch might come from a wrong rule or from code that implements the rule incorrectly, and these need different fixes.
- Rule revision — if revision is chosen, the LLM rewrites the natural-language rulebook to better explain the evidence. The rulebook describes entities, action semantics, dynamics, goals, and termination conditions, and can deliberately leave unobserved mechanics unspecified.
- Compilation — a rulebook-guided compiler generates executable code from the previous executable, the proposed rulebook, and the updated history. The code defines an internal state space, an initializer, a transition function, an observation function, and a goal set.
- Verification — two checks must both pass. The LLM judges whether the code faithfully implements the rulebook, and replay executes the recorded action sequences to confirm every predicted next observation matches what was actually observed. A repair must explain the earlier evidence, not just the latest discrepancy. If verification fails, the loop returns to reflection.
For planning, the code is treated as a deterministic model with unit action costs. The agent replays the current trajectory to reconstruct its internal state, then computes the shortest path from that state to the inferred goal set; each planned action reduces the remaining distance by one, so a nonempty plan cannot loop. When no plan exists, the agent explores or requests a model revision. The paper notes that exploration can be valuable in itself, because a hypothesis-discriminating action can raise the overall probability of reaching the goal even if it makes no immediate progress.
The population variant keeps several rulebook–executable pairs alive at once, samples one for each planning episode, and checks the transitions that result against all members, so different models can steer exploration while remaining consistent with shared evidence.
Why This Matters
The paper's core claim is that an agent can improve its own competence on an unfamiliar task without updating its underlying language model — the learning lives in an external, inspectable artifact (a rulebook plus code). This matters for research on recursive self-improvement because the "self-improvement" target is a verifiable, auditable object rather than opaque weights, and the paper argues that replay-based verification can reject inconsistent models where LLM confidence alone cannot. It also connects two normally separate lines of work: code-as-world-model methods for planning, and memory-as-learning-state methods from the Memento series.
Potential real-world applications (these are implications of the approach rather than results reported in the paper):
- Robotics and industrial control, where an agent must infer dynamics from limited interaction and can plan against an explicit, human-readable model of the plant.
- Software and UI automation agents that must discover the rules of an unfamiliar application, form hypotheses about its behavior, and act toward a goal under a step budget.
- Scientific and engineering workflows, where agent behavior must be auditable — a natural-language rulebook is inspectable in a way that a fine-tuned model's weights are not.
- Simulation and testing environments, where a learned executable transition model can be replayed and checked against recorded trajectories before being trusted.
Industry relevance: The reported setup uses a frozen LLM behind an application programming interface and stores learning in memory and generated code, which fits deployment patterns where retraining or fine-tuning a frontier model is impractical. The paper's affiliations — University College London, Huawei Noah's Ark Lab (UK), and the University of Liverpool — point to direct interest from industrial research labs in agents that adapt through external artifacts. The reported ARC-AGI-3 numbers (RHAE 100.0, 44% of human actions) and the 59.3-point RHAE advantage over Claude Opus 5 with the ARC Prize Standard harness are the kind of benchmark-facing results that drive adoption decisions in that space.
Future Directions
- Quantifying the population extension. The paper motivates maintaining N world models to avoid premature commitment to one hypothesis and says this is compared against the single-model setting, but the specific N values, their effect on exploration, and the full comparison results are not in the available content.
- Scaling the verification gate. Verification combines an LLM's semantic judgment of rulebook fidelity with exact replay. Since fidelity is a semantic condition assessed by the frozen LLM, the reliability and failure modes of that judgment — beyond replay-checkable observations — remain an open question.
- Extending beyond discrete planning. The paper reports a learned feedback controller for Atari Pong that wins 21-0 in three evaluated episodes without further LLM calls. How far the same rulebook-and-code machinery extends to continuous control, and how it performs across a wider range of control tasks, is not established by this case study.
- Cost and budget accounting. The paper counts every real action toward the budget B but explicitly excludes model construction, simulated rollouts, and replay. What those stages cost in LLM calls, tokens, or wall-clock time is not reported in the available content, and would matter for practical deployment.
Target Audience
Researchers working on LLM-based agents, world models, and program synthesis; reinforcement learning and planning researchers interested in Bayes-adaptive POMDP approximations and belief-space control; and engineers building agents that must adapt to unfamiliar software or control environments without model fine-tuning. Readers need comfort with POMDP formalism, belief updates, and Markov decision process planning to follow the theory sections, though the Code as Model loop itself — rulebook, compile, verify, plan — is describable in plain terms. Readers interested primarily in benchmark numbers will find ARC-AGI-3 and Atari Pong results; readers interested in the recursive self-improvement framing will find the formal definition and belief-space objective in Sections 2.1 through 2.3.
Authors’ abstract
Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects underspecified. It compiles this rulebook into executable code for prediction and planning. Through a continual loop of observation, reflection, rule revision, compilation, and verification, the agent uses prediction errors to refine both the rulebook and its code. Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions. We investigate this process as a model-based route to recursive self-improvement (RSI): the agent autonomously explores the environment, revises its world model, and uses verified updates to guide subsequent interaction and learning, while the underlying LLM remains fixed. A population extension maintains multiple world models in parallel, sharing interaction evidence and using their predictions to guide exploration. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.