Research
RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents
Overview Research area: Robotics and embodied AI — specifically LLM/MLLM-based embodied agents, agent self-improvement, long-horizon memory, and cross-embodiment transfer. Technical level: Advanced. T

- arXiv
- 2609.32862
- Published
- 2026-09-26
- Authors
- Jingsong Liang, Shuhao Liao, Shizhe Zhang, Diyuan Hou, Yuxin Cai, Xinjian Deng, Chengyang He, Wenhui Huang, Runjia Tan, Zhidong Wang, Lan Yu, Xuesong Tian, Guillaume Sartoretti, Jie Luo, Yao Mu, Wenjun Wu, Wanhua Li, Chen Lv
AI summary
Overview
Research area: Robotics and embodied AI — specifically LLM/MLLM-based embodied agents, agent self-improvement, long-horizon memory, and cross-embodiment transfer.
Technical level: Advanced. The paper assumes familiarity with foundation models as decision makers, code-as-policy methods, vision-language-action (VLA) backends, filesystem-based agent memory, and embodied benchmarks.
Scope: RoboFoundry reframes the entire supporting system around a frozen foundation model — not just the model weights — as the policy to be evolved from embodied execution experience, and demonstrates the resulting gains across simulation benchmarks and real robots.
What This Paper Is About
Existing embodied agents improve one component at a time — a better interaction harness, a better memory module, a better skill library, or a program-repair loop — while the rest of the stack stays fixed, so improvements do not propagate and the same weaknesses recur. The authors argue that the executable embodied policy is actually the whole system (model plus what information reaches it, what experience persists, and how decisions become physical actions), and that interaction alone does not produce self-improvement unless execution traces are converted into persistent, validated system changes. RoboFoundry is presented as the first embodied agent framework that formulates this process as Self-Evolving System-as-Policy, diagnosing capability gaps and evolving the supporting system through contained, validated edits.
Key Contributions
-
System-as-policy formulation. RoboFoundry treats the agent stack itself as the policy and makes it self-evolving, extending the field's progression from language-model planning to code-as-policy execution to system-as-policy evolution. The foundation model remains frozen; the supporting system
H(M)is edited through filesystem operations (cat,grep,add,modify,delete). -
Capability-guided context–skill co-evolution. Execution traces are attributed to one of six capabilities —
perceive,reason,plan(decision-making) andsave,retrieve,utilize(memory management) — and turned into targeted edits to either the context surface or the skill surface. Task-specific repairs are validated in place; recurring improvements are promoted to the general system only after passing a held-out, capability-conditioned check. -
A shared semantic interface for cross-embodiment reuse. A semantic binding layer separates embodiment-invariant task logic from embodiment-specific execution bindings, so an evolved system can be reused on different robots by swapping bindings, and the same evolved support can attach to different foundation models.
-
Broad empirical validation. Reported gains span embodied decision-making, long-horizon memory, robustness under perturbation, and real-world zero-shot transfer and online evolution — reproduced across multiple foundation models rather than a single backbone.
Main Findings
-
Embodied decision-making (EmbodiedBench, 1,128 tasks across EB-ALFRED, EB-Habitat, EB-Navigation, EB-Manipulation). RoboFoundry improves all five tested backbones. It improves GPT-5.5 by 27.8% (average 56.9 → 72.7), and brings Qwen3.7-Plus to near parity with RoboFoundry (GPT-5.5) at 70.3% vs. 72.7%. Even the frontier-tier GPT-6 Astra gains 8.5% relatively (71.9 → 78.0). Full RoboFoundry beats RoboFoundry-Lite by 11.2% relatively on Qwen3.7-Plus (70.3 vs. 63.2), isolating general-scope evolution as the critical ingredient.
-
A larger context budget alone does not help. Under a 16K-token budget, GLM5.3-Flash drops from 62.0% at 4K to 61.3% at 16K, while RoboFoundry still improves all backbones at 16K by 7.9%–13.1% relative to their baselines.
-
Evolution is data-efficient and traceable. On the EB-Habitat spatial subset, RoboFoundry (GPT-5.5) rises from 21.4% to 78.0% held-in success over six candidate evaluations: three committed edits lift success from 21.4% to 50.0%, two candidates are discarded without changing the system, and the final one is tagged for promotion. Using one quarter of held-in data already reaches 78.0% on EB-ALFRED and 66.0% on EB-Habitat, while three quarters matches the full-data EB-Habitat score of 74.0%.
-
Long-horizon memory (RoboMemArena, 26 tasks in Transfer, Occlusion, Counting, Sequence). With Qwen3.7-Plus as backbone, RoboFoundry reaches 53.5% TSR / 72.8% CSR, outperforming all baselines overall and in all four categories. Against the strongest baseline, PrediMem (38.5/55.2), this is at least a 39.0% relative TSR improvement, with the largest relative TSR gain on Transfer at 213.3% (70.5 vs. 22.5).
-
Context evolution transfers across memory architectures. Applying RoboFoundry to standalone π0.5 raises its average from 26.8/45.6 to 56.2/72.2 (TSR/CSR), and applying it to PrediMem raises its average from 42.0/60.6 to 57.7/75.9, including a 200.0% relative TSR gain for PrediMem on Transfer (22.5 → 67.5).
-
Robustness under distribution shift (LIBERO-PRO, 30 tasks, 50 trials per task and perturbation). RoboFoundry achieves the highest average success rate across the six settings among methods without privileged object poses (91.9% vs. 87.8% for Harness VLA (CC)) and outperforms CaP-Agent0 by 243.8%–679.7% relative across perturbation types, with the largest relative gain (679.7%) on Spatial position perturbations (92.0 vs. 11.8).
-
Perception remains a bottleneck. Given privileged simulator object poses, RoboFoundry† improves across all six settings, reaching 100% on both Object and Goal task perturbations.
-
Real-world zero-shot transfer and online evolution. On an AgileX COBOT MAGIC platform, nesting-doll manipulation reaches 70% success vs. 30% for standalone π0.5. A folding skill evolved only on a blue towel transfers directly to held-out green and pink towels at over 80% success, compared with below 20% for π0.5. On a Unitree G1 humanoid, the shared semantic interface maps "Find drinking water" to navigation actions, and a six-subtask chemistry experiment is completed by preserving and retrieving subtask state across many steps, which standalone π0.5 struggles to finish.
Methodology in Plain English
RoboFoundry keeps the foundation model frozen and treats everything wrapped around it as the thing that gets edited. The setup is an inner–outer loop. In the inner loop, the agent runs tasks using the current general system plus a task-specific system, producing a trace that records system and tool calls, observations, robot states, outcomes, and feedback. In the outer loop, the system reads that trace and asks why it failed or was inefficient, assigning blame to one of six capabilities — perceive, reason, plan, save, retrieve, or utilize. It then edits the responsible surface: the context surface (active in-episode context plus persistent filesystem memory, with operations SAVE, RETRIEVE, UTILIZE) or the skill surface (atomic skills, reusable compositions, and a recovery tree whose nodes are key states and whose edges are skill executions including failures). Candidate edits are re-run and kept only if they repair the failure.
To avoid overfitting to one task, a successful repair is abstracted into a candidate general rule, then tested on held-out traces from other tasks that share the same capability label; it is promoted to the general system only if the capability-conditioned transfer measure is non-negative. This makes self-evolution a controlled loop of diagnosis, intervention, validation, and promotion rather than unconstrained self-rewriting. A semantic binding layer separates task semantics from robot-specific execution, so evolved capabilities can be re-instantiated on other bodies and backbones. The execution backend can be frozen VLA policies, code generated by coding agents, or visuomotor APIs such as CuRobo.
Why This Matters
Impact on research. The paper shifts the unit of optimization in embodied agents from a single module to the whole supporting system, and provides a validation-and-promotion protocol for turning one-off repairs into reusable system capabilities. It also shows that gains come from the evolved system rather than from a stronger backbone, since an open-source backbone (Qwen3.7-Plus) is lifted to near parity with a frontier model running the same framework.
Real-world applications (as demonstrated or indicated in the paper):
- Household and tabletop manipulation, including nesting-doll ordering and precise insertion, and shape-sorter placement.
- Deformable-object manipulation, such as towel folding that generalizes to unseen fabrics with different appearance, geometry, and material.
- Memory-heavy industrial procedures, illustrated by a six-subtask chemistry experiment where earlier steps become persistent prerequisites for later actions.
- Semantic navigation on legged humanoids, demonstrated on a Unitree G1 mapping a natural-language goal to embodiment-specific navigation actions.
- Multi-robot deployment, where a system evolved on one platform is reused on another by swapping embodiment bindings.
Industry relevance. The framework improves agent behavior without retraining model weights, which is relevant to teams that deploy frozen commercial or open foundation models and need robustness to
Authors’ abstract
A foundation model should not act in isolation as an embodied agent. Yet, existing methods often optimize individual components of the agent stack, such as memory, context, skills, or action interfaces, rather than treating the supporting system itself as a unified policy. Moreover, interaction alone does not yield self-improvement unless execution experience is converted into persistent, validated system changes. We therefore propose RoboFoundry, the first embodied agentic framework that formulates this process as Self-Evolving System-as-Policy. RoboFoundry diagnoses capability gaps in decision-making and memory management, converts execution traces into validated task-specific system updates, and promotes recurring improvements to the general system. Evolution operates over two complementary surfaces: a context system that manages active internal context and persistent file-system memory, and a hierarchical skill system that organizes atomic skills, reusable compositions, and failure-conditioned recovery. A shared semantic interface separates embodiment-invariant decisions from embodiment-specific execution, allowing evolved system capabilities to transfer across heterogeneous robots. On EmbodiedBench, RoboFoundry achieves state-of-the-art performance, notably improving GPT-5.5 by 27.8%. It also brings Qwen3.7-Plus to near parity with GPT-5.5 (70.3% vs. 72.7%), showing consistent gains from system-as-policy evolution across foundation models. For long-horizon memory, RoboFoundry outperforms all baselines on RoboMemArena by at least 39.0%, even against methods assisted by external foundation models. On LIBERO-PRO, it further outperforms Cap-Agent0 by 243.8%-679.7% across all perturbation types. In real-world deployments, RoboFoundry demonstrates zero-shot transfer and online evolution across robots and tasks, highlighting its potential for fully autonomous embodied agents.