Research
Recursive Harness Distillation across Agents for Robot Manipulation
Overview Research area: Robotics — vision-language-action (VLA) manipulation, agentic control harnesses, and cross-agent knowledge distillation. Technical level: Advanced (assumes familiarity with VLA

- arXiv
- 2609.33378
- Published
- 2026-09-27
- Authors
- Seungyeon Kim, Junhoo Lee, Minkyu Kim, Baekseung Kim, Nojun Kwak
AI summary
Overview
- Research area: Robotics — vision-language-action (VLA) manipulation, agentic control harnesses, and cross-agent knowledge distillation.
- Technical level: Advanced (assumes familiarity with VLA policies, policy optimization notation, and LLM agent pipelines), though the core idea is explained in accessible terms.
- Scope: The paper introduces Recursive Harness Distillation (RHD), a method in which a strong language-model agent distills its experience operating a frozen VLA policy into a reusable "playbook" for a lighter agent, and recursively refines that playbook using the light agent's execution feedback, validated on SimplerEnv Bridge and three real-world Franka Panda tasks.
What This Paper Is About
Vision-language-action models give robots broad manipulation abilities, but they can fail when a task requires diagnosing what went wrong mid-execution and adjusting behavior. A capable language-model agent can learn how to intervene on such a policy through trial and error, but running that expensive agent for every deployment is costly, and simply handing a lighter agent the same intervention tools does not work — in the real-world experiments here, Luna with the policy but no playbook succeeded on 0 of 25 trials in each of three tasks. The paper's goal is to accumulate a strong agent's intervention expertise as external, reusable guidance so a cheaper agent can operate the frozen policy effectively without any model parameters being updated.
Key Contributions
- Recursive Harness Distillation (RHD): a teacher–playbook–recipient loop in which a strong agent (Astra) interacts with a frozen VLA, distills its experience into a playbook, and then revises that playbook using the light agent's (Luna's) execution feedback across rounds — with all model parameters held fixed.
- A concrete intervention interface: three sites along the VLA's computation — instruction editing (up to 512 characters), attention-logit bias on visual tokens (strength 0–4), and output/action correction over up to five control steps — each with an identity setting that leaves the VLA unchanged.
- A policy-optimization framing: the playbook is treated as the optimizing variable over the deployment agent's induced execution policy, with a finite-horizon performance-difference identity used to argue that revisions should be judged on the histories the recipient actually encounters, not on reproducing the teacher's decisions.
- Empirical validation across simulation and hardware: results on four SimplerEnv Bridge tasks and three real-world Franka Panda tasks, including ablations of each intervention site and a teacher-capability comparison.
Main Findings
- Real-world gains from the playbook: On a Franka Panda with a fine-tuned π0.5 policy, the refined playbook raised Luna's overall success to 64.0%, versus 37.3% for π0.5 alone and 0% for Luna with the policy but no playbook. Per task: Cube-to-Tray 18/25, Cube Stacking 13/25, Button Pressing 17/25, over 25 trials per task (75 total).
- Simulation gains on SimplerEnv Bridge: GR00T alone achieved 41.7%; adding Luna without a playbook raised this only to 43.8%; with the refined playbook Luna reached 66.7%, a 22.9 percentage-point gain over its unguided execution.
- The playbook also helps the strong agent: Astra with the same refined playbook reached 79.2% success, meaning the light agent with the playbook outperforms the strong agent without one.
- Naive distillation can backfire: The initial playbook reduced Luna's success from 43.8% (no playbook) to 31.3%, while the refined playbook raised it to 66.7% — indicating that recursive refinement, not one-shot distillation, drives the gains.
- A strong teacher is necessary: With Luna as both teacher and recipient (separate contexts, same development split, interface, playbook length limit, and rollout/revision limits), the refined playbook achieved only 22.9% on 48 held-out Bridge instances, against 66.7% for an Astra-derived playbook, and below Luna's 43.8% unguided rate.
- Instruction editing matters most among tested sites: Disabling one intervention site at a time with the refined playbook fixed showed the largest drop when instruction editing was removed, suggesting gains extend beyond direct action correction.
- Refinement does not inflate cost: Across the four Bridge tasks (evaluated on the same 12 held-out instances per task), refinement improved success while reducing mean tokens per episode in most tasks relative to the initial playbook.
- Deployment cost gap: Over the full held-out Bridge test set, Luna incurred an estimated $30.03 in inference cost versus $164.61 for Astra ($0.63 vs. $3.43 per episode), making Astra approximately 5.5 times more expensive, based on per-million-token rates of $0.20/$0.25/$0.02/$1.20 (standard input/cache writes/cache reads/output) for Luna and $10/$1/$50 (uncached input/cached input/output) for Astra.
Methodology in Plain English
The researchers keep a pretrained VLA policy frozen and give an external language-model agent a set of "levers" it can pull while the robot runs. One lever rewrites the language instruction given to the policy; a second biases the policy's attention toward a chosen image region; a third edits the proposed end-effector and gripper commands before they are executed. Each lever has a no-op setting, so the agent can also decide to leave the policy alone, and a single decision can combine levers or just inspect candidate actions without advancing the robot.
A strong agent (Astra) first operates the policy directly on a set of development tasks and writes down what worked — when to intervene, how, and what to check afterward — producing an initial playbook. A lighter agent (Luna) then runs tasks using that playbook. Astra reviews recordings of Luna's observations, chosen interventions, and outcomes, and proposes revisions to the guidance. Each proposed revision is frozen and tested on a development batch of n task instances; the first candidate whose empirical success meets or exceeds the teacher's own success rate on the same instances is accepted. This loop repeats, so the guidance adapts to how the specific recipient actually behaves.
Evaluation is kept separate: development used 12 initial configurations per Bridge task and 12 were held out for evaluation (48 instances per split). Episode limits were 120 steps for spoon-on-towel, carrot-on-plate, and cube stacking, and 240 for eggplant-in-basket. For the real world, the researchers collected 50 human-teleoperated demonstrations per task on a 7-DoF Franka Panda with wrist, front, and right-side cameras, fine-tuned π0.5, then froze it. Object positions were randomized before every trial within a 50 cm × 50 cm workspace, and Astra adapted the GR00T-derived playbook to the physical tasks for Luna to use.
Why This Matters
The paper argues that what limits deployed robot manipulation is often not the policy itself but knowing how to steer it — and that this know-how can be written down, improved through use, and shared between agents of different capabilities, all without retraining anything. It offers a practical alternative to running an expensive model on every robot step, and shows that the cheaper agent plus distilled guidance can outperform the expensive agent alone.
Real-world applications:
- Warehouse and logistics picking: a low-cost agent supervising a frozen manipulation policy on randomized object placement, as in the Cube-to-Tray setup.
- Assembly and stacking operations: the Cube Stacking task illustrates recovery from failed grasps and verification of placement before release.
- Facility safety and equipment handling: the Button Pressing task, which involved a red emergency stop button, points to supervised physical interactions where reliability matters.
- Fleet deployment economics: since guidance lives in a playbook rather than in model weights, the same distilled artifact can be shipped to many robots running the same policy.
Industry relevance: the reported cost comparison ($30.03 versus $164.61 over the test set, $0.63 versus $3.43 per episode) speaks directly to the economics of running LLM-driven control at scale, and the fact that no parameters are updated means the approach can be layered on top of an existing deployed policy checkpoint.
Future Directions
- Whether the playbook generalizes to policies other than GR00T-N1.7-SimplerEnv-Bridge and the fine-tuned π0.5, and to tasks outside the four Bridge and three real-world tasks studied.
- How the recursion terminates in practice: the paper notes the acceptance rule does not guarantee termination, and that if no candidate meets the teacher's success target, no update is accepted under that rule.
- What exactly makes a teacher "strong enough" — the Luna-to-Luna result (22.9%) versus Astra-to-Luna (66.7%) shows the hierarchy matters, but the paper does not report a systematic sweep over teacher capability.
- Whether a fixed playbook length limit and the current interface (512-character instructions, attention-bias strength 0–4, five-step action edits) constrain what can be distilled, and whether richer intervention sites would change the results.
Target Audience
Robotics and embodied-AI researchers working on VLA policies and agentic control; engineers building LLM-driven robot deployment pipelines who are weighing inference cost against task success; and machine learning practitioners interested in distillation where the transferred object is external guidance rather than model weights. Readers unfamiliar with policy optimization notation may find Section 3.3 the most demanding part, though the method itself can be followed from Sections 3.2 and 3.4.
Authors’ abstract
A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong agents can discover effective interventions through interaction with these policies. We propose Recursive Harness Distillation to accumulate this experience as reusable guidance across agents. A strong agent distills its experience into a playbook for a light agent, then recursively refines the playbook using the light agent's execution feedback. The resulting playbook enables agents to reuse accumulated intervention knowledge in new task instances without updating model parameters. In real-world manipulation, the harness improves success from 37.3% to 64.0%. On SimplerEnv Bridge, the light agent with the playbook achieves 66.7% success, compared with 41.7% for the GR00T-only baseline, and outperforms the strong agent without a playbook. The same playbook also benefits the strong agent, which reaches 79.2% success. These results demonstrate the feasibility of harness distillation for robotics: intervention experience can be accumulated, refined through execution, and reused across agents to improve manipulation.