Skip to content
AI.info

Research

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Overview Research area: Robotics — long-horizon robot manipulation, action-conditioned world models, modular vision-language-action (VLA) policies, and active imitation learning. Technical level: Adva

RoboCoach: World Models as Active Coaches for Compositional Robot Skills
arXiv
2609.39685
Published
2026-09-30
Authors
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

AI summary

Overview

Research area: Robotics — long-horizon robot manipulation, action-conditioned world models, modular vision-language-action (VLA) policies, and active imitation learning.

Technical level: Advanced. The paper assumes familiarity with VLA backbones, LoRA adapters, flow-matching video generation, camera calibration, and imitation-learning evaluation protocols.

Scope: The paper proposes and evaluates RoboCoach, a world-model-guided coaching loop that uses imagined rollout failures to decide which subtask demonstrations to collect and which reusable skill experts to update.

What This Paper Is About

Long-horizon manipulation tasks (the paper's examples include preparing tea and setting a table) are sequences of reusable skills, and each skill changes the state the next one runs in, so a single local error can derail the whole task. Collecting more end-to-end demonstrations to fix this is expensive, especially on physical robots, where every trajectory costs hardware time, environment resets, and safety oversight. The paper's goal is to turn a world model into an "active coach" that answers two questions per improvement round: which skill should the next demonstrations target, and which policy component should receive the update.

Key Contributions

  1. Coupled self-improvement decisions. The paper formalizes embodied self-improvement around two linked choices — which demonstrations to acquire and which policy components to update — and instantiates them with subtask–expert pairs, making reusable skills explicit targets for both data acquisition and adaptation.
  2. The RoboCoach framework. RoboCoach runs a Route–Imagine–Diagnose–Improve (RIDI) loop that connects CoachWorld (a shared action-conditioned world model), progress-based diagnosis, and selective expert adaptation, converting recurring imagined failures into targeted supervision.
  3. CoachWorld as a cross-embodiment predictive model. CoachWorld conditions future-video prediction on sparse visual history, task instructions, and calibrated two-slot end-effector trajectories across single-arm and bimanual embodiments, and is trained on a mixture described as about 500 hours of data.
  4. Evidence that targeting and update location both matter. Experiments separate data acquisition from update location and show that applying identical acquired demonstrations to corresponding skill experts outperforms a shared global adapter, and that coached experts can be recombined into held-out compositions.

Main Findings

  • Imagination tracks deployment. Across 22 frozen task–policy checkpoint pairs, imagined and deployed success rates correlate with a Pearson correlation of r = 0.820 and a Spearman correlation of ρ = 0.840.
  • Real-robot gains under a fixed budget. With 150 additional subtask demonstrations per platform, success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX. Under the same budget, Single VLA + Uniform reaches 30.0% and 47.5% respectively.
  • Normalized Budget AUC. RoboCoach scores 53.1 versus 18.9 on Franka and 72.7 versus 46.3 on AgileX.
  • Simulation results. RoboCoach reaches 71.2% on LIBERO and 68.0% on RoboTwin, the highest final success among the four compared conditions.
  • Targeted acquisition alone is not enough. With a shared adapter, targeted acquisition changes final success relative to uniform acquisition by +0.8 pp on LIBERO and −4.8 pp on RoboTwin. Applying the same targeted demonstrations to selected experts instead adds 3.4 pp and 13.2 pp.
  • Target selection matters within the modular system. World-model targeting exceeds random target selection by 2.8 pp and 12.8 pp.
  • Held-out composition generalization. Coached experts average 35.0% success on four held-out compositions, versus 0% for the shared-policy baseline updated with uniformly acquired demonstrations. Individual routes: 13/20 (65%) on the AgileX continuation combining table setting with tea making, 3/20 (15%) on the Franka green-button-pressing-plus-lamp-cord-pulling combination, 5/20 (25%) on the AgileX shoe-placement reversal, and 7/20 (35%) on the Franka button-order reversal using the same press expert.
  • World-model fidelity. On DROID-180, mixed-domain CoachWorld achieves the best LPIPS (0.0996), FVD (51.71), all three trajectory metrics, and instruction-following score (0.880) among compared models (Ctrl-World, Cosmos 3, OSCAR-2B, and the two CoachWorld variants). On Lab Franka-180 it achieves the best LPIPS (0.1712). Mixed-domain training improved most visual and trajectory measures relative to single-domain training at matched training volume.
  • Progress-judge reliability. RoboMeter achieves the best scores on reference video, with Switch MAE of 0.414 s, 82.73% Recall@1.0s, and 88.21% outcome F1, and is used as RoboCoach's progress judge. On CoachWorld-generated observations it reaches terminal-stage Spearman ρ = 0.638, outcome F1 83.84%, Switch MAE 0.815 s, and 62.73% Recall@1.0s.
  • Stated limitations. Physical modeling degrades when the manipulated object is occluded or when contact and collision occur outside the camera's view; agreement across generation seeds reduces sensitivity to stochastic variation but cannot rule out systematic world-model bias; and the expert library is predefined by skill semantics.

Methodology in Plain English

The system decomposes a long-horizon instruction into a sequence of atomic skills — such as picking, placing, or pressing — and assigns each skill to a reusable expert. Every expert is the shared VLA backbone plus its own LoRA adapter, so one expert can serve multiple task-specific subtasks (for example, the same press expert handles pressing a red button and a green button).

The coaching loop has four stages. Route takes the instruction, current observation, skill vocabulary, and the list of already-completed subtasks, and selects the next unfinished subtask and its expert. Imagine rolls that expert out in closed loop inside CoachWorld: the expert predicts an action in end-effector space, a lightweight adapter turns it into a future end-effector trajectory, and CoachWorld generates the next video chunk for each calibrated camera independently; each chunk is committed before the next policy or judge query. Diagnose uses a progress judge to estimate whether the active subtask has been completed; if it reaches the completion threshold, control returns to the router, and if it hits its time limit, the trial terminates and the active subtask–expert pair is recorded as a timeout. Each initial-state trial is run with three world-model seeds, and the diagnosis is retained only when at least two seeds agree. Improve ranks candidate subtask–expert pairs by a task-balanced first-timeout mass (so each task contributes equally), selects the top M = 2 pairs, requests B_q / M accepted subtask demonstrations for each, and fine-tunes the corresponding LoRA adapters using the new demonstrations pooled with a smaller replay sample from existing data. The backbone and unselected adapters stay frozen, and updated experts re-enter the next round.

Experiments use ten LIBERO-Long tasks and five RoboTwin 2.0 tasks in simulation, plus real-robot tasks on a Franka Research 3 and an AgileX dual-Piper platform. π_0.5 serves as the shared VLA backbone for LIBERO and Franka, and MolmoAct2 for RoboTwin and AgileX. Each task is evaluated over 50 trials in simulation or 20 trials on real robots. Coaching rounds add B_sim = 100 accepted subtask demonstrations in simulation or B_real = 50 per real platform, and results are reported at cumulative budgets of {0, B, 2B, 3B}. Four conditions are compared: Single VLA + Uniform, Single VLA + WM-targeted, Modular + Random, and RoboCoach. CoachWorld is initialized from Wan2.2 TI2V-5B, runs at 5 Hz with 512 × 768 images and a 5:3 history-to-future latent ratio, uses AdamW with cosine decay (learning rate 1×10⁻⁵, then 5×10⁻⁶ in the late stage) and batch size 40, and is trained by conditional flow matching. Main training and offline evaluation used eight NVIDIA H100 GPUs.

Why This Matters

The paper reframes the role of a world model: instead of only simulating extra experience, the model is used to decide where scarce real experience is most valuable. It also provides evidence that the location of an update matters independently of the data: identical targeted demonstrations produced different outcomes when routed to skill experts versus a shared adapter.

Real-world applications:

  • Household and service robotics — the paper's own real-robot tasks include table setting, tea making, and lamp switch and cord operation, all of which are long-horizon, multi-skill routines.
  • Warehousing and logistics — placing and stacking tasks (the RoboTwin tasks include arranging blocks by color and by size, placing bottles into a dustbin, and stacking bowls and blocks) map to pick-place-and-arrange workflows.
  • Manufacturing and machine tending — button pressing and drawer or box closing are repeated atomic skills that benefit from reusable experts.
  • Cost-constrained robot deployment — the central premise is reducing human data-collection effort, hardware time, environment resets, and safety oversight per additional trajectory.

Industry relevance: The approach targets the practical bottleneck of scaling robot data collection. Because it operates at the level of reusable skills and only trains small LoRA adapters, it suggests a maintenance-style workflow in which a deployed fleet requests particular demonstrations for particular skills rather than full retraining. The expert library is deliberately predefined by skill semantics, which the authors identify as a limitation and an open direction.

Future Directions

  • Handling occlusion and out-of-view interaction. The authors state that prediction and physical modeling become difficult when the manipulated object is occluded or when contact and collision occur outside the camera's view, and that preserving task-relevant evidence remains challenging in those regimes.
  • Reducing systematic world-model bias. Seed agreement across three generation seeds reduces sensitivity to stochastic variation but cannot rule out systematic bias, leaving open how to detect and correct it.
  • Learning the expert library structure. The expert library is currently predefined by skill semantics; learning its structure, and reducing reliance on human demonstrations, are named as open directions.
  • Extending budget and round analysis. The reported real-robot comparisons only include Single VLA + Uniform versus RoboCoach, attributed to hardware effort, and the paper points to per-round results and acquisition traces in its appendix.

Target Audience

Researchers and practitioners in robot learning who work on long-horizon manipulation, world models, VLA policies, modular or mixture-of-experts adaptation, and active or corrective imitation learning. The paper is most useful to readers already comfortable with policy evaluation protocols and adapter-based fine-tuning, since it combines simulator fidelity measurement, progress-judge calibration, and fixed-budget real-robot comparisons into a single empirical study. Readers looking for an accessible introduction to world-model-based robotic control will find the RIDI loop conceptually clear, but the evaluation details and appendix material assume a robotics research background.

Authors’ abstract

Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. A practical self-improving system must decide both what to teach next and where to apply that supervision. We present ROBOCOACH, a world-model-guided coaching framework that uses imagined failures to guide demonstration requests and expert updates. Its Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts inside COACHWORLD, our shared action-conditioned world model, and uses a progress judge to record the first subtask that fails to complete. Aggregated records select which subtask demonstrations to acquire and which expert adapters to update. Across two simulation suites and two real-robot platforms, imagined and deployed success correlate over 22 task-policy pairs (rho = 0.840). Controlled comparisons show that our coaching method outperforms matched baselines under matched data budgets and update schedules. With only 150 additional subtask demonstrations, success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX. The coached experts also transfer to four held-out compositions, achieving an average success of 35.0%, compared with 0% for a shared-policy baseline updated with uniformly acquired demonstrations. Together, these results show that world models can serve as active coaches, turning imagined failures into targeted supervision for modular policy improvement. Project Page: https://robocoach-ai.github.io/

Read the original paper