Research
Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation
Overview Research area: Robotics — long-horizon compositional manipulation, world-action models (WAMs), vision-language-action (VLA) models, and in-context learning for robot policies. Technical level

- arXiv
- 2610.02368
- Published
- 2026-10-01
- Authors
- Shukai Gong, Xuanran Zhai, Yintianrun Zhang, Ruopeng Cui, Ye Huang, Yiyang Fu, Dexuan Lyu, Chaojie Li, Xinyi Song, Peiwen Lin, Chuang Wang, Mingyuan Jia, Yufan Deng, Jiaxin Fang, Bo Liang, Jiaxin Li, Yuxiang Gao, Hao Liu, Daquan Zhou
AI summary
Overview
- Research area: Robotics — long-horizon compositional manipulation, world-action models (WAMs), vision-language-action (VLA) models, and in-context learning for robot policies.
- Technical level: Advanced (flow-matching generative objectives, VAE latent conditioning, mixture-of-transformers world-model backbones, multi-stage training recipes).
- Scope: The paper proposes ViGAR, a hierarchical framework that splits compositional manipulation into visual subgoal planning and subgoal-conditioned joint video-action generation, and evaluates it on the RoboTwin Clean2Random benchmark and on real AgiBot A2 robot tasks.
What This Paper Is About
World-action models predict short-horizon visual futures and actions well, but they are not built to organize behavior across a sequence of subtasks or to accept an explicit task-level instruction about what state the robot should reach. ViGAR addresses this by making a subgoal image an explicit decision variable: a subgoal planner first predicts the visual target state for the next subtask, and a subgoal-guided world-action model then jointly generates future frames and actions toward it. The same framework also lets a user-supplied global goal image steer the robot toward task compositions never seen in training, without any parameter update.
Key Contributions
- Factorization of compositional manipulation into "long-term future selection" and "short-term future realization," implemented as ViGAR — a hierarchical framework that predicts the next subgoal as an explicit decision variable and conditions joint video-action generation on it.
- In-context robotic manipulation as a capability that emerges from explicit goal conditioning: with scene and instruction held fixed, a new global goal image elicits a different subtask decomposition and a correspondingly reorganized action sequence at deployment, with no parameter update.
- Strong empirical results on the RoboTwin Clean2Random benchmark (82.00% Clean, 67.02% Random, 74.51% average) and on five real-robot compositional tasks, outperforming VLA and WAM baselines.
- Implementation designs for subgoal supervision — a subgoal look-ahead rule for temporal boundary handling and end-effector region-of-interest (ROI) weighting for spatial supervision — both validated by ablation.
Main Findings
- Subgoal prediction quality: On RoboTwin, ViGAR's planner 𝒢_θ achieves LPIPS 0.13, DINO-cos 0.80, ΔIoU 0.65, and RetAcc 0.94, versus VISTA (0.38, 0.54, 0.22, 0.82) and RxBrain (0.73, 0.19, 0.18, 0.79). The same ranking holds on real-robot data: ViGAR 0.11 / 0.81 / 0.66 / 0.76, VISTA 0.38 / 0.61 / 0.24 / 0.31, RxBrain 0.60 / 0.29 / 0.18 / 0.22.
- Simulation success rates: ViGAR reaches 82.00% (Clean), 67.02% (Random), 74.51% (Avg.) on RoboTwin Clean2Random with the Aloha AgileX dual-arm embodiment. The next-highest average among the listed baselines is 4D-WAM at 61.65%, so ViGAR's average is 12.86 percentage points higher, matching the paper's headline claim.
- Other baselines: StarVLA 46.50 / 3.20 / 24.85; Abot-M0 57.40 / 30.40 / 43.90; X-VLA 68.00 / 20.90 / 44.45; π0.5 70.70 / 46.00 / 58.35; Fast-WAM 77.80 / 1.90 / 39.85; LingBot-VA 80.70 / 34.60 / 57.65; 4D-WAM 81.50 / 41.80 / 61.65.
- Clean performance is largely saturated; the gain is in Random: The paper states Clean performance is largely saturated across WAMs and that the benefit of explicit subgoal selection materializes when scene layout can no longer be memorized from demonstrations.
- Subgoal guidance versus backbone alone: Cosmos3-Nano-RoboTwin — the same policy adaptation recipe but without subgoal guidance — reaches only 22.46% in the Random setting (77.06% Clean, 49.76% average).
- Real-robot compositional tasks: On an AgiBot A2 robot (two 7-DoF arms, two wrist cameras, one front-facing camera), ViGAR achieves the best overall performance across five long-horizon tasks measured by task progress score, outperforming π0.5 and Cosmos3-Nano-Policy.
- In-context steering: On two in-context tasks (Fruit Arrangement, Desktop Item Storage), ViGAR achieves average task progress of 77.5% in-domain and 47.5% out-of-distribution, each over 20 rollouts.
- Goal injection route ablation: Using ground-truth goal images and 30k training steps on the same 2500 RoboTwin Clean trajectories — None 76.58 / 21.78 / 49.18; Reasoner 78.42 / 23.52 / 50.97; Generator 78.92 / 45.42 / 62.17; Generator+Reasoner 82.16 / 41.28 / 61.72. Generator-only injection gives the best average (62.17%); adding the reasoner route improves Clean but reduces robustness under Random, motivating the geometric route used in ViGAR.
- Subgoal supervision ablation: Full ViGAR 82.00 / 67.02 / 74.51; without subgoal look-ahead 79.80 / 65.92 / 72.86; without EEF ROI weighting 79.28 / 65.64 / 72.46; without both 72.68 / 52.40 / 62.54. EEF ROI weighting contributes the most to the gain; look-ahead prevents regressing to already-completed subtasks.
Methodology in Plain English
ViGAR has two parts, both initialized from the same pretrained Cosmos3-Nano world model (a 16B-A8B-parameter omnimodal world model with a dual-branch mixture-of-transformers architecture: a reasoner branch over language and semantic visual tokens and a generator branch that produces images or video by flow matching, interacting through a shared multimodal attention layer).
1. Subgoal planner. Given the current observation and the global instruction, the planner predicts the image of the target state for the next subtask, framed as an image-editing problem (the backbone's image-to-video capability is adapted to image editing). The global instruction is kept unchanged for every subtask rather than being decomposed into subtask text, because the subgoal image already grounds the target state and this avoids error accumulation from text generation. Two supervision tricks are used: a subgoal look-ahead rule, which redirects observations falling in the last p% of a subtask to the next subtask's goal image so visually similar frames near boundaries get consistent targets; and an end-effector ROI mask, built by projecting expert end-effector positions through calibrated cameras, which up-weights latent tokens near where manipulation happens. The planner is trained with a weighted flow-matching loss.
2. Subgoal-guided world-action model. Conditioned on the observation, proprioceptive state, instruction, and the predicted subgoal, the policy jointly predicts a short visual future and an action chunk with a joint flow-matching objective. The subgoal is injected through the geometric route — VAE-encoded into a clean latent fed to the generator branch, tagged with a zero-initialized learnable goal-role embedding — rather than as vision-encoder tokens to the reasoner, because this grounds the goal at the pixel level. Action horizon and video horizon are both 48 with a video downsample rate of 4, so each control query yields 48 actions and 12 video frames. To reduce train–test mismatch, the trained planner is run on the policy's own training demonstrations, and any semantically inconsistent or visually distorted predictions are replaced with ground-truth subgoals before the policy is trained on this mixture.
3. In-context specification. The inference-time context is a global goal image — the terminal world state the task should produce. Training uses the terminal frame of each demonstration trajectory as this context. The context enters only the subgoal planner; the downstream policy still conditions only on the predicted subgoal, so a global goal unseen in training can induce a never-demonstrated subtask decomposition with θ and φ unchanged.
Data and training. The subgoal planner has two variants (Base and ROI) trained for 100k steps on a large robot manipulation corpus spanning single-arm and bimanual embodiments; ROI uses the standard objective for the first 30k steps then 70k steps with EEF ROI weighting at a 4:1 foreground-to-background weight ratio. Both use AdamW with learning rate 2×10⁻⁵ (50-step warm-up, then constant) and global batch size 32; the generator branch and its visual projections are trainable while the reasoner branch and visual tokenizer are frozen. The policy is post-trained per the Cosmos3 robot-policy recipe with action encoder, action decoder, and action-modality embedding introduced and zero-initialized. Simulation training uses 2,500 Clean expert RoboTwin demonstrations, with 19 tasks segmented as long-horizon and 31 treated as single-stage. Real-robot data comprises roughly 900 hours of teleoperation on AgiBot A2 covering 752 tasks, more than 200 distinct objects, and over 10 backgrounds, with human-in-the-loop subtask annotation.
Why This Matters
Impact on research. The paper argues that short-horizon world-action prediction is not sufficient for long-horizon composition, and offers an explicit task-level interface — a visual subgoal — that links planning to physical action generation while keeping both stages inside one shared world-model representation. It also reframes in-context robot learning: instead of demonstrating how a task is done (as with demonstration-video context in prior work), the context specifies what the scene should become, leaving the decomposition to the model.
Real-world applications (from the paper's own tasks):
- Warehouse and household tidying, such as the Tidy-up Desktop task that collects scattered trash and objects into a bin (four subtasks).
- Food service and handling, such as Take-out Steam Buns, retrieving buns from each tier of a two-tier bamboo steamer onto a plate (five subtasks).
- Sorting and organizing, such as Stack Colored Bowls, grouping bowls of the same color into separate stacks (three subtasks).
- Reconfigurable pick-and-place steered by a goal image, as in Fruit Arrangement (distributing fruits across two trays) and Desktop Item Storage (placing specified item types into one box).
Industry relevance. Deployment environments need policies that hold progress across subtask transitions and adapt to unseen objects and layouts instead of replaying a fixed sequence. A goal-image interface means a task composition can be re-specified at inference time without data collection or fine-tuning, which is directly relevant to operators who want to redirect a deployed robot cheaply. The results emphasizing the Random setting also matter commercially: gains appear exactly when layout cannot be memorized from demonstrations.
Future Directions
- Closing the in-context generalization gap: Task progress drops from 77.5% in-domain to 47.5% out-of-distribution on the two in-context tasks, leaving substantial room for improvement in following unseen global goals.
- Broadening the goal-injection study: The route ablation is run with ground-truth goal images for only 30k steps on the 2500 RoboTwin Clean trajectories; extending it to predicted subgoals and to real-robot settings is an open step.
- Extending beyond the evaluated embodiments and tasks: Evaluation covers RoboTwin with the Aloha AgileX dual-arm embodiment plus the AgiBot A2 with two 7-DoF arms; behavior on other embodiments, camera setups, and control modes is not established here.
- Reducing reliance on subtask annotation: Both the simulation (19 segmented tasks) and real-robot (human-in-the-loop segmentation) pipelines depend on annotated subtask structure, which raises the question of how the framework scales where such segmentation is unavailable.
Target Audience
Robotics and embodied-AI researchers working on manipulation policies, world models, and VLA systems; engineers building long-horizon robot deployments who need task-level control interfaces; and readers interested in in-context learning or goal-conditioned generation applied to physical systems. The paper assumes familiarity with flow matching, VAE latent spaces, and standard manipulation benchmarks, so it is best suited to readers with a graduate-level or practitioner background in robot learning.
Authors’ abstract
Long-horizon compositional manipulation has become increasingly important for real-world robot deployment, where a single task involves multiple coordinated subtasks. Existing world-action models (WAMs) jointly predict short-horizon visual futures and actions, but typically lack explicit subtask-level reasoning. We propose Visual Goal-conditioned Action Reasoning (ViGAR), a hierarchical framework that factorizes manipulation into a visual subgoal planner and a subgoal executor. Given the current observation and global instruction, the subgoal planner predicts a visual subgoal for the next subtask. The subgoal executor then jointly generates future visual trajectories and actions conditioned on the predicted subgoal. Both components share a pretrained world-model representation, enabling task-level planning and action generation to benefit from common physical knowledge. Moreover, our framework naturally supports in-context learning: using a global goal image as context can induce different subtask decompositions and behaviors without parameter updates. On the RoboTwin Clean2Random benchmark, ViGAR achieves 82.00% and 67.02% success rates under the Clean and Random settings, respectively, surpassing the strongest baseline by 12.86 percentage points in average success rate. Real-world robot experiments on five compositional and two in-context learning tasks further confirm the effectiveness of ViGAR.