Research
EVO-WAM: Evolving World Action Models through Video-Action Verification
Overview Research area: Robot learning / computer vision — specifically world action models (WAMs) that jointly predict future videos and robot actions, and their adaptation to tasks not seen during b

- arXiv
- 2609.38057
- Published
- 2026-09-29
- Authors
- Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang, Songsong Yu, Yijun Yang, Jianhui Liu, Haoze Sun, Senqiao Yang, Li Jiang, Jingyong Su, Haoyang Huang, Zhuotao Tian
AI summary
Overview
- Research area: Robot learning / computer vision — specifically world action models (WAMs) that jointly predict future videos and robot actions, and their adaptation to tasks not seen during base-model training.
- Technical level: Advanced. The paper assumes familiarity with vision-language-action (VLA) policies, world models, inverse dynamics models (IDMs), autoregressive video generation, and supervised fine-tuning of large pretrained backbones.
- Scope: The paper introduces EVO-WAM, a generate-verify-improve self-training framework that lets a WAM improve on unseen tasks using only its own generated video-action rollouts, validated by a vision-language model and an inverse dynamics model, with no additional expert demonstrations and no execution of candidate actions in an external environment during self-evolution.
What This Paper Is About
World action models can produce videos and actions for tasks they were never trained on, but those generations are unreliable: a video might not depict the task being completed, and even a visually successful video may be paired with actions that do not actually realize the depicted behavior. Collecting new expert demonstrations to fix this is expensive, so the authors ask whether a WAM can instead improve by mining reliable trajectories out of its own generated rollouts. EVO-WAM answers this by generating complete autoregressive rollouts, filtering them through two verification stages, and iteratively fine-tuning the model on the survivors.
Key Contributions
- A framework for improving WAMs with generated experience. EVO-WAM enables WAMs to generate autoregressive video-action rollouts and improve on unseen tasks through verification and iterative self-training, without additional expert demonstrations or action execution in an external environment during self-evolution.
- Verification of task completion and video-action consistency. A two-stage verification process in which a vision-language model identifies task-completing prefixes and an inverse dynamics model assesses their video-action consistency, selecting generated rollouts for iterative self-training.
- Generality across WAM backbones, tasks and VLMs. Improvements are demonstrated across Cosmos3 and DreamZero on unseen RoboTwin 2.0 tasks, and with Cosmos3 on real-world long-horizon composite tasks, with robustness to verifier choice (Qwen3.8-Flash-Next and the smaller Qwen3.5-27B).
- Two enabling modifications for rollout generation. The WAM is augmented with state prediction (predicting the robot state at the end of each chunk) and anchored multi-frame context so complete rollouts can be produced without external execution feedback.
Main Findings
- Simulation gains on unseen tasks: On seven RoboTwin 2.0 tasks unseen during base-model training, EVO-WAM raises average success from 26.9% to 68.0% for Cosmos3 (approximately 2.5 times the initial rate) and from 28.5% to 46.4% for DreamZero (approximately 1.6 times).
- Comparison to baselines: The strongest baseline in the simulation table is Cosmos3 at 31.6% average; DreamZero reaches 27.2%, π0.5 16.9%, LingBot-VLA 11.9%, LingBot-VA 10.4%, Fast-WAM 8.4%, and StarVLA-OFT 4.4%. EVO-WAM DreamZero reaches 46.4% and EVO-WAM Cosmos3 68.0%.
- Real-world gains: On three unseen long-horizon composite tasks with a Franka robot, EVO-WAM Cosmos3 reaches 76.7% average success versus 20.0% for both Cosmos3 and DreamZero and 6.7% for π0.5 — a gain of 56.7 percentage points over the Round 0 starting point. Per task it scores 80.0% (stack bowls), 60.0% (place ducks) and 90.0% (load the air fryer), exceeding the strongest baseline by 20, 50 and 80 percentage points respectively.
- Most improvement comes early: EVO-WAM Cosmos3 gains most in Round 1 (26.9% to 58.3%). After four rounds, EVO-WAM Cosmos3 and EVO-WAM DreamZero reach 68.0% and 46.4%. The Cosmos3 variant temporarily declines in Round 3 (63.6%) and recovers to 68.0% in Round 4. On the real robot, EVO-WAM Cosmos3 reaches 76.7% in Rounds 2 and 4, with 73.3% in Round 3.
- Visual completion alone is not enough: In the ablation, VLM-only verification reaches 43.7% in Round 4, while VLM + IDM reaches 68.0%. A simulator-based verifier (VLM + Simulator), which directly tests whether actions complete the task, reaches 72.7% but requires a simulator of the target task and scene; IDM-based verification also works on real-world tasks without one.
- Robustness to verifier strength: Using the smaller Qwen3.5-27B for task-completion assessment still yields 65.7% in Round 4 versus 68.0% with Qwen3.8-Flash-Next.
- Generalization to new scenes: On 100 newly sampled scenes per task, EVO-WAM Cosmos3 reaches 70.4% success versus 24.9% for the 34K Cosmos3 baseline.
- Seen-task retention: On the 43 seen tasks with 50 Clean and 50 Randomized trials each, EVO-WAM Cosmos3 at Round 4 achieves 84.8% versus 85.8% for the 34K Cosmos3 baseline — a small decrease.
- Accumulating data helps: Latest-round-only training declines after Round 1 (26.9, 58.6, 60.5, 57.1, 53.6 across Rounds 0–4), while accumulating verified prefixes across rounds yields 26.9, 58.3, 66.6, 63.6, 68.0, measured over 1,400 trials per round.
- Concrete rejection examples: In one bowl-stacking case, a rollout had a prefix IDM score of 0.008170 against a threshold of 0.008111 and was rejected; in an air-fryer case, a rollout scored 0.009082 against the same threshold and was rejected.
Methodology in Plain English
The framework runs a loop over rounds, starting from a base WAM checkpoint.
- Generate. The WAM produces video-action chunks one after another. To do this without a robot in the loop, the authors trained the model to also predict the robot state at the end of each chunk, supplying the configuration needed to continue. For visual context, the rollout starts from a single frame that is kept as a persistent "anchor" alongside the most recent generated frames — the recent frames carry motion history, the anchor keeps the objects and scene stable. Repeating this for a task-specific budget of chunks yields a full candidate trajectory.
- Verify. A vision-language model first splits its assessment into a visual description (objects, spatial relations, robot configuration from multiple camera views) and a task judgment checked against the instruction and active subgoal. Four criteria must all be satisfied: goal satisfaction, required gripper release, object consistency, and robot structural consistency. Predefined time points are scanned chronologically and subgoals are checked in sequence; each endpoint must receive at least two Accept judgments out of three, and a missing endpoint or failed check discards the candidate. Separately, an inverse dynamics model — trained on recorded video-action pairs, including unsuccessful executions — reconstructs actions from the generated video and compares them with the WAM's paired actions in the same normalized action space. The mean squared error across verification windows must fall at or below a threshold that is fixed per backbone and dataset.
- Improve. Prefixes passing both stages are combined with all retained data from earlier rounds and mixed with the original training data, then the WAM is updated by supervised fine-tuning. The updated model generates candidates for the next round.
Simulation setup: 43 RoboTwin 2.0 tasks for base-model training and seven held out for self-training and evaluation, 100 Clean and 100 Randomized trials per task. Four self-training rounds from 30K-step checkpoints, each generating 2,800 candidate rollouts followed by 1K training updates at global batch size 256; baselines were trained for 34K steps. Real-world setup: three unseen long-horizon composite tasks on a Franka robot, four rounds using an initial batch of 800 candidate rollouts per round and 500 training updates, ten trials per task per policy.
Why This Matters
This work shows that a world action model's own predictions can serve as a training signal for tasks it cannot yet perform — turning unreliable imagination into usable supervision through verification rather than through more human-collected data. It also isolates what makes generated experience useful: visual task completion and video-action consistency must both be checked, since visually successful videos can carry actions that fail on execution.
Real-world applications:
- Household and service robotics: long-horizon composite chores such as stacking bowls, sorting objects into matching containers, and operating appliances (opening a drawer before placing bread inside), the exact tasks evaluated here.
- Warehouse and logistics manipulation: picking and placing objects into boxes or onto scales, including placement relative to other objects ("place A to the left of B").
- Rapid task onboarding in factories: adding new manipulation tasks to a deployed policy without collecting new demonstrations or risking hardware during adaptation, since candidate actions are never executed in the external environment during self-evolution.
- Simulation-to-real pipelines: the verifier design transfers to physical robots without requiring a simulator of the target task and scene, where a simulator-based verifier would.
Industry relevance: the approach targets the main cost driver in robot deployment — collecting expert demonstrations for every new task — and works across two different WAM backbones and two different vision-language models, which suggests the pattern can be applied to other world-model architectures rather than being tied to one system.
Future Directions
- What to do when generated candidates lack the needed behavior: the authors note that DreamZero's limited improvement on empty-cup placement and block stacking may reflect constraints on the useful behaviors available in its generated candidates, implying the ceiling depends on the base model's imagination.
- Why later rounds regress: the Cosmos3 simulation variant drops in Round 3 and real-world success dips in Round 3, showing that additional self-training does not always improve performance; the cause of these occasional regressions is not reported.
- Closing the gap to execution-verified supervision: the VLM + Simulator criterion reached 72.7% in Round 4 versus 68.0% for VLM + IDM, but requires a simulator; finding simulator-free verification that matches it is an open problem.
- Scaling the verification and data budget: the paper reports specific round counts, rollout budgets and update counts for RoboTwin and the real robot but does not report what happens with substantially larger generation budgets or more rounds.
Target Audience
Robotics and embodied-AI researchers working on policy learning, world models, VLA systems, and self-improvement from generated data. It is also relevant to practitioners who deploy manipulation policies and need to add new tasks without new demonstrations, and to readers interested in how verification (rather than reward modeling) can gate self-training. Because it combines video generation, inverse dynamics, and vision-language judgment, it suits readers comfortable with modern generative video models and robot learning pipelines; newcomers will find the high-level framing accessible but the method details demanding.
Authors’ abstract
Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generated videos may fail to depict task completion, and even visually successful videos may be paired with inconsistent actions that lead to execution failure. We propose EVO-WAM, a framework that adapts WAMs to unseen tasks by learning from their own generated video-action trajectories, without executing candidate actions in an external environment. First, we augment WAM training with state prediction and anchored multi-frame context to enable complete autoregressive rollouts without external execution feedback. Second, we identify reliable training experience by selecting task-completing prefixes with a vision-language model and verifying their video-action consistency with an inverse dynamics model. Third, we iteratively train the WAM on verified prefixes and generate new rollouts with the updated model. On seven unseen RoboTwin 2.0 tasks, EVO-WAM increases average success rates from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, reaching approximately $2.5\times$ and $1.6\times$ their initial success rates. On three unseen long-horizon composite tasks in the real world, it improves Cosmos3's average success rate from 20.0% to 76.7%, a gain of 56.7 percentage points. Project Page: https://evo-wam.github.io/.