Research
UniWAM: Unified World-Action Model
UniWAM: Unified World-Action Model Overview Research area: Robotics — robot foundation models, specifically the combination of vision-language-action models (VLAs) and world-action models (WAMs) for m

- arXiv
- 2610.02054
- Published
- 2026-10-01
- Authors
- Jiayi Chen, Wenxuan Song, Jingbo Wang, Shuai Zhou, Xicheng Gong, Zehua Fan, Ziyang Zhou, Junwu E, Haodong Yan, Fuhao Li, Qize Yu, Xu Huang, Pengwei Wang, Wen Chen, Shunbo Zhou, Haoang Li
AI summary
UniWAM: Unified World-Action ModelOverview
Research area: Robotics — robot foundation models, specifically the combination of vision-language-action models (VLAs) and world-action models (WAMs) for manipulation.
Technical level: Advanced. The paper assumes familiarity with vision-language models, video generation models, flow matching, and Mixture-of-Transformers architectures.
Scope (one sentence): UniWAM is a single Mixture-of-Transformers architecture that unifies a vision-language physical reasoner, a video-generation world generator, and a flow-matching action predictor, pre-trained on robot data, human egocentric video, and visual question answering data, and evaluated on LIBERO, RoboTwin 2.0, and LIBERO-Plus.
What This Paper Is About
Vision-language-action models inherit strong semantic understanding and reasoning from pretrained VLMs, but supervising them with actions alone gives them limited grounding in how the physical world actually changes. World-action models, built on video generation backbones, inherit spatiotemporal priors from large-scale video pretraining, but they remain weaker at semantic understanding and reasoning, especially under distribution shifts. UniWAM asks how to combine both strengths — the reasoning and generalization of VLMs with the dynamics understanding of video generation models — inside one architecture trained on a mixture of human and robot data.
Key Contributions
-
A unified architecture: A Mixture-of-Transformers (MoT) that connects a physical reasoner (built on Qwen3-VL-2B-Instruct), a world generator (built on the Wan2.2-TI2V-5B backbone), and an action predictor through joint attention, allowing tokens of understanding, video, and action streams to attend to each other at every layer.
-
A three-source pretraining scheme with complementary supervision: VQA data (8.293M QA pairs) supervises the physical reasoner to preserve pretrained knowledge; human egocentric data supervises both the physical reasoner and the world generator while avoiding low-precision action labels; robot data supervises all three experts because it has the highest physical consistency with downstream post-training and the most accurate action annotations. The paper reports source-level ablation studies evaluating each data source.
-
Post-training designs for efficiency and robustness: Introduction of future-frame noise augmentation (which perturbs only future visual latents, with probability 0.5 per sample, with the augmentation scale drawn from U(0.5, 1)) and history-conditioned flow matching (which initializes action generation from a perturbed chunk of previously executed actions with η = 0.02 rather than from Gaussian noise), together reported to substantially reduce denoising steps while maintaining performance.
-
A rigorous data pipeline: Quality control for robot and human data, including temporal/statistical trajectory screening, geometry-aware consistency checks, and visual validity filtering, plus EgoANT, a VLM-based pipeline that automatically segments untrimmed egocentric video into atomic manipulation subtasks and annotates them with language descriptions. EgoANT uses Qwen3.6-27B for global segmentation and local refinement and Qwen3.5-397B for segment annotation and text-only candidate selection.
Main Findings
-
In-distribution LIBERO performance: UniWAM reaches 99.2% average success across the LIBERO suites (Spatial 99.6, Object 99.6, Goal 99.2, Long 98.4), which the paper states outperforms the previous best listed baseline, Xiaomi-Robotics-0 at 98.7%, by 0.5%.
-
RoboTwin 2.0 results: On the 50-task RoboTwin 2.0 evaluation, UniWAM scores 75.14% on Clean2Clean (C2C) and 68.32% on Clean2Random (C2R), for an average of 71.73%, the highest average among the listed methods. It reports the highest C2R value among listed methods.
-
Much smaller degradation under domain randomization: UniWAM's success rate drops by 6.82 percentage points from C2C to C2R, compared with 24.70 percentage points for π0.5 and 50.46 percentage points for Spatial Forcing.
-
Robustness on LIBERO-Plus: UniWAM achieves 92.6% overall success on LIBERO-Plus, the highest reported among the listed methods, and reports the highest listed values under robot (89.5%), language (92.2%), light (97.9%), and background (97.6%) perturbations across seven perturbation dimensions (camera, robot, language, light, background, noise, layout).
-
A scaling law for human-robot co-training: The paper reports uncovering a log-linear scaling law of unified human-robot co-training, presented as evidence that large-scale pretraining on a mixture of human and robot data is effective. The truncated content does not report the fitted parameters or the scaling exponent.
-
Physical language as supervision: Low-level actions are represented in natural language (following LAP), so that action supervision matches the VLM's pretraining input-output distribution. The paper attributes improved language robustness (92.2% under language perturbation on LIBERO-Plus) partly to this design, and poses the question of how physical language shapes learned representations as an explicit research question.
-
Data scale: Table 1 reports roughly 4,958 hours of robot data, roughly 5,072 hours of human egocentric data, and roughly 10,013 hours in total. Human sources are EgoVerse (4,003 hours), EgoDex (829 hours), and VITRA (240 hours); robot sources include AgiBot World 2026 (891), AgiBot World Alpha (595), Bridge (80), Droid (365), Fractal (340), Robocoin (1,088), and Interdata-A1 (1,600). Section 3.1.1 separately states that the selected robot data total approximately 4,363 hours.
-
Real-world results: The paper poses questions about instruction following and long-horizon manipulation in the real world and includes a Section 5.2 and Figure 5, but the supplied content is truncated at that point, so no real-world numbers are reported here.
Methodology in Plain English
The model is a single transformer stack with three specialized "experts" that share attention. The understanding expert takes the current image and the task instruction and produces language; the video expert predicts future visual latents in the frozen latent space of a video autoencoder; the action expert produces continuous actions via flow matching. Because the experts attend to each other's tokens at every layer, action prediction can draw on both the language semantics and the evolving visual prediction, while each expert keeps its own normalization, projections, and feed-forward network.
Training has two phases. In pretraining, the team mixes three data sources and assigns each a different kind of supervision. VQA questions train the understanding expert to retain its pretrained knowledge. Human egocentric video trains the understanding and video experts, since human video has rich visual and motion content but imprecise action labels. Robot demonstrations train all three experts because their actions are accurate. A key trick is to write actions out as natural-language sentences describing end-effector motion, so that teaching the VLM about actions looks like the text tasks it was pretrained on; the loss is computed only on answer tokens.
In post-training, the model is adapted to specific embodied setups. Two changes matter. First, future visual latents are partially corrupted with noise so the action expert cannot simply lean on a near-perfect future prediction — it must pull control-relevant information out of coarse visual representations. Second, instead of starting action generation from pure Gaussian noise, the flow-matching process starts from the recent action history perturbed by small noise, which gives the model a structured prior and lets it reach a good action chunk in fewer denoising steps.
Data quality is handled separately: robot trajectories are screened for abrupt transitions, state-action temporal consistency, and distribution outliers; joint and end-effector measurements are cross-checked by forward kinematics where metadata allow; and invalid visual frames are removed. Human video is segmented into short manipulation subtasks and described in language by the EgoANT pipeline, with wrist poses transformed into the head-camera frame and concatenated into a 14-dimensional motion representation.
Why This Matters
Impact on research. The paper argues that action-only supervision and video-only supervision are each insufficient, and offers a concrete recipe for combining a pretrained VLM, a pretrained video generator, and an action head under one attention scheme with data-specific supervision. It also presents human-robot co-training as a scalable axis, reporting a log-linear scaling law, which addresses the practical problem that robot demonstrations are far scarcer than human video. The data-cleaning and automatic annotation pipelines are contributions in their own right for anyone assembling heterogeneous robot and egocentric datasets.
Real-world applications:
- General-purpose manipulation in homes or laboratories, where lighting, backgrounds, object layouts, and initial robot poses vary from the training setup.
- Instruction-following assistants that must respond to rephrased natural-language commands, since language perturbation was one of the dimensions tested.
- Dual-arm and mobile robot platforms (the pretraining mixture covers single-arm, dual-arm, and mobile embodiments).
- Long-horizon task execution, where the model must carry a plan across many manipulation subtasks.
Industry relevance. The model builds on existing open backbones (Qwen3-VL-2B-Instruct and Wan2.2-TI2V-5B), which lowers the barrier to reproducing or adapting the approach. The emphasis on reducing denoising steps through history-conditioned flow matching targets inference cost, which is a practical constraint for deployed robots. The paper lists public code, checkpoints, and a project page.
Future Directions
-
Report and characterize the log-linear scaling law in more detail — fitted parameters, exponents, and the point at which human data stops substituting for robot data.
-
Quantify real-world instruction following and long-horizon manipulation, which the paper sets up as research questions but whose content is not included in the supplied material.
-
Test how far the physical-language action representation extends across embodiments, given that it is the mechanism the paper credits for language robustness and shared motion semantics.
-
Determine whether the future-frame noise augmentation and history-conditioned flow matching trade off against anything else, and how few denoising steps the model can use before performance degrades — the paper states denoising steps are significantly reduced but the truncated content does not give the step counts.
-
Extend the EgoANT annotation pipeline (and its subtask-level labels) to more egocentric sources, and evaluate how annotation quality propagates into policy performance.
Target Audience
Robotics and embodied-AI researchers working on robot foundation models, vision-language-action models, and world models; engineers building manipulation policies who want to know which pretraining data mixtures and post-training tricks pay off; and dataset builders interested in the quality-control and automatic ego-centric annotation pipelines. Readers need working knowledge of flow matching, video diffusion backbones, and vision-language model training to follow the method section in detail; the experimental tables are readable without it.
Authors’ abstract
Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. To ensure the quality of the training data, we developed a rigorous data cleaning and annotation pipeline for both human egocentric data and robot data. To adapt the vision-language component to embodied tasks while preserving its inherited language capabilities, we represent low-level actions in natural language and introduce a pre-training recipe that assigns complementary supervision from visual question answering (VQA) data, human egocentric data, and robot demonstrations to the appropriate model components. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while history-conditioned flow matching uses encoded action history to initialize action generation. Together, these designs significantly reduce denoising steps while maintaining performance. UniWAM achieves state-of-the-art (SOTA) performance across multiple evaluations, including in-distribution performance, robustness, generalization, instruction following, and long-horizon task execution. Furthermore, we uncover a log-linear scaling law of unified human-robot co-training, demonstrating the effectiveness of large-scale pre-training on a mixture of human and robot data.