Skip to content
AI.info

Research

Rethinking Video Generation Model for the Embodied World

Overview Research area: Computer vision — video generation models applied to embodied AI and robotics, spanning benchmark design, dataset construction, and evaluation methodology. Technical level: Adv

arXiv
2601.15282
Published
2026-01-21
Authors
Yufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li, Ruoqing Hu, Yufei Ding, Yiming Zou, Yan Zeng, Daquan Zhou

AI summary

Overview

  • Research area: Computer vision — video generation models applied to embodied AI and robotics, spanning benchmark design, dataset construction, and evaluation methodology.
  • Technical level: Advanced
  • Scope: Introduces RBench, a 650-sample benchmark for robot-oriented video generation with automated metrics, plus RoVid-X, a 4-million-clip robotic video dataset built to train such models.

What This Paper Is About

Video generation models are increasingly used to synthesize robot data for embodied AI, but there has been no standardized way to judge whether the videos they produce actually show physically plausible, task-correct robotic behavior. The authors argue that existing evaluations lean on perceptual quality metrics and therefore give overly optimistic scores to videos with unnatural motion, object penetration, or missing actions. To fix this, they build both a benchmark that grades task-level correctness and visual fidelity, and a large-scale training dataset to help models close the gap the benchmark exposes.

Key Contributions

  1. RBench, a benchmark for robotic video generation. It contains 650 image–text pairs: 250 task-oriented pairs across five tasks (Common Manipulation, Long-horizon Planning, Multi-entity Collaboration, Spatial Relationship, Visual Reasoning — 50 samples each) and 400 embodiment-specific pairs across four robot types (Dual-arm, Humanoid, Single-arm, Quadruped — 100 samples each), evaluated with reproducible automated sub-metrics.
  2. A systematic evaluation of 25 video models. The study covers open-source, commercial, and robotics-specific models, revealing where current video foundation models fall short on physically realistic robot behavior.
  3. RoVid-X, a large-scale robotic video dataset. Built through a four-stage pipeline, it contains 4 million annotated video clips with standardized task descriptions and physical property annotations, drawn from internet video platforms and more than 20 open-source embodied video datasets.
  4. Validation of both artifacts. A human preference study with 30 participants shows a Spearman rank correlation of ρ = 0.96 (two-sided p < 10⁻³) between human scores and RBench scores, and finetuning experiments on RoVid-X show consistent gains across tasks and embodiments.

Main Findings

  • Automated metrics track human judgment closely. On the ten-model subset used in the human preference study, RBench scores correlate with human scores at ρ = 0.96 (two-sided p < 10⁻³). Participants were shown two model outputs side by side for the same prompt and chose "A is better," "B is better," or "Tie"; a win contributed 5, a loss 1, and a tie 3 to both models.
  • Commercial closed-source models dominate. The top seven leaderboard positions are all commercial models: Wan 2.6 (Rank 1, average 0.607), Seedance 1.5 pro (Rank 2, 0.584), Wan 2.5 (Rank 3, 0.570), Hailuo v2 (Rank 4, 0.565), Veo 3 (Rank 5, 0.563), Seedance 1.0 (Rank 6, 0.551), and Kling 2.6 pro (Rank 7, 0.534). The best open-source model is Wan2.2_A14B at Rank 8 with 0.507.
  • Iterative model scaling produces large physical-reasoning jumps. The Wan series goes from Wan 2.1 (Rank 14, 0.399) to Wan 2.6 (Rank 1, 0.607), and Seedance moves from 1.0 (Rank 6) to 1.5 Pro (Rank 2).
  • Consumer-oriented models underperform. Sora v2 Pro lands at Rank 17 with 0.362 and Sora v1 at Rank 22 with 0.266, which the authors attribute to a "domain gap" — models tuned for cinematic smoothness sacrifice physical fidelity and precise motion control. The paper notes that approximately 50 of the 650 videos could not be generated for Sora v2 Pro due to review limitations in the official Sora API.
  • Robotics-specific models show a trade-off. Cosmos 2.5 reaches Rank 9 with 0.464, outperforming much larger open-source general models. Models fine-tuned on narrow robot entities rank at the bottom: Vidar at Rank 24 (0.206) and UnifoLM-WMA-0 at Rank 25 (0.123). The authors frame this as a trade-off between domain-specific control precision and general "world knowledge" from large-scale pretraining.
  • Cognitive and fine-grained control are the bottlenecks. Even the top model, Wan 2.6, drops in Visual Reasoning (0.531). Models consistently score higher on coarse locomotion embodiments than on fine-grained manipulation: for example, Wan 2.6 scores 0.723 on Quadruped and 0.667 on Humanoid versus 0.666 on Single arm and 0.681 on Dual arm.
  • RoVid-X finetuning improves performance. With MSE loss and a random 200k-instance sample of RoVid-X, Wan2.1_14B improves from 0.399 to 0.446 overall, and Wan2.2_5B from 0.380 to 0.439. Task-level gains for Wan2.1_14B include Manipulation 0.344 to 0.376, Long-horizon 0.335 to 0.389, Multi-entity 0.282 to 0.295, Spatial 0.268 to 0.314, and Reasoning 0.205 to 0.298.
  • Characteristic failure modes are identified. Figure 2 highlights robot shape distortion, object attribute drift, and non-contact attachment, alongside floating or penetrating parts, spontaneous emergence of entities, and omitted key actions.
  • The benchmark avoids test contamination. Videos used in the evaluation set are excluded from the training database, and new task prompts are written for each reference image, with all samples verified by human annotators.
  • Qualitative differences across models. In the visual reasoning case, Seedance 1.0 and Hailuo correctly identify the blue clothing and hollow basket while Wan 2.5 mistakes the woven basket for the hollow one; in long-horizon planning, Wan 2.5 completes all actions in order while Hailuo omits the "turn-on" action; in spatial relationship, Hailuo places the bok choy to the left of the pan while other models put it inside the pan, and LongCat-Video introduces an unrealistic human arm intervention.

Methodology in Plain English

The authors approach the problem from both sides — measurement and data. First, they assemble an evaluation set of 650 robot image–text pairs by extracting keyframes from high-quality public datasets and online videos, then manually verifying each image and writing a fresh task prompt for it. Cases are split by task type and by robot embodiment so results can be broken down meaningfully. Each sample carries metadata such as the manipulated object, embodiment type, and camera viewpoint.

To grade generated videos without relying on humans for every run, they use multimodal large language models — the open-source Qwen3-VL and the closed-source GPT-5 — as zero-shot evaluators. Task completion is graded along two axes. Physical-Semantic Plausibility uses a VQA-style protocol over uniformly sampled frames from a temporal grid to catch failures like floating or penetrating parts, entities that appear or vanish without cause, and objects that move with the robot without contact or with a badly closed gripper. Task-Adherence Consistency checks whether the video follows the prompt's intent and action order, flagging missing actions, wrong ordering, semantic drift, and non-responsiveness.

Visual quality is handled with three quantitative metrics. Motion Amplitude measures how much the robot subject actually moves while discounting apparent motion from camera movement; the authors follow VMBench, localizing active subjects with GroundingDINO, producing temporally stable masks with GroundedSAM, and tracking salient points with CoTracker, then averaging a capped mean displacement over frames. Robot-Subject Stability uses a contrastive VQA setup comparing a reference frame against a generated frame to detect shape drift, extra or missing manipulators, joint inversion, and object attribute drift. Motion Smoothness measures frame-to-frame quality stability using the Q-Align aesthetic score, flagging temporal anomalies when the change in score exceeds an adaptive threshold tied to the subject's motion.

For the dataset, they run a four-stage pipeline. They collect raw video from internet platforms and more than 20 open-source embodied datasets, using GPT-5 to filter for robot-relevant content, yielding roughly 3 million raw clips. They then apply scene segmentation to drop non-robot footage and score remaining clips on clarity, dynamic effects, aesthetics, and OCR. A video understanding model segments videos into task segments with timestamps and generates standardized captions naming the action subject, the manipulated object, and the operation. Finally, FlashVSR improves resolution, AllTracker annotates a unified optical flow, and Video Depth Anything generates relative depth maps.

For benchmarking, open-source models run with official default configurations while closed-source models are called through official APIs. Three videos are generated per image–text pair and averaged into the final sample score.

Why This Matters

Impact on research. The paper reframes how robot-oriented video generation should be judged, shifting emphasis from pixel-level aesthetics toward physical plausibility and task-level correctness. The 0.96 correlation with human rankings gives the automated metrics credibility, meaning progress in this area can be tracked reproducibly rather than through inconsistent manual review. The finding that specialized robotics models with narrow training data underperform general foundation models also raises a concrete design question for the field about how to balance domain data against broad pretraining.

Real-world applications.

  • Synthetic robot trajectory generation to supplement or replace expensive human teleoperation data collection.
  • Video world models that predict future states and help initialize or co-train robot policies.
  • Controlled simulation of manipulation and locomotion tasks for testing policies before deploying hardware.
  • Quality screening tools for teams deciding which video generators are reliable enough to be used in a robotics data pipeline.

Industry relevance. The leaderboard's top seven positions being commercial models, and the gap between Wan 2.6 (0.607) and the best open-source model Wan2.2_A14B (0.507), gives a concrete picture of where proprietary capability currently stands. The dataset comparison table places RoVid-X against RoboTurk, RoboNet, BridgeData, RH20T, DROID, Open X-Embodiment (1.4M videos, 217 skills), RoboMIND, RoboCOIN, Galaxea, InternData-A1, Fourier ActionNet, Humanoid Everyday, and Agibot World (1M videos, 87 skills), positioning 4M clips and 1300+ skills at 720P as the largest open resource of its kind with optical flow, diverse robotic forms, and diverse captions all present.

Future Directions

  • Closing the loop from video to action. The authors plan to use Inverse Dynamics Models to recover executable actions from generated videos, enabling closed-loop control experiments in both simulation and on real hardware.
  • More physically grounded metrics. They intend to develop more automated evaluation metrics that rigorously assess the kinematic and dynamic feasibility of generated behaviors, extending beyond the current plausibility and consistency checks.
  • Training for physical capability. They aim to train video generation models that generate high-fidelity actions, targeting the manipulation and visual-reasoning gaps identified in the benchmark.
  • Balancing domain data against world knowledge. The contrast between Cosmos 2.5's resilience and the weak results of entity-specific models like Vidar and UnifoLM-WMA-0 leaves open the question of how to combine proprietary robot data with generalizable representations.

Target Audience

Researchers and engineers working on video generation, robot learning, and embodied AI who need a defensible way to evaluate robot-oriented video models or a large-scale dataset to train them on. It is also relevant to teams building data pipelines for robotics, and to benchmark designers interested in using multimodal large language models as automated evaluators with validated human agreement. Readers without background in diffusion-based video generation or robot learning will find the evaluation design accessible, but the full metric definitions and mathematical details are placed in the paper's appendices.

Authors’ abstract

Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and action in the physical world. However, synthesizing high-quality videos that accurately reflect real-world robotic interactions remains challenging, and the lack of a standardized benchmark limits fair comparisons and progress. To address this gap, we introduce a comprehensive robotics benchmark, RBench, designed to evaluate robot-oriented video generation across five task domains and four distinct embodiments. It assesses both task-level correctness and visual fidelity through reproducible sub-metrics, including structural consistency, physical plausibility, and action completeness. Evaluation of 25 representative models highlights significant deficiencies in generating physically realistic robot behaviors. Furthermore, the benchmark achieves a Spearman correlation coefficient of 0.96 with human evaluations, validating its effectiveness. While RBench provides the necessary lens to identify these deficiencies, achieving physical realism requires moving beyond evaluation to address the critical shortage of high-quality training data. Driven by these insights, we introduce a refined four-stage data pipeline, resulting in RoVid-X, the largest open-source robotic dataset for video generation with 4 million annotated video clips, covering thousands of tasks and enriched with comprehensive physical property annotations. Collectively, this synergistic ecosystem of evaluation and data establishes a robust foundation for rigorous assessment and scalable training of video models, accelerating the evolution of embodied AI toward general intelligence.

Read the original paper