Skip to content
AI.info

Research

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

Overview Research area: Computer vision, specifically egocentric (first-person) video generation and embodied world-model evaluation. Technical level: Advanced. The paper builds on video diffusion mod

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
arXiv
2610.01092
Published
2026-10-01
Authors
Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana, Qinrong Cui, Erland Hilman Fuadi, Zayd M. K. Zuhri, Nanda Ryaas Absar, Ahmed Elshabrawy, Wilfried Ariel Mulyawan, Shoubin Yu, Yue Zhang, Mohit Bansal, Alham Fikri Aji

AI summary

Overview

  • Research area: Computer vision, specifically egocentric (first-person) video generation and embodied world-model evaluation.
  • Technical level: Advanced. The paper builds on video diffusion models, vision-language model judging pipelines, and inter-annotator agreement statistics.
  • One-sentence scope: The paper introduces Ego2Act, a 110-case, 2,640-video benchmark for testing whether video generation models can produce physically plausible, multi-step, goal-directed egocentric manipulation videos, together with Ego2ActJudge, a reference-free automated evaluator for those outputs.

What This Paper Is About

Video generation models are being proposed as "world simulators" that could help robots and agents plan by imagining the future. For that to work, a model must do more than produce pretty frames: given a first-person photo of a real scene and a high-level goal such as "pack the toothpaste and toothbrush into the travel case," it must invent the intermediate steps and simulate realistic hand-object physics until the goal state is reached. The paper argues that existing benchmarks either evaluate exocentric settings, score only visual fidelity, or rely on step-by-step prompt guidance, leaving multi-step goal-directed egocentric planning untested; Ego2Act fills that gap and also supplies an automatic evaluator that agrees with human judgement without needing reference videos.

Key Contributions

  1. Ego2Act benchmark. Described as the first benchmark for goal-directed, multi-step egocentric video generation, containing 110 real-world cases and 2,640 videos (660 human-collected real videos and 1,980 model-generated videos) spanning diverse domains, three clutter levels, and varying action complexity.
  2. Ego2ActJudge. A reference-free, rubric-derived automated evaluation pipeline that assesses task completion and physical plausibility in open-ended manipulation rollouts, reporting a Pearson correlation of r = 0.69 with human consensus and outperforming the tested automated baselines.
  3. Comprehensive analysis. An examination of rubric design (sequential gating versus independent gating), model failure modes in egocentric physical execution, and the reliability of the proposed automated evaluation, including backbone and prompt ablations.

Main Findings

  • Large human-AI gap. On a 25-case human-inspected panel of 600 videos, the highest-scoring model, Seedance-2.0, reaches only 64.0 overall points under human judgement. Human successful controls score 100.0 aggregate and human unsuccessful controls score 73.4, while model aggregates on the same human-judge column are 59.6 (Kling-v3-Pro), 55.3 (MiniMax-H3), 49.9 (Grok-1.5), 44.3 (Wan-2.7), and 3.9 (Cosmos-3).
  • Models skip or partially execute steps. Qualitative failure analysis shows rollouts frequently proceed to later operations, or cut straight to a completed state, without performing prerequisite actions, leaving later steps missing the dependent states they require, so the goal is unfulfilled.
  • Task success and physical correctness are only moderately coupled. The two axes show a moderate positive correlation but a wide per-sample spread, meaning one axis is not a reliable predictor of the other, and scores fluctuate considerably across the three generation seeds.
  • Contact-heavy actions are the bottleneck. At the subgoal level, attachment/connection subgoals reach only 38.6% Task completion and 66.7% Physics validity, versus 51.1% and 75.5% for relocation/arrangement, while material transfer/mixing is most often left incomplete. This pattern holds across all models.
  • Task length is a weak predictor. Task length has a weak negative correlation with performance (Spearman rho = -0.29), and no significant association was found from coarse scene features such as clutter level or object count.
  • Recurring failure taxonomy. Task errors include skipped prerequisites, incomplete outcomes, and domain-specific operations. Physics errors include object inconsistencies (unexplained changes in shape, material, identity, assembly, count, appearance, disappearance, duplication, or merging), boundary violations (objects passing through closed lids or solid surfaces), and mechanism violations (articulated objects rotating about incorrect axes, coupled parts moving independently, rigid structures deforming).
  • The evaluator aligns with humans and beats baselines. Ego2ActJudge reaches r = 0.69, Kendall tau = 0.46, CCC = 0.61, and MAE = 22.0, compared with r = 0.42 (WRA), 0.40 (RBench), 0.32 (Simple VQA), 0.31 (PQSG), 0.16 (WMBench), and -0.10 (VideoScore). Humans and Ego2ActJudge rank the six generators almost identically (Spearman rho = 0.94), swapping only the adjacent Grok-1.5 and MiniMax-H3.
  • Prior evaluators have specific blind spots. Perception-heavy evaluators such as VideoScore score positive and negative human controls almost identically, and action-oriented evaluators such as PQSG and RBench correlate less well overall because their scores stay insensitive to drops in physical plausibility.
  • Strict sequential gating matters. Removing the primary physics continuity gate (P1) causes a large error increase, and bypassing the task initiation gate (T1) also degrades alignment. Scoring all gates independently instead of sequentially inflates Final scores by 6.8 points and lowers Final r from 0.69 to 0.65 and CCC from 0.62 to 0.51.
  • Backbone choice affects agreement. Replacing Gemini-3.7-Flash with GPT-5.6-Terra keeps the judge aligned with humans but less closely on every axis: Task r 0.580 versus 0.624 (MAE 21.79 versus 19.49), Physics r 0.514 versus 0.568 (MAE 32.08 versus 29.37), and Final r 0.609 versus 0.694 (MAE 24.70 versus 22.10). The judge catches within-frame violations such as object duplication or identity changes but largely misses those that unfold between frames, so its Physics scores should be read as optimistic.
  • Subgoal plans are stable. Although the plan is re-derived for every video, 91% of plans keep the same subgoal count and name nearly the same objects.
  • Reliability of the rubric. Inter-annotator agreement is ICC = 0.863 (Task 0.831, Physics 0.834) and Krippendorff's alpha = 0.745 (Task 0.691, Physics 0.697).

Methodology in Plain English

The researchers hand-built a benchmark from real life. Nine annotators recorded first-person videos of everyday activities across five domains: kitchen and food preparation, household organization and storage, personal care and utilities, office and study workspace, and other everyday activities. Each case pairs an initial scene image with an outcome-oriented goal that states the desired end state but not the steps, and each case must admit at least three genuinely distinct execution sequences. Every case includes six human recordings (three successful attempts and three unsuccessful attempts as controls). Each scene was also fed to six video generation models (Grok Imagine 1.5, Seedance 2.0, Cosmos 3, Kling v3 Pro, MiniMax H3, and Wan 2.7), producing three videos per model per case at seeds 101, 202, and 303, for 18 generated videos per case. Successful human executions require more than five steps on average and take a median of 20.4 seconds.

For scoring, they wrote human rubrics and turned them into a multi-stage question-answering pipeline. First the judge looks only at the goal and the initial scene to derive a minimal list of observable subgoals. Then each subgoal is checked against three task gates (T1 action initiation, T2 coarse interaction process, T3 post-process state matches the outcome) and four physics gates (P1 state and object continuity, P2 physically valid causation at interaction onset, P3 plausible contact during interaction, P4 post-interaction physical stability). Gates are applied in order with early exit at the first failure, producing discrete subgoal scores. Task score T is the mean over subgoals; Physics score P is averaged only over subgoals that were actually attempted (unattempted subgoals are marked N/A); and the final score is the geometric mean S = sqrt(T·P), with an unattempted rollout assigned S = 0. Human evaluation used three annotators who scored anonymized videos from a 25-case panel (600 videos) while blind to which model produced them, and the automated judge was benchmarked against Simple VQA, PQSG, RBench, WR-Arena, WorldModelBench, and VideoScore.

Why This Matters

  • Impact on research: The paper provides a testbed and a reference-free scoring pipeline for a capability that prior egocentric benchmarks either skipped or evaluated only with fidelity metrics, and it documents concrete failure modes that future world-model work can target.
  • Robotics and imitation learning: Generated manipulation rollouts used as training data or planning predictions need to be functionally valid, and this benchmark measures exactly that.
  • Home and assistive robotics: Tasks like packing a bag, organizing a workspace, or preparing food are the everyday, multi-step goals the benchmark covers.
  • Evaluation tooling: The judge offers a scalable alternative to reference-video comparison for open-ended generation, where many valid action paths exist.
  • Industry relevance: The evaluated systems are commercial and open-weight video generators, and the leaderboard shows a wide spread from 64.0 down to 3.9 on human-judged aggregates, giving model developers a diagnostic target for contact-rich dynamics and object persistence.

Future Directions

  • Improving judges with denser temporal evidence. The authors state that the most direct path to a stronger judge is denser temporal evidence, since the current judge misses violations that unfold between frames rather than within a frame.
  • Closing the step-completion gap. Future models must execute every required step rather than omitting prerequisite states or leaving actions incomplete.
  • Persistent world modeling. Maintaining object identity and state under dynamic egocentric motion and frequent hand occlusions remains an open problem.
  • Grounded contact dynamics. Contact-rich interactions such as attachment/connection (38.6% Task, 66.7% Physics) and material transfer/mixing need physically plausible simulation before these models can serve as reliable simulators for embodied planning.

Target Audience

Researchers and engineers working on video generation, world models, embodied AI, and robot learning; benchmark and evaluation designers interested in rubric-based, reference-free judging; and practitioners who want to know how far current commercial and open-weight video generators are from producing usable goal-directed manipulation simulations.

Authors’ abstract

Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilities is crucial, existing benchmarks focus mainly on single short actions or step-by-step instructions. This leaves multi-step physical reasoning underexplored, especially in egocentric video generation that requires planning to simulate proper execution to accomplish high-level goals by carrying out multiple real-world manipulations. We introduce Ego2Act, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity. Given an initial scene image and a high-level goal, Ego2Act evaluates whether video generation models can produce realistic egocentric videos of a hand manipulating objects to carry out the task. To support scalable evaluation, we also introduce Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines. Our findings reveal that models' generated simulations often skip or partially execute steps, leaving later steps missing dependent states, which leads to unfulfilled goal. Furthermore, models consistently fail at fine-grained physical dynamics, particularly during complex object manipulation and persistent world modeling. We hope Ego2Act provides a rigorous testbed for advancing video models toward physically plausible, goal-directed simulation.

Read the original paper