Research
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning Overview Research area: Computer vision, specifically multimodal large language models (MLLMs) applied to egocentric (first-person)

- arXiv
- 2511.18242
- Published
- 2025-11-23
- Authors
- Yogesh Kulkarni, Pooyan Fazli
AI summary
EgoVITA: Learning to Plan and Verify for Egocentric Video ReasoningOverview
Research area: Computer vision, specifically multimodal large language models (MLLMs) applied to egocentric (first-person) video understanding, procedural reasoning, and reinforcement learning for multimodal reasoning.
Technical level: Advanced. The paper assumes familiarity with MLLM architectures, policy-gradient reinforcement learning (GRPO), preference-style objectives, and video question-answering benchmarks.
Scope in one sentence: The paper introduces EgoVITA, a reinforcement-learning framework that splits egocentric video reasoning into an egocentric planning stage and an exocentric verification stage, trained with GRPO and two dense rewards on 52k samples across three MLLM backbones.
What This Paper Is About
Egocentric video is hard for current MLLMs because the camera wearer's viewpoint shifts continuously, objects are frequently occluded, and the scene is only partially observable, so models must reason procedurally about how actions and object states unfold over time. Existing MLLMs often produce plausible but visually inconsistent or weakly grounded responses, and supervised fine-tuning on egocentric data improves in-domain accuracy while degrading performance on standard exocentric (third-person) benchmarks. EgoVITA's goal is to teach first-person procedural reasoning without losing third-person general ability, by having the model first generate a first-person action plan and then audit that plan through third-person reasoning over the same video, with no paired ego-exo video data required.
Key Contributions
-
A plan-then-verify decomposition. EgoVITA separates reasoning into an egocentric planner that predicts a causal, temporally ordered sequence of first-person actions (the
<ego_plan>block) and an exocentric verifier that audits that plan from a third-person perspective (the<exo_verify>block) before producing the final<answer>. This creates explicit cross-perspective feedback without needing synchronized ego-exo video pairs. -
Two dense reward signals for GRPO. Anticipatory Cross-Modal Grounding (ACMG) grounds each predicted plan clause against subsequent visual observations, and a confidence reward reinforces consistent third-person verification through a teacher-guided warm-up followed by self-ranking.
-
State-of-the-art egocentric results with preserved exocentric performance. EgoVITA improves Qwen2.5-VL-7B by +7.7 on EgoBlind and +4.4 on EgoOrient, using 52k training samples, while maintaining or improving exocentric benchmark scores.
-
Exocentric regularization against catastrophic forgetting. Periodic interleaving of GRPO updates with a lightweight cross-entropy step on the exocentric MSR-VTT dataset (λ_exo = 0.05) is used to preserve third-person capability.
Main Findings
-
Gains replicate across three backbones. For Qwen2.5-VL, EgoVITA improves EgoBlind by +7.7 (± 0.9), EgoPlan by +2.7 (± 0.6), EgoThink by +3.7 (± 0.7), EOC-Bench by +2.5 (± 0.6), and EgoOrient by +4.4 (± 0.8). InternVL-3.5-8B improves by +3.6 (± 0.5) on EgoBlind and +3.8 (± 0.7) on EgoOrient. Qwen3-VL-8B improves by +3.5 (± 0.7) and +2.3 (± 0.6) on the same tasks.
-
Exocentric performance is not sacrificed. Under EgoVITA, InternVL-3.5 gains +2.7 (± 0.6) on Video-MME and +2.7 (± 0.7) on LVBench. Qwen2.5-VL gains +1.6 (± 0.5) on MVBench, +2.4 (± 0.6) on Video-MME, +0.8 (± 0.5) on LVBench, and +0.9 (± 0.5) on TOMATO. The Qwen3-VL LVBench gain of +0.4 (± 0.4) is marked as not statistically significant at the 0.05 level.
-
The framework, not the teacher, drives the improvement. A self-trained variant (EgoVITA ♢) using Qwen3-VL-8B to generate both the SFT data and the teacher confidence signals still outperforms the Qwen3-VL baseline on all egocentric benchmarks: EgoBlind 50.2 (+1.8), EgoPlan 34.1 (+0.4), EgoThink 63.0 (+0.3), EOC-Bench 47.4 (+0.6), EgoOrient 61.5 (+0.7). It is roughly flat on exocentric tasks (MVBench -0.1, Video-MME +0.1, LVBench -0.1, TOMATO -0.1).
-
Reinforcement learning, not just SFT, is what recovers exocentric ability. Qwen2.5-VL after SFT alone drops by -2.0 on MVBench and -2.1 on Video-MME. Applying EgoVITA's RL recovers MVBench and Video-MME by +3.6 and +4.5 points over SFT while adding +2.6 on EgoBlind over SFT. For InternVL-3.5, RL improves Video-MME by +6.2 and EgoBlind by +2.2 over SFT.
-
Dense rewards matter beyond format and answer rewards. GRPO with only format and answer rewards gives limited gains over SFT; adding ACMG and the confidence reward adds +1.8 on EgoBlind and +3.4 on EgoOrient for Qwen2.5-VL, and +1.6 / +1.7 on the same tasks for InternVL-3.5.
-
Teacher-guided initialization is not necessary. Removing the 200-step teacher phase still yields an average gain of +2.9 across the evaluated egocentric benchmarks (EgoBlind 36.0, +6.3; EgoPlan 31.8, +1.6; EgoThink 50.5, +2.3; EOC-Bench 43.0, +1.4). The authors attribute this to the SFT policy already being initialized from human-annotated data.
-
ACMG and confidence rewards are complementary. On Qwen2.5-VL, ACMG alone primarily helps temporal grounding (EgoPlan +0.6), the confidence reward alone helps verification consistency (EOC-Bench +1.3), and the combination gives a larger EOC-Bench gain (+1.9) than either alone.
-
Both reasoning stages are needed. Removing
<ego_plan>degrades EgoPlan by -1.8; removing<exo_verify>reduces Video-MME by -1.7. -
Anticipatory grounding beats present-frame grounding. A present-frame baseline (grounding on frame t) improves EgoBlind by +4.1, EgoPlan by +1.0, EgoThink by +1.9, and EOC-Bench by +1.2, while ACMG yields +7.7, +2.7, +3.7, and +2.5 respectively.
-
Category-specific gains over SFT. On EgoBlind, the largest improvements over SFT are in Safety Warning (+5.4) and Other Resources (+3.8). On EOC-Bench, Future reasoning benefits most (+3.2). On EgoOrientBench, the Choose task improves by +5.5.
-
The reward responds to action dynamics, not static scenery. In a controlled analysis on 850 EgoPlan samples, correct action clauses score a mean ACMG similarity of μ = 0.68 while same-scene, wrong-action pairs drop to μ = 0.35 (Δ = 0.33). Wrong-action frames were identified using Qwen3-VL-30B-A3B.
-
Favorable comparison against a prior method. Against EgoThinker (built on Qwen2-VL and trained with approximately 5M samples), EgoVITA reaches 46.8 on EgoBlind, +5.8 (± 0.3) over the base Qwen2-VL, while EgoThinker shows degradation on exocentric tasks (-1.8 on MMMU, -10.5 on DocVQA, -33 on MME, -2.3 on VQAv2). EgoVITA also gains +0.3 (± 0.4) on MMMU, +0.6 (± 0.3) on DocVQA, +5 (± 7) on MME, and +0.8 (± 0.6) on VQAv2 over the base model.
-
Attention becomes more task-focused. EAGLE visualizations
Authors’ abstract
Egocentric video understanding requires procedural reasoning under partial observability and continuously shifting viewpoints. Current multimodal large language models (MLLMs) struggle with this setting, often generating plausible but visually inconsistent or weakly grounded responses. We introduce $\textbf{EgoVITA}$, a framework that decomposes egocentric video reasoning into a structured $\textit{plan-then-verify}$ process. The model first generates an $\textbf{egocentric plan}$: a causal sequence of anticipated actions from a first-person perspective. This plan is then evaluated by an $\textbf{exocentric verification}$ stage that uses third-person reasoning over the same video to verify its spatiotemporal and logical consistency, without exocentric video input. This decomposition enables cross-perspective feedback without requiring paired ego-exo supervision. To train this reasoning process, we adopt Group Relative Policy Optimization (GRPO) with two dense reward signals: one that grounds anticipated actions in subsequent visual observations and another that reinforces consistent third-person verification. $\textbf{EgoVITA}$ achieves state-of-the-art performance on egocentric reasoning benchmarks, outperforming Qwen2.5-VL-7B by $\mathbf{+7.7}$ on EgoBlind and $\mathbf{+4.4}$ on EgoOrient, while maintaining strong generalization on exocentric video tasks with only $52k$ training samples.