Research
EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
Overview Research area: Computer vision and multimodal video-language models, specifically egocentric (first-person) video understanding and tool-use reasoning. Technical level: Advanced. The paper as

- arXiv
- 2609.39378
- Published
- 2026-09-30
- Authors
- Shulin Tian, Junsu Kim, Shuai Liu, Hao Li, Yujiao Shen, Sihan Li, Zhe Yang, Yeongon Kim, Feiyu Li, Jialin Wu, Yichi Zhang, Wenhui Wang, Runmao Yao, Yuhao Dong, Zhaoxi Chen, Fangzhou Hong, Antonino Furnari, Jingkang Yang, Hongyuan Zhu, Ziwei Liu
AI summary
Overview
- Research area: Computer vision and multimodal video-language models, specifically egocentric (first-person) video understanding and tool-use reasoning.
- Technical level: Advanced. The paper assumes familiarity with video-language models, supervised fine-tuning, egocentric datasets, and 3D reconstruction; the high-level argument is still accessible to a general reader.
- Scope: The paper introduces EgoTools, a combined dataset (EgoTools-Data, 100 hours of tool-centric egocentric video) and diagnostic benchmark (EgoTools-Bench, 1,000 multiple-choice QA pairs over four reasoning tracks), and shows that current multimodal models score far below human experts on tool-mediated reasoning while fine-tuning on the dataset improves results.
What This Paper Is About
Many everyday and professional tasks are mediated by tools, and understanding them requires reasoning about what a tool affords, how hands, tools and objects are arranged in space, where a procedure stands, and what effects an action has on a target object. Current multimodal video models do well on tasks such as captioning and general video QA, but the authors argue they remain weak at this kind of tool-centric embodied reasoning. The paper's goal is to close the gap on both fronts by supplying dense real-world egocentric training data and a benchmark that isolates tool-use reasoning from general activity recognition.
Key Contributions
- EgoTools-Data: A corpus of approximately 100 hours of tool-centric egocentric recordings with synchronized audio, dense hierarchical captions, reasoning-heavy tool-centric narrations with 2D grounding points, and supplementary 3D information derived from reconstruction and long-horizon object tracking. Videos are 1024 x 1024 at 20 FPS, captured with the HOMIE head-mounted device (four synchronized fisheye camera views) across seven tool-use domains (kitchen, classroom, research lab, repair workshop, craft, office, household).
- EgoTools-Bench: A diagnostic benchmark of 1,000 8-way multiple-choice QA pairs built from a 40.34-hour benchmark-reserved pool, comprising 900 human-crafted questions and 100 human-verified spatial questions, split into four tracks: Affordance & Causality (363), Perception & Grounding (236), Procedural Dynamics (222), and Spatial Reasoning (179).
- A video instruction-tuning corpus of 184,679 examples derived only from the training pool, labeled by the capability each example supervises, with the benchmark held out at the source-video level so no benchmark clip, caption, narration, or synthetic QA enters training.
- An empirical study showing that current models (human experts reach 83.2% overall; Gemini-3.1-Pro reaches 66.9%) struggle on the benchmark, and that single-stage full supervised fine-tuning of Qwen3-VL-8B-Instruct on the corpus raises its accuracy from 50.0% to 60.9%.
Main Findings
- Models fail to ground tool use in visual evidence: Gemini-3.1-Pro achieves 66.9% overall accuracy but only 51.7% on Perception & Grounding, which the authors note is 18-23 points below its scores on Affordance & Causality, Procedural Dynamics, and Spatial Reasoning.
- Open-source models are far below human level: Open-source instruct models range from 38.9% to 50.0% overall against a human expert score of 83.2%. The strongest instruct result is Qwen3-VL-8B-Instruct at 50.0%, followed by InternVL3.5-8B-Instruct at 48.7% and Qwen2.5-Omni-7B-Instruct at 47.1%. The strongest open-source thinking model is GLM4.1V-Thinking at 50.7%. Chance performance is 12.5% in the 8-way format.
- Broad first-person understanding does not transfer to tool reasoning: Qwen3-VL-8B-Instruct scores 69.0% on EgoSchema and 61.0% on EgoThink but 50.0% on EgoTools-Bench; Qwen3-VL-4B-Instruct scores 69.6% on EgoSchema, 61.9% on EgoThink, and 45.7% on EgoTools-Bench.
- EgoTools-Data is usable supervision: Fine-tuning improves Qwen3-VL-8B-Instruct from 50.0% to 60.9% (a 10.9 point gain), with +9.3 points on Affordance & Causality, +13.8 on Perception & Grounding, and +11.0 on Procedural Dynamics, while Spatial Reasoning drops by 11.0 points. Gains generalize to held-out source videos rather than memorized benchmark clips.
- Results on other benchmarks after fine-tuning are mixed: The tuned model reaches 67.6% on EgoSchema (1.4 points below the 69.0% backbone) and 64.0% macro accuracy on EgoThink (3.0 points above the 61.0% backbone), but only 35.2% on EgoPlan-Bench (7.1 points below the 42.3% backbone). The authors limit their primary claim to tool-centric reasoning rather than uniform improvement across all egocentric benchmarks.
- Video input matters, but language priors remain: In an input ablation, temporally ordered 64 frames give 51.7% overall, versus 49.3% for shuffled frames, 46.9% for a single middle frame, and 42.7% for text-only input — a 9.0 point gap between ordered video and text-only. Text-only performance is still above the 12.5% chance level, which the authors interpret as residual commonsense or answer-option cues rather than evidence that video is unnecessary.
- Perception & Grounding is the persistent bottleneck: Open-source instruct models score between 35.6% and 45.9% on this track, and it also shows the largest improvement from visual evidence over text-only input.
- Model variants show different track-wise strengths: Qwen3-VL-8B-Instruct improves over its 4B counterpart by 6.9 points on Procedural Dynamics and 5.4 on Spatial Reasoning, but only 0.5 on Perception & Grounding. Qwen2.5-Omni-7B-Instruct improves over Qwen2.5-VL-7B-Instruct on Affordance & Causality, Perception & Grounding and Procedural Dynamics but is slightly lower on Spatial Reasoning; Qwen3-VL-8B-Thinking underperforms its instruct counterpart on three tracks and only matches it on Spatial Reasoning.
- Annotation reliability was independently checked: On a stratified sample of 250 benchmark questions (100 of them Affordance & Causality), five annotators produced 500 judgments. Agreement with the original answer key was 89.8% overall and 86.5% on Affordance & Causality, exact agreement between the two assigned annotators was 88.4%, and 94.4% of items were judged to have a uniquely best answer supported by visible evidence. Fourteen items (5.6%) were flagged as ambiguous or multi-answer and were reviewed, revised, or replaced.
- The benchmark is tool-dense: A word-cloud comparison with EPIC-KITCHENS-100 and Ego4D indicates EgoTools-Bench places more lexical emphasis on tools, manipulated objects, actions, and state changes than on general egocentric activity recognition. The authors note the benchmark is strongly long-tailed across domains, with most questions drawn from kitchen videos, so they focus primary analysis on reasoning capabilities.
Methodology in Plain English
The authors collected raw first-person video using HOMIE, a head-mounted device with four synchronized fisheye cameras plus audio and motion signals. The front-left stream was rectified into a canonical first-person view (1024 x 1024, 20 FPS). Collection covered two modes: structured task-guided recordings, where participants followed predefined task sequences with clear steps and task boundaries, and open-ended participant-driven recordings, where they received only a broad topic and worked freely, allowing spontaneous tool choice, repeated attempts, error recovery, and substitution. Data was gathered across multiple kitchens, workshops, laboratories, and daily-living spaces at universities and research sites in Asia, with domain expertise matched to the task for expertise-intensive procedures.
Annotations come in three layers. Textually, Gemini-3-Flash generates dense captions in a streaming hierarchical pipeline: each video is split into 5-minute clips, 32 uniformly sampled frames produce a global context caption, and each clip is then captioned sequentially in 5-second chunks, with later chunks using the previous caption as memory. Chunk captions are aggregated into 1-minute windows and 5-minute summaries. Separately, the participants who performed the tasks write tool-centric narrations emphasizing tool selection, usage intent, and effects on target objects, and mark 2D grounding points on keyframes for focal tools or objects. Geometrically, raw depth and camera poses are reconstructed with 3D Gaussian Splatting (following Holi-Spatial) for multi-view-consistent scene geometry and temporally coherent depth. Object tracking uses a recursive VLM-guided segmentation framework: narrations seed SAM2 masks, and when later frames contain additional grounding points a VLM verifies tracking consistency and re-initializes SAM2 from the point of drift. Tracked masks are lifted into 3D to yield object boxes and long-horizon 4D trajectories, which in turn drive spatial QA templates.
The recordings were partitioned at the source-video level into a training pool and a 40.34-hour benchmark pool before any instruction data was built. The benchmark was annotated by original data collectors (procedural and causal questions) plus external annotators (visually inferable but model-challenging points), formatted as 8-way multiple choice with one correct answer and seven distractors, with clips extracted as localized temporal windows around an anchor moment, most lasting at least three minutes and averaging 4.19 minutes per QA. A "fix-first" quality-control pipeline audits structural validity, video grounding, and answer choices, repairs repairable issues, and hardens flagged items by replacing distractors with visually plausible, length-balanced alternatives. Training used a single stage of full-parameter supervised fine-tuning of Qwen3-VL-8B-Instruct's language-model component, keeping the visual encoder and multimodal aligner frozen. Evaluation reports overall and track-wise accuracy; most open-source instruct models received 64 uniformly sampled frames, thinking models used 64 or 512 frames, and Gemini models received 1 FPS video input, with only synchronized non-narration audio for audio-capable models.
Why This Matters
Impact on research. The paper argues that tool use has been effectively invisible in existing annotations: prior benchmarks may label an action as "cook eggs" while the spatula and pan that mediate it go unmentioned. By foregrounding tools and providing source-video-disjoint training and evaluation resources, EgoTools gives the field a way to measure and train a capability that general activity recognition benchmarks do not isolate. The measured gap between high EgoSchema/EgoThink scores and low EgoTools-Bench scores is direct evidence that broad first-person understanding is not the same as tool-centric embodied reasoning.
Real-world applications.
- Assistive and wearable systems that help a user select, hold, or substitute the right tool during a task, drawing on affordance and spatial-relation reasoning.
- Robotics and embodied agents that must track object state and procedural progress while manipulating tools in kitchens, workshops, or labs.
- Instructional video and training platforms, including expert procedures such as wet-lab protocols, where tool choice and step ordering carry correctness implications.
- Automatic documentation, quality-checking, or review of procedural work in craft, repair, and laboratory settings, based on dense captions, narrations, and 4D object trajectories.
Industry relevance. The 10.9-point improvement from fine-tuning shows a practical path: a general 8B video-language model can be adapted to tool-centric reasoning using this corpus, and the authors release both the data and a model. Because the annotation pipeline combines synthetic captioning with human narrations and MLLM-assisted checking, the workflow is also a blueprint for building similar domain-specific corpora in industrial, clinical, or manufacturing settings.
Future Directions
- Spatial reasoning is unresolved. Spatial Reasoning decreased by 11.0 points after fine-tuning despite explicit spatial supervision in the training corpus (17,887 examples, 9.7%), and the authors state it remains a limitation of the current training recipe.
- Perception & Grounding remains the diagnostic bottleneck even for the strongest proprietary model, raising the question of what training signals or architectures would let models resolve visually similar tools, contact points, and local state changes.
- Generalization to other egocentric benchmarks is uneven. Gains on EgoThink came with regressions on EgoSchema and EgoPlan-Bench, and the authors explicitly note that their format-matched training subsets do not isolate format coverage from other effects of fine-tuning.
- Domain balance. EgoTools-Bench is strongly long-tailed across the seven domains, with the majority of questions from kitchen videos, so the authors focus on capability-level rather than domain-level analysis; extending balanced coverage is an open direction.
- Residual language priors. Text-only input still exceeds chance on the benchmark, meaning some questions remain answerable through commonsense or option cues despite anti-shortcut filtering.
Target Audience
This paper is most useful to researchers and engineers working on multimodal video-language models, egocentric and first-person vision, embodied reasoning, and robot or assistive manipulation. It also serves benchmark and dataset builders interested in annotation pipelines, source-video-disjoint evaluation design, and quality control for reasoning-heavy QA. Practitioners adapting general video-language models to specialized domains (laboratory, workshop, kitchen, office) will find the training-utility results and the released dataset and model directly actionable.
Authors’ abstract
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding them requires reasoning about affordances, hand-tool-object geometry, procedural progress, and causal effects on target objects. Yet despite strong performance on perception-oriented video tasks such as captioning and general video QA, current multimodal video models remain limited in this form of tool-centric embodied reasoning. Progress in this direction has been limited by the lack of real-world egocentric data and diagnostic benchmarks. To address this gap, we introduce EgoTools, the first comprehensive suite for egocentric tool-use understanding. It consists of two complementary components: EgoTools-Data, a large-scale corpus of 100 hours of tool-centric egocentric recordings with synchronized audio, dense captions, reasoning-heavy narrations, and supplementary 3D information; and EgoTools-Bench, a diagnostic benchmark of 1,000 QA pairs across four tracks that cover tool-use understanding from perception and geometry to procedure and causal reasoning. Experimental results show that current models still struggle to ground tool use in visual evidence: Gemini-3.1-Pro achieves 66.9% overall accuracy but only 51.7% on Perception & Grounding. Beyond evaluation, we validate EgoTools-Data as a training resource. On the full 1,000-question benchmark, full supervised fine-tuning improves Qwen3-VL-8B-Instruct from 50.0% to 60.9%, under strict source-video separation. Together, these results establish EgoTools as a unified resource for both training and diagnostic evaluation of real-world egocentric tool-use understanding.