EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
Authors: Shulin Tian, Junsu Kim, Shuai Liu, Hao Li, Yujiao Shen, Sihan Li, Zhe Yang, Yeongon Kim, +12 more
Organizations: S-Lab, Nanyang Technological University · A*STAR · KAIST · PKU · FDU · School of Biological Sciences, Nanyang Technological University · University of Catania
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding them requires reasoning about affordances, hand-tool-object geometry, procedural progress, and causal effects on target objects. Yet despite strong performance on perception-oriented video tasks such as captioning and general video QA, current multimodal video models remain limited in this form of tool-centric embodied reasoning. Progress in this direction has been limited by the lack of real-world egocentric data and diagnostic benchmarks. To address this gap, we introduce EgoTools, the first comprehensive suite for egocentric tool-use understanding. It consists of two complementary components: EgoTools-Data, a large-scale corpus of 100 hours of tool-centric egocentric recordings with synchronized audio, dense captions, reasoning-heavy narrations, and supplementary 3D information; and EgoTools-Bench, a diagnostic benchmark of 1,000 QA pairs across four tracks that cover tool-use understanding from perception and geometry to procedure and causal reasoning. Experimental results show that current models still struggle to ground tool use in visual evidence: Gemini-3.1-Pro achieves 66.9% overall accuracy but only 51.7% on Perception & Grounding. Beyond evaluation, we validate EgoTools-Data as a training resource. On the full 1,000-question benchmark, full supervised fine-tuning improves Qwen3-VL-8B-Instruct from 50.0% to 60.9%, under strict source-video separation. Together, these results establish EgoTools as a unified resource for both training and diagnostic evaluation of real-world egocentric tool-use understanding.
Figures & tables
Table 1: Comparison with representative egocentric datasets and benchmarks. A checkmark indicates that the feature is a central design component; indicates partial support. Signals denote sensor/video inputs and exclude QA text or narrations used for benchmark construction: = RGB/video, = 3D/depth/pose, = inertial signals, and = gaze; non-RGB signals may be available only for subsets of large datasets. Hours are reported total video hours or computed from source-reported clip counts and durations; “–” denotes not applicable or not reported. Embodied interaction benchmarks are included as scope contrast rather than directly comparable egocentric video corpora.
Figure 2: Benchmark distribution and annotation examples. (a) Distribution of 1,000 QA pairs across four tool-use reasoning tracks and their fine-grained subtracks. (b) Example annotations, including captions, tool-centric narrations, and finalized QA pairs produced through human annotation with MLLM-assisted checking. (c) Word-cloud comparison with EPIC-KITCHENS-100 and Ego4D, showing that EgoTools-Bench is more tool-dense and centered on tools, actions, objects, and state changes.
Figure 3: Textual annotation pipeline. (a) Tool-centric narrations are annotated from grounded egocentric video segments. (b) Benchmark QA pairs are constructed and refined through an annotation interface with MLLM-assisted checking.
Figure 4: 3D annotation pipeline. Raw geometry from the HOMIE device is reconstructed with 3DGS to obtain consistent depth, then combined with tool-centric narrations for recursive VLM-guided object tracking. The resulting masks are lifted to 3D for object boxes and spatial QA generation.
Model
Frames
Audio
EgoSchema [ 32 ]
EgoPlan [ 8 ]
EgoThink [ 9 ]
EgoTools
AC
PG
PD
SR
Overall
Human Baseline
Human Expert
–
✓
∼ 76
–
–
82.1
85.2
83.3
82.7
83.2
Proprietary Models
Gemini-3.1-Pro [ 19 ]
1fps
✓
–
–
–
69.7
51.7
72.1
74.9
66.9
Gemini-3-Flash [ 17 ]
1fps
✓
–
–
–
64.5
50.0
73.0
72.6
64.4
Table 2: Performance comparison across egocentric video understanding benchmarks. We compare human performance, proprietary models, and open-source video-language models on three existing egocentric benchmarks and our EgoTools-Bench benchmark. The last five columns report results on EgoTools-Bench across four research-facing tracks and the overall average. Track abbreviations are: AC = Affordance & Causality, PG = Perception & Grounding, PD = Procedural Dynamics, and SR = Spatial Reasoning.
Figure 5: Composition of the EgoTools-Data instruction-tuning corpus. (a) The inner ring groups the 184,679 examples by what they supervise, and the outer ring gives the corresponding categories; segment angles are schematic rather than strictly proportional, so that the smallest categories remain readable. (b) Shares drawn to scale; the four tool-centric categories are nested under their group total. Exact numbers are listed in Table 10 .
Model
EgoTools-Bench
Existing Egocentric Benchmarks
AC
PG
PD
SR
Overall
EgoSchema
EgoPlan
EgoThink
Qwen3-VL-8B-Instruct
46.1
45.9
57.1
55.4
50.0
69.0
42.3
61.0
EgoTools-8B
55.4
59.7
68.1
44.4
60.9
67.6
35.2
64.0
Table 3: Training utility of EgoTools-Data. EgoTools-8B is Qwen3-VL-8B-Instruct after a single stage of full supervised fine-tuning on 184,679 EgoTools-derived instruction examples. On EgoTools-Bench, both models use the same standard 64-frame MP4-based evaluation protocol. The instruction-tuning corpus and EgoTools-Bench are disjoint at the source-video level. EgoTools-Bench results are reported on the full quality-controlled 1,000-question benchmark; EgoSchema is reported on the full 5,031-question set, EgoPlan-Bench on its 3,343-question validation set, and EgoThink as the macro-averaged score over its twelve single-image subtasks.
Input condition
Overall
AC
PG
PD
SR
Ordered 64 frames
51.7
47.4
50.0
57.1
57.6
Shuffled 64 frames
49.3
45.5
47.4
53.9
56.5
Middle frame
46.9
44.7
45.9
51.0
45.7
Text-only
42.7
39.0
36.1
51.4
47.8
Table 4: Input-evidence ablation on the full 1,000-question EgoTools-Bench. All conditions use the same Qwen3-VL-8B-Instruct backbone, deterministic decoding, and scoring configuration. The visual conditions use the same pre-extracted frame-list pipeline, with the ordered condition serving as the matched within-protocol control.
Model
Full Bench.
Implicit
Explicit
Δ
Gemini-3.1-Pro
66.9
65.2
72.9
+7.7
Gemini-3-Flash
64.4
63.1
68.0
+4.9
Gemini-3.1-Flash-Lite
55.4
54.3
59.5
+5.2
Qwen3-VL-8B-Instruct
50.0
49.1
53.1
+4.0
Qwen3-VL-8B-Thinking
45.7
45.8
45.2
-0.6
Table 5: Performance on implicit and explicit tool-use questions. Δ denotes Explicit − Implicit.
Figure 6: Jaccard overlap of model errors on EgoTools-Bench. This highlights shared failure patterns within and across model families.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Role
# Contributors
Main Responsibility
Participant / narrator
16
Record egocentric videos and provide tool-centric narrations
QA annotator
10
Write human-crafted benchmark QA pairs
QA reviewer
2
Review answer validity, visual evidence, and distractor quality
Data curator
3
Curate videos, manage annotation progress, and organize benchmark data
Grounding annotator
1
Annotate focal-tool grounding for selected instances
Trajectory annotator
2
Construct or verify trajectory annotations for selected clips
Appendix
Table 6: Contributor roles in the documented subset. Counts are non-exclusive.
Table 7: Contributor backgrounds. Aggregate (a) degree and field and (b) age-group distributions for the documented subset.
Figure 7: Canonical-view rectification. Example input views and the rectified first-person frame.
Statistic
Value
Total EgoTools videos
646
Total video duration
100.37 hours
Hierarchical captions
361,332
Tool-centric narrations
6,519
Benchmark-reserved source duration
40.34 hours
Benchmark QA pairs
1,000
Appendix
Table 8: EgoTools dataset and benchmark statistics.
EgoTools-Data
EgoTools-Bench
Domain
# Videos
Hours
# Captions
# Narrations
# QAs
Kitchen
276
32.37
116,532
2,113
837
Classroom
74
6.86
24,696
467
15
Research Lab
61
30.92
111,312
1,887
99
Repair Workshop
84
12.34
44,424
763
16
Craft
67
10.17
36,612
657
23
Appendix
Table 9: Domain composition of EgoTools-Data and EgoTools-Bench. The benchmark is not domain-balanced.
Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate reasoning steps, and most provide answers only in the text domain. We introduce Minerva-Ego, a benchmark for evaluating complex egocentric visual reasoning. We extend recent high-quality video data sources recorded from egocentric / embodied settings with a set of challenging, multi-step multimodal questions and spatiotemporally-dense human-annotated reasoning traces. Benchmarking experiments show that state-of-the-art models still have a large gap to human performance. To investigate this gap in detail, we annotate each reasoning trace in the dataset with the objects of interest required to solve the question, as spatiotemporal mask annotations. Through extensive evaluations, we identify that prompting frontier models with hints of 'where' and 'when' to look yields substantial improvements in performance. Minerva-Ego can be downloaded at https://github.com/google-deepmind/neptune.
The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the ability for MLLMs to recognize fine-grained hand-object interactions, track object state changes over time, and reason about manipulative processes in dynamic environments from a first-person perspective. However, existing egocentric video benchmarks suffer from \textbf{limited grounded rationale evaluation}, offering limited support for fine-grained operation-centric reasoning and rarely examining whether model rationales are grounded in explicit spatio-temporal evidence. To address this gap, we introduce \textbf{EgoCoT-Bench}, a fine-grained egocentric benchmark for grounded and verifiable operation-centric reasoning with explicit step-by-step rationale annotations. Overall, EgoCoT-Bench comprises 3,172 verifiable QA pairs over 351 egocentric videos separated into four task groups for a total of 12 sub-task groups, encompassing perception and retrospection, anticipation, and high-level reasoning. The benchmark is constructed through a spatio-temporal scene graphs (STSG) guided generation framework and is further refined by human annotators to ensure correctness, egocentric relevance and fine-grained quality. Experimental results show continuing difficulties with egocentric fine-grained reasoning and further reveal that many multimodal models produce explanations that are answer-correct, but have evidence that is inconsistent with the answer. We hope EgoCoT-Bench can serve as a useful testbed for grounded and verifiable reasoning in egocentric video understanding. Project page and supplementary materials are available at: https://dstardust.github.io/EgoCoT/.
Egocentric video understanding requires procedural reasoning under partial observability and continuously shifting viewpoints. Current multimodal large language models (MLLMs) struggle with this setting, often generating plausible but visually inconsistent or weakly grounded responses. We introduce EgoVITA, a framework that decomposes egocentric video reasoning into a structured plan-then-verify process. The model first generates an egocentric plan: a causal sequence of anticipated actions from a first-person perspective. This plan is then evaluated by an exocentric verification stage that uses third-person reasoning over the same video to verify its spatiotemporal and logical consistency, without exocentric video input. This decomposition enables cross-perspective feedback without requiring paired ego-exo supervision. To train this reasoning process, we adopt Group Relative Policy Optimization (GRPO) with two dense reward signals: one that grounds anticipated actions in subsequent visual observations and another that reinforces consistent third-person verification. EgoVITA achieves state-of-the-art performance on egocentric reasoning benchmarks, outperforming Qwen2.5-VL-7B by +7.7 on EgoBlind and +4.4 on EgoOrient, while maintaining strong generalization on exocentric video tasks with only 52k training samples.