EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
Authors: Shulin Tian, Junsu Kim, Shuai Liu, Hao Li, Yujiao Shen, Sihan Li, Zhe Yang, Yeongon Kim, +12 more
Organizations: S-Lab, Nanyang Technological University · A*STAR · KAIST · PKU · FDU · School of Biological Sciences, Nanyang Technological University · University of Catania
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding them requires reasoning about affordances, hand-tool-object geometry, procedural progress, and causal effects on target objects. Yet despite strong performance on perception-oriented video tasks such as captioning and general video QA, current multimodal video models remain limited in this form of tool-centric embodied reasoning. Progress in this direction has been limited by the lack of real-world egocentric data and diagnostic benchmarks. To address this gap, we introduce EgoTools, the first comprehensive suite for egocentric tool-use understanding. It consists of two complementary components: EgoTools-Data, a large-scale corpus of 100 hours of tool-centric egocentric recordings with synchronized audio, dense captions, reasoning-heavy narrations, and supplementary 3D information; and EgoTools-Bench, a diagnostic benchmark of 1,000 QA pairs across four tracks that cover tool-use understanding from perception and geometry to procedure and causal reasoning. Experimental results show that current models still struggle to ground tool use in visual evidence: Gemini-3.1-Pro achieves 66.9% overall accuracy but only 51.7% on Perception & Grounding. Beyond evaluation, we validate EgoTools-Data as a training resource. On the full 1,000-question benchmark, full supervised fine-tuning improves Qwen3-VL-8B-Instruct from 50.0% to 60.9%, under strict source-video separation. Together, these results establish EgoTools as a unified resource for both training and diagnostic evaluation of real-world egocentric tool-use understanding.
Figures & tables
Table 1: Comparison with representative egocentric datasets and benchmarks. A checkmark indicates that the feature is a central design component; indicates partial support. Signals denote sensor/video inputs and exclude QA text or narrations used for benchmark construction: = RGB/video, = 3D/depth/pose, = inertial signals, and = gaze; non-RGB signals may be available only for subsets of large datasets. Hours are reported total video hours or computed from source-reported clip counts and durations; “–” denotes not applicable or not reported. Embodied interaction benchmarks are included as scope contrast rather than directly comparable egocentric video corpora.
Figure 2: Benchmark distribution and annotation examples. (a) Distribution of 1,000 QA pairs across four tool-use reasoning tracks and their fine-grained subtracks. (b) Example annotations, including captions, tool-centric narrations, and finalized QA pairs produced through human annotation with MLLM-assisted checking. (c) Word-cloud comparison with EPIC-KITCHENS-100 and Ego4D, showing that EgoTools-Bench is more tool-dense and centered on tools, actions, objects, and state changes.
Figure 3: Textual annotation pipeline. (a) Tool-centric narrations are annotated from grounded egocentric video segments. (b) Benchmark QA pairs are constructed and refined through an annotation interface with MLLM-assisted checking.
Figure 4: 3D annotation pipeline. Raw geometry from the HOMIE device is reconstructed with 3DGS to obtain consistent depth, then combined with tool-centric narrations for recursive VLM-guided object tracking. The resulting masks are lifted to 3D for object boxes and spatial QA generation.
Model
Frames
Audio
EgoSchema [ 32 ]
EgoPlan [ 8 ]
EgoThink [ 9 ]
EgoTools
AC
PG
PD
SR
Overall
Human Baseline
Human Expert
–
✓
∼ 76
–
–
82.1
85.2
83.3
82.7
83.2
Proprietary Models
Gemini-3.1-Pro [ 19 ]
1fps
✓
–
–
–
69.7
51.7
72.1
74.9
66.9
Gemini-3-Flash [ 17 ]
1fps
✓
–
–
–
64.5
50.0
73.0
72.6
64.4
Table 2: Performance comparison across egocentric video understanding benchmarks. We compare human performance, proprietary models, and open-source video-language models on three existing egocentric benchmarks and our EgoTools-Bench benchmark. The last five columns report results on EgoTools-Bench across four research-facing tracks and the overall average. Track abbreviations are: AC = Affordance & Causality, PG = Perception & Grounding, PD = Procedural Dynamics, and SR = Spatial Reasoning.
Figure 5: Composition of the EgoTools-Data instruction-tuning corpus. (a) The inner ring groups the 184,679 examples by what they supervise, and the outer ring gives the corresponding categories; segment angles are schematic rather than strictly proportional, so that the smallest categories remain readable. (b) Shares drawn to scale; the four tool-centric categories are nested under their group total. Exact numbers are listed in Table 10 .
Model
EgoTools-Bench
Existing Egocentric Benchmarks
AC
PG
PD
SR
Overall
EgoSchema
EgoPlan
EgoThink
Qwen3-VL-8B-Instruct
46.1
45.9
57.1
55.4
50.0
69.0
42.3
61.0
EgoTools-8B
55.4
59.7
68.1
44.4
60.9
67.6
35.2
64.0
Table 3: Training utility of EgoTools-Data. EgoTools-8B is Qwen3-VL-8B-Instruct after a single stage of full supervised fine-tuning on 184,679 EgoTools-derived instruction examples. On EgoTools-Bench, both models use the same standard 64-frame MP4-based evaluation protocol. The instruction-tuning corpus and EgoTools-Bench are disjoint at the source-video level. EgoTools-Bench results are reported on the full quality-controlled 1,000-question benchmark; EgoSchema is reported on the full 5,031-question set, EgoPlan-Bench on its 3,343-question validation set, and EgoThink as the macro-averaged score over its twelve single-image subtasks.
Input condition
Overall
AC
PG
PD
SR
Ordered 64 frames
51.7
47.4
50.0
57.1
57.6
Shuffled 64 frames
49.3
45.5
47.4
53.9
56.5
Middle frame
46.9
44.7
45.9
51.0
45.7
Text-only
42.7
39.0
36.1
51.4
47.8
Table 4: Input-evidence ablation on the full 1,000-question EgoTools-Bench. All conditions use the same Qwen3-VL-8B-Instruct backbone, deterministic decoding, and scoring configuration. The visual conditions use the same pre-extracted frame-list pipeline, with the ordered condition serving as the matched within-protocol control.
Model
Full Bench.
Implicit
Explicit
Δ
Gemini-3.1-Pro
66.9
65.2
72.9
+7.7
Gemini-3-Flash
64.4
63.1
68.0
+4.9
Gemini-3.1-Flash-Lite
55.4
54.3
59.5
+5.2
Qwen3-VL-8B-Instruct
50.0
49.1
53.1
+4.0
Qwen3-VL-8B-Thinking
45.7
45.8
45.2
-0.6
Table 5: Performance on implicit and explicit tool-use questions. Δ denotes Explicit − Implicit.
Figure 6: Jaccard overlap of model errors on EgoTools-Bench. This highlights shared failure patterns within and across model families.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Role
# Contributors
Main Responsibility
Participant / narrator
16
Record egocentric videos and provide tool-centric narrations
QA annotator
10
Write human-crafted benchmark QA pairs
QA reviewer
2
Review answer validity, visual evidence, and distractor quality
Data curator
3
Curate videos, manage annotation progress, and organize benchmark data
Grounding annotator
1
Annotate focal-tool grounding for selected instances
Trajectory annotator
2
Construct or verify trajectory annotations for selected clips
Appendix
Table 6: Contributor roles in the documented subset. Counts are non-exclusive.
Table 7: Contributor backgrounds. Aggregate (a) degree and field and (b) age-group distributions for the documented subset.
Figure 7: Canonical-view rectification. Example input views and the rectified first-person frame.
Statistic
Value
Total EgoTools videos
646
Total video duration
100.37 hours
Hierarchical captions
361,332
Tool-centric narrations
6,519
Benchmark-reserved source duration
40.34 hours
Benchmark QA pairs
1,000
Appendix
Table 8: EgoTools dataset and benchmark statistics.
EgoTools-Data
EgoTools-Bench
Domain
# Videos
Hours
# Captions
# Narrations
# QAs
Kitchen
276
32.37
116,532
2,113
837
Classroom
74
6.86
24,696
467
15
Research Lab
61
30.92
111,312
1,887
99
Repair Workshop
84
12.34
44,424
763
16
Craft
67
10.17
36,612
657
23
Appendix
Table 9: Domain composition of EgoTools-Data and EgoTools-Bench. The benchmark is not domain-balanced.