Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
Authors: Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu, Xiaoye Wang, Yufan Chen, Junwei Zheng, +8 more
Organizations: INSAIT, Sofia University “St. Kliment Ohridski” · University of Würzburg · Technical University of Munich · Institute of Information Engineering · Karlsruhe Institute of Technology
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 567 single-video and 1,487 multi-video questions derived from 126 human-collected egocentric recordings covering 39 outdoor routes. Repeated traversals across movement speeds and lighting conditions ground comparisons in shared physical environments; 531 questions require alignment across independent recordings. Single-video questions measure the local visual, spatial, and motion evidence available to a model, while multi-video questions test whether evidence remains bound to the correct observation and can be composed into consistent route relationships. We report 29 single-video and 31 multi-video MLLM configurations across six model families in the main leaderboard. Among the 20 configurations evaluated comparably on both splits, every model performs worse on multi-video questions, with a mean decrease of 22.5 percentage points, and the gap persists when answer format and scoring are held fixed. The gap is not explained simply by additional videos or recording boundaries. The central bottlenecks are observation--evidence binding and ordered route-state tracking. The code and benchmark are publicly available at https://github.com/lei-qi-233/EgoGears.
Figures & tables
Benchmark
Video source
#QA
Video hours
QA input
Route understanding and acquisition design
Route-structure QA
Repeat-path protocol
Multiple speed conditions
Day/night counterparts
EgoSchema ( Mangalam et al., 2023 )
Ego4D
>5 K
>250
S
—
—
—
—
EgoTempo ( Plizzari et al., 2025 )
Ego4D
500
∼ 4.6
S
—
—
—
—
EgoCross ( Li et al., 2026b )
Five video datasets
957
NR
S
—
—
—
—
EgoNight ( Zhang et al., 2026 )
Real + synthetic
3,658
NR
M
—
—
—
✓
EgoExoMem ( Liu et al., 2026 )
Ego-Exo4D + LEMMA
∼ 2.6K
NR
Synchronized ego–exo
—
—
—
—
Table 1: Comparison with representative egocentric and cross-video QA benchmarks. S: single-video QA; M: joint reasoning over multiple videos or clips. NR: a reliable total video duration was not established from the paper.
Figure 1: Overview of the EgoGears collection and annotation pipeline. Controlled outdoor route recordings feed separate candidate-generation paths for EgoGears-Single and EgoGears-Multi. Automated gates screen structure, leakage, guessability, and visual support before human review. Independent post-review verification and adjudication then retain, revise, repair, or remove questions to produce the final paired benchmark.
Figure 2: EgoGears statistics. (a) Route coverage: of the 39 routes, 29 have multiple recordings, 21 cover all three movement speeds, 12 have both day and night recordings that together cover all three speeds, and 9 are fully crossed, with one recording under each of the six speed × lighting conditions. (b) EgoGears-Single question types ( N=567 ). (c) Number of clips per EgoGears-Multi question ( N=1,487 ). (d) EgoGears-Multi question types; the seven most frequent of the 29 fine-grained types are shown individually, and the remaining 22 are grouped.
Figure 3: Mean exact-set accuracy on the 9 fully crossed routes (29 single-video and 31 multi-video configurations).
Figure 4: Qualitative cases illustrating the benefits and limitations of explicit reasoning. (a) On single-video ego-relative spatial QA, three configurations fail without reasoning but recover the exact three-option answer when reasoning is enabled. (b) Cross-recording correspondence also requires matching the same place across opposite travel directions.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Single-video ( N=567 )
Multi-video ( N=1,487 )
Model
Mode
Exact
Jaccard
Exact
Jaccard
κ531
Gemini (API)
3.8-Flash
API
–
–
78.3
83.1
82.7
3.1-Pro
API
–
–
66.5
76.5
68.1
2.5-Pro
API
–
–
56.4
68.6
55.4
Qwen3.5
Appendix
Table 4: Exact-set versus partial-credit accuracy (%). Jaccard is the mean per-question ∣P∩G∣/∣P∪G∣ between the predicted set P and the reference set G ; failed or missing answers score 0 . κ531 is chance-normalized exact accuracy on the 531 questions spanning independent recordings (random baseline 4.47% ). Configurations and settings match Tables 3 and 3 ; – marks a configuration not evaluated on that split.
All (exact)
SC (exact)
MS (exact)
All (Jaccard)
Model
Mode
S
M
Δ
S
M
Δ
S
M
Δ
S
M
Δ
Qwen3.5-27B
T
68.6
40.8
27.8
69.5
45.2
24.3
67.4
38.7
28.7
73.1
56.3
16.8
Qwen3.5-122B-A10B
T
68.6
37.8
30.8
68.9
42.7
26.2
68.2
35.5
32.7
74.1
53.9
20.2
Qwen3.5-35B-A3B
T
66.5
37.9
28.6
67.7
46.1
21.6
64.9
34.1
30.8
70.9
53.6
17.3
Qwen3.5-9B
T
58.6
21.3
37.3
59.7
34.6
25.1
57.0
15.2
41.8
63.0
38.4
24.6
Gemma-4-31B
T
68.4
56.2
12.2
69.5
62.8
6.7
66.9
53.1
13.8
74.4
69.5
4.9
Appendix
Table 5: Single- versus multi-video accuracy (%) for the 20 configurations evaluated in comparable modes on both splits. Both splits are drawn from the same recordings. SC and MS restrict each split to its single-choice or multi-select questions; Jaccard is the partial-credit score of Table 4 . Δ is the single-video minus multi-video score.
Model
Mode
Land. (337)
Event (287)
Turn (239)
Route (239)
Spatial (154)
Match (118)
Env. (113)
MST (87)
Random
–
7.3
4.9
6.8
10.4
2.4
6.6
6.6
15.5
Gemini (API)
3.8-Flash
API
79.8
77.7
64.0
82.0
85.7
85.6
79.6
33.3
3.1-Pro
API
67.7
65.9
50.6
77.4
66.2
73.7
68.1
29.9
2.5-Pro
API
53.7
55.1
39.7
71.1
60.4
61.0
61.9
16.1
Qwen3.5
Appendix
Table 6: Multi-video QA by capability ( N=1,487 ). Exact-set accuracy (%) on the seven capability groups, with question counts in the header. MST is the multi-step-turns question type over whole-video evidence ( n=87 ), a subset of Turn. Configurations and modes match Table 3 .
Model
Mode
All
Human
95% CI
Gemini (API)
3.8-Flash
API
78.3
90.6
[87.6, 93.5]
3.1-Pro
API
66.5
81.3
[77.5, 84.9]
2.5-Pro
API
56.4
70.8
[66.5, 74.6]
Qwen3.5
122B-A10B
T
37.8
49.0
[44.3, 53.5]
Appendix
Table 7: Multi-video exact accuracy on human-confirmed questions (%). All: the full set ( N=1,487 , as in Table 3 ). Human: the 445 questions whose reference answers the reviewers confirmed ( 65 after option repair), with 95% bootstrap confidence intervals.
Figure 5: Cross-recording side and order reasoning. The same physical structures must be aligned across observations before direction-dependent relations can be reversed and combined consistently.
Figure 6: Whole-video start–end correspondence. The models must match distant route phases through a shared landmark; reasoning leaves both the correct and incorrect model predictions unchanged.
Figure 7: Single-video cases where reasoning has opposite effects across models. Correct direct answers may be destabilized even when another model is rescued on the same question.
Figure 8: Multi-video regressions and corrections under reasoning. The intervention can improve evidence binding for one model while inducing an inconsistent or wrong-cardinality answer in another.
Figure 9: Reasoning rescue versus persistent failure in multi-video QA. Panel (a) contrasts successful and unsuccessful cross-clip evidence integration; panel (b) shows that added computation does not repair an incorrectly inferred turn sequence.