Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
Authors: Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu, Xiaoye Wang, Yufan Chen, Junwei Zheng, +8 more
Organizations: INSAIT, Sofia University “St. Kliment Ohridski” · University of Würzburg · Technical University of Munich · Institute of Information Engineering · Karlsruhe Institute of Technology
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 567 single-video and 1,487 multi-video questions derived from 126 human-collected egocentric recordings covering 39 outdoor routes. Repeated traversals across movement speeds and lighting conditions ground comparisons in shared physical environments; 531 questions require alignment across independent recordings. Single-video questions measure the local visual, spatial, and motion evidence available to a model, while multi-video questions test whether evidence remains bound to the correct observation and can be composed into consistent route relationships. We report 29 single-video and 31 multi-video MLLM configurations across six model families in the main leaderboard. Among the 20 configurations evaluated comparably on both splits, every model performs worse on multi-video questions, with a mean decrease of 22.5 percentage points, and the gap persists when answer format and scoring are held fixed. The gap is not explained simply by additional videos or recording boundaries. The central bottlenecks are observation--evidence binding and ordered route-state tracking. The code and benchmark are publicly available at https://github.com/lei-qi-233/EgoGears.
Figures & tables
Benchmark
Video source
#QA
Video hours
QA input
Route understanding and acquisition design
Route-structure QA
Repeat-path protocol
Multiple speed conditions
Day/night counterparts
EgoSchema ( Mangalam et al., 2023 )
Ego4D
>5 K
>250
S
—
—
—
—
EgoTempo ( Plizzari et al., 2025 )
Ego4D
500
∼ 4.6
S
—
—
—
—
EgoCross ( Li et al., 2026b )
Five video datasets
957
NR
S
—
—
—
—
EgoNight ( Zhang et al., 2026 )
Real + synthetic
3,658
NR
M
—
—
—
✓
EgoExoMem ( Liu et al., 2026 )
Ego-Exo4D + LEMMA
∼ 2.6K
NR
Synchronized ego–exo
—
—
—
—
Table 1: Comparison with representative egocentric and cross-video QA benchmarks. S: single-video QA; M: joint reasoning over multiple videos or clips. NR: a reliable total video duration was not established from the paper.
Figure 1: Overview of the EgoGears collection and annotation pipeline. Controlled outdoor route recordings feed separate candidate-generation paths for EgoGears-Single and EgoGears-Multi. Automated gates screen structure, leakage, guessability, and visual support before human review. Independent post-review verification and adjudication then retain, revise, repair, or remove questions to produce the final paired benchmark.
Figure 2: EgoGears statistics. (a) Route coverage: of the 39 routes, 29 have multiple recordings, 21 cover all three movement speeds, 12 have both day and night recordings that together cover all three speeds, and 9 are fully crossed, with one recording under each of the six speed × lighting conditions. (b) EgoGears-Single question types ( N=567 ). (c) Number of clips per EgoGears-Multi question ( N=1,487 ). (d) EgoGears-Multi question types; the seven most frequent of the 29 fine-grained types are shown individually, and the remaining 22 are grouped.
Figure 3: Mean exact-set accuracy on the 9 fully crossed routes (29 single-video and 31 multi-video configurations).
Figure 4: Qualitative cases illustrating the benefits and limitations of explicit reasoning. (a) On single-video ego-relative spatial QA, three configurations fail without reasoning but recover the exact three-option answer when reasoning is enabled. (b) Cross-recording correspondence also requires matching the same place across opposite travel directions.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Single-video ( N=567 )
Multi-video ( N=1,487 )
Model
Mode
Exact
Jaccard
Exact
Jaccard
κ531
Gemini (API)
3.8-Flash
API
–
–
78.3
83.1
82.7
3.1-Pro
API
–
–
66.5
76.5
68.1
2.5-Pro
API
–
–
56.4
68.6
55.4
Qwen3.5
Appendix
Table 4: Exact-set versus partial-credit accuracy (%). Jaccard is the mean per-question ∣P∩G∣/∣P∪G∣ between the predicted set P and the reference set G ; failed or missing answers score 0 . κ531 is chance-normalized exact accuracy on the 531 questions spanning independent recordings (random baseline 4.47% ). Configurations and settings match Tables 3 and 3 ; – marks a configuration not evaluated on that split.
All (exact)
SC (exact)
MS (exact)
All (Jaccard)
Model
Mode
S
M
Δ
S
M
Δ
S
M
Δ
S
M
Δ
Qwen3.5-27B
T
68.6
40.8
27.8
69.5
45.2
24.3
67.4
38.7
28.7
73.1
56.3
16.8
Qwen3.5-122B-A10B
T
68.6
37.8
30.8
68.9
42.7
26.2
68.2
35.5
32.7
74.1
53.9
20.2
Qwen3.5-35B-A3B
T
66.5
37.9
28.6
67.7
46.1
21.6
64.9
34.1
30.8
70.9
53.6
17.3
Qwen3.5-9B
T
58.6
21.3
37.3
59.7
34.6
25.1
57.0
15.2
41.8
63.0
38.4
24.6
Gemma-4-31B
T
68.4
56.2
12.2
69.5
62.8
6.7
66.9
53.1
13.8
74.4
69.5
4.9
Appendix
Table 5: Single- versus multi-video accuracy (%) for the 20 configurations evaluated in comparable modes on both splits. Both splits are drawn from the same recordings. SC and MS restrict each split to its single-choice or multi-select questions; Jaccard is the partial-credit score of Table 4 . Δ is the single-video minus multi-video score.
Model
Mode
Land. (337)
Event (287)
Turn (239)
Route (239)
Spatial (154)
Match (118)
Env. (113)
MST (87)
Random
–
7.3
4.9
6.8
10.4
2.4
6.6
6.6
15.5
Gemini (API)
3.8-Flash
API
79.8
77.7
64.0
82.0
85.7
85.6
79.6
33.3
3.1-Pro
API
67.7
65.9
50.6
77.4
66.2
73.7
68.1
29.9
2.5-Pro
API
53.7
55.1
39.7
71.1
60.4
61.0
61.9
16.1
Qwen3.5
Appendix
Table 6: Multi-video QA by capability ( N=1,487 ). Exact-set accuracy (%) on the seven capability groups, with question counts in the header. MST is the multi-step-turns question type over whole-video evidence ( n=87 ), a subset of Turn. Configurations and modes match Table 3 .
Model
Mode
All
Human
95% CI
Gemini (API)
3.8-Flash
API
78.3
90.6
[87.6, 93.5]
3.1-Pro
API
66.5
81.3
[77.5, 84.9]
2.5-Pro
API
56.4
70.8
[66.5, 74.6]
Qwen3.5
122B-A10B
T
37.8
49.0
[44.3, 53.5]
Appendix
Table 7: Multi-video exact accuracy on human-confirmed questions (%). All: the full set ( N=1,487 , as in Table 3 ). Human: the 445 questions whose reference answers the reviewers confirmed ( 65 after option repair), with 95% bootstrap confidence intervals.
Figure 5: Cross-recording side and order reasoning. The same physical structures must be aligned across observations before direction-dependent relations can be reversed and combined consistently.
Figure 6: Whole-video start–end correspondence. The models must match distant route phases through a shared landmark; reasoning leaves both the correct and incorrect model predictions unchanged.
Figure 7: Single-video cases where reasoning has opposite effects across models. Correct direct answers may be destabilized even when another model is rescued on the same question.
Figure 8: Multi-video regressions and corrections under reasoning. The intervention can improve evidence binding for one model while inducing an inconsistent or wrong-cardinality answer in another.
Figure 9: Reasoning rescue versus persistent failure in multi-video QA. Panel (a) contrasts successful and unsuccessful cross-clip evidence integration; panel (b) shows that added computation does not repair an incorrectly inferred turn sequence.
Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate reasoning steps, and most provide answers only in the text domain. We introduce Minerva-Ego, a benchmark for evaluating complex egocentric visual reasoning. We extend recent high-quality video data sources recorded from egocentric / embodied settings with a set of challenging, multi-step multimodal questions and spatiotemporally-dense human-annotated reasoning traces. Benchmarking experiments show that state-of-the-art models still have a large gap to human performance. To investigate this gap in detail, we annotate each reasoning trace in the dataset with the objects of interest required to solve the question, as spatiotemporal mask annotations. Through extensive evaluations, we identify that prompting frontier models with hints of 'where' and 'when' to look yields substantial improvements in performance. Minerva-Ego can be downloaded at https://github.com/google-deepmind/neptune.
Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on Video Question Answering (VideoQA). Nevertheless, existing benchmarks are predominantly evaluated through answer correctness, while the faithfulness of predicted evidence supporting those answers remains insufficiently evaluated. This disconnect between answer generation and evidence verification motivates the construction of the Evidence-Grounded Video Question Answering Benchmark (EG-VQA), a large-scale open-ended benchmark in which each QA pair is annotated with temporally localized textual evidence, requiring models to jointly produce answers and verifiable evidence. EG-VQA comprises 2,067 videos and 11,838 QA pairs with fine-grained evidence annotations. To evaluate predicted evidence, Evidence-Grounded F1 (EG-F1) is introduced as a metric that jointly measures temporal alignment and semantic consistency between predicted and ground-truth evidence. Experiments reveal a substantial discrepancy between answer correctness and evidence faithfulness: even strong proprietary models can achieve high answer accuracy while exhibiting gaps in temporal-semantic grounding. To investigate evidence-aware learning under EG-VQA, we develop EG-Reasoner, an evidence-aware model trained with explicit evidence supervision. EG-Reasoner achieves strong evidence grounding performance among open-source models while maintaining competitive answer generation results compared with proprietary systems. These findings highlight the importance of explicitly modeling evidence grounding and suggest that evidence-aware supervision provides an effective direction toward more reliable and interpretable VideoQA systems.
Linpeng Huang, Weixing Chen, Zexin Chen +2
Sun Yat-sen University · Shenzhen University · Peng Cheng Laboratory
Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to assess whether models can maintain consistent memory across days or weeks of real-world experience. We introduce EgoMonth, the first month-level egocentric video understanding benchmark. EgoMonth comprises over 300 hours of first-person daily-life recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice question-answer pairs. We design a cognitively grounded 14-task evaluation framework organized into three hierarchical cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that even the best-performing model, Gemini 2.5 Pro, achieves only 71.8% macro-average accuracy, remaining 22.4 percentage points below the corrected human baseline of 94.2%. Several models perform near or below the 25% chance level on tasks such as Route Reasoning, Cross-view Spatial Reasoning, and Direction Judgement, while even the strongest closed-source model remains substantially below human performance. These results indicate that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.
Weitao Chen, Hu Jiaxin, Xie Tianyidan +15
Nanjing University, China · Huawei Technologies Co., Ltd., China · Tianjin University, China