Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.
Figures & tables
Figure 1: Moments record events; persistent entities connect them into biographies. To answer whether the mug used for coffee ended up in the dishwasher, a memory must know which physical mug took part in each event. A descriptive memory may retrieve “coffee is poured into a red mug” and “a red mug is placed in the dishwasher” yet cannot tell whether the two mugs are the same.
Figure 2: Overview of Grounded Entity Biographies (GEB). Visually grounded observations of each instance are associated across clips into a persistent biography; here the blue hand mixer is linked across days, each observation keeping its description, its source frames and its episode context. At question time a controller retrieves biography excerpts together with episodic and visual evidence, and the answer model identifies Shure from them.
EgoLifeQA ↑
Ego-R1 ↑
MM-Lifelong ↑
Model
EL
ER
HI
RM
TM
Avg.
Manual
Gemini
Avg.
Week
Day
General MLLMs
GPT-5 †
47.2
42.1
47.5
53.6
55.6
48.6
–
–
–
15.00 ∗
15.25 ∗
Qwen3-VL-235B
44.0
44.4
50.8
44.8
50.8
46.0
40.0
62.7
51.3
15.63 ∗
12.44 ∗
Qwen3.5-35B
43.2
46.8
49.2
50.4
61.9
49.0
38.7
65.3
52.0
13.75
7.50
Long Video MLLMs
Table 1: GEB achieves the highest overall accuracy on each benchmark split shown. Accuracy ( % ). EL, ER, HI, RM, TM: EntityLog, EventRecall, HabitInsight, RelationMap, TaskMaster; Manual/Gemini: human-written/model-generated Ego-R1 questions. ∗ on a model name: EgoLifeQA and Ego-R1 results from Li et al. (2026a) ; ∗ on a cell: Chen et al. (2026) ; † : Yeo et al. (2026) . Other entries are our runs. GEB, MAGIC-Video, and WorldMM use the same Qwen3.5-35B controller and answer model. Bold / underline : best/second-best per column. Full comparisons in Tables 16 and 10 .
Temporal grounding
Answering
Model
# Frames
mIoP ↑
mIoG ↑
IoU@0.3 ↑
mIoU ↑
Sim. ↑
Score ↑
Human ‡
Full
71.8
81.0
87.0
61.8
74.3
–
Whole-clip MLLMs
GPT-4o ‡
60
18.9
24.4
12.0
12.2
73.7
–
InternVL2-8B ‡
30
11.8
24.0
6.3
6.6
71.9
3.33
LLaVA-NeXT-Video-7B ‡
32
–
–
–
–
62.1
2.91
Table 2: GEB leads the compared memory frameworks on MultiHop-EgoQA. Whole-clip models are separate references; Qwen3.5-35B reads 60 frames without memory. Full : complete video for memory construction, with evidence retrieved at question time. IoU@0.3 averages all questions; mIoP, mIoG, and mIoU average those predicting intervals. Sim. : sentence similarity. Score : 1–10 grading by gpt-oss-120b, averaged over three runs for memory frameworks. ‡ : grounding and Sim. from Chen et al. (2025) , Score from released models using the same judge (Appendix C ). Bold / underline : best/second-best within each block.
Figure 3: GEB improves access to annotated evidence over MAGIC-Video, with larger relative gains when evidence spans multiple intervals. (a) Percentage of the 484 annotated EgoLifeQA questions whose evidence window overlaps a received unit, overall and by question family (abbreviations in Table 1 ). (b) Percentage of MultiHop-EgoQA questions for which received units overlap every annotated interval, grouped by interval count. Brackets show relative gains over MAGIC-Video, and n gives the number of questions in each group.
Evidence reached
Read
Outcome
vs. ours
Memory
All ↑
All, ≥ 3 ↑
Clip ↓
Score ↑
mIoU ↑
Δ All [ 95% CI]
MAGIC-Video-Qwen3.5-35B
0.354
0.116
0.230
2.58 ± 0.07
17.6
−0.167[−0.190,−0.145]
6 units/round
0.449
0.182
0.319
2.75 ± 0.05
18.5
−0.071[−0.096,−0.048]
WorldMM-Qwen3.5-35B
0.526
0.244
0.399
2.74 ± 0.06
18.9
+0.005[−0.019,+0.030]
GEB-Qwen3.5-35B (ours)
0.521
0.235
0.343
3.11 ± 0.03
20.5
Ours with one decision removed
Table 3: GEB exceeds MAGIC-Video in complete coverage and is comparable to WorldMM with lower retrieved duration. All : fraction of questions with overlap for every annotated interval; All, ≥ 3 : the same over the 285 questions with three or more intervals. Clip : union of received intervals divided by clip duration. Score and mIoU follow Table 2 ; Score gives mean ± std across runs. Δ All: row minus GEB, with a bootstrap 95% confidence interval over questions. The lower block removes biography text or association.
Memory
Acc. ↑
Δ
Memory
Acc. ↑
Δ
GEB (full)
72.0
–
The index
The reading
w/o association
68.6
−3.4
w/o biography text (index only)
68.0
−4.0
identity keyed by name
69.2
−2.8
w/o biography text for the answer model
70.8
−1.2
descriptions appended to captions
68.2
−3.8
w/o identity notes
68.2
−3.8
w/o observation → timeline edges
68.2
−3.8
w/o unsearched-observation line
70.2
−1.8
Table 4: GEB benefits from physical-instance organization, contextual connections, and biography reading. EgoLifeQA accuracy (%) on 500 questions with fixed controller and answer model. Left block: the index. Right block: the reading. Δ : variant minus full GEB, in percentage points. Variant definitions: Section 4.4 ; additional edge ablations: Table 15 .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Count
Thirty-second video clips
6,266
Approximate video duration
52 hours
Tracked object observations
308,244
Entities after association
77,716
of which join two or more observations
27,446
of which span more than one day
15,365
Appendix
Table 5: Scale of the memory built for the EgoLife week. Entities are counted once each, including those carried across several days; the seven named people are one entity each.
WorldMM- Qwen3.5-35B
MAGIC-Video- Qwen3.5-35B
GEB-Qwen3.5-35B (ours)
Nodes
Episode nodes (captions at 4 granularities)
7,624
7,625
7,625
Visual clip nodes
6,223 †
6,223
6,223
Named-entity nodes extracted from text
55,468
2,669
–
Semantic triple nodes
3,821 †
3,821
–
Observation nodes
–
–
308,244
Appendix
Table 6: Memory statistics for the EgoLife A1 week , counted on the memory a question asked at the end of the week retrieves over. WorldMM is not one graph: it keeps a HippoRAG graph per caption granularity (passage and extracted-entity vertices, weighted relation edges) beside separate semantic-triple and visual-clip indices. GEB and MAGIC-Video share the episodic memory (episode, clip and temporal edges). GEB adds one node per tracked observation and one per entity with two or more observations. MAGIC-Video adds its text-derived entity and triple layer.
Input
Test@Week ↑
Test@Day ↑
1536 frames, longer side 512 pixels (the Qwen3.5-35B row of Table 1 )
13.75
7.50
1536 frames at the processor’s default budget, 160 × 96 per frame
9.25
6.50
256 frames as images at 151,200 pixels each, greedy
12.50
7.25
64 frames, longest edge 512
11.50
7.00
64 frames and their 64 captions
9.00
4.75
512 captions, no frames
9.00
3.25
Appendix
Table 7: Qwen3.5-35B alone on MM-Lifelong The answer model of every memory framework we run reads the recording directly, its inputs sampled uniformly over the whole recording, 200 questions per cell. The first two rows send 1536 frames as one video through the model’s own video processor, at 512 pixels on the longer side and at the processor’s default budget of 160 × 96 per frame. The third row sends 256 frames as separate images. The 64-frame rows use the protocol of Li et al. (2026a) for general MLLMs. Captions are the memory frameworks’ 30-second captions at the sampled positions, and the frame-only rows carry no transcript.
Memory
Tok
Fr
GEB-Qwen3.5-35B (ours)
130.0k
62.9
w/o visual frames
8.2k
0.0
Appendix
Table 8: Answering-context size with and without visual frames. Tok and Fr : mean answering-context tokens and frame references per question on EgoLifeQA. The second row is the w/o visual frames row of Table 4 ; controller, prompts, caps and video are the same.
MM-Lifelong ↑
Memory
Week
Δ
Day
Δ
GEB-Qwen3.5-35B (ours)
36.83 ± 0.88
–
17.58 ± 0.14
–
w/o biography text (index only)
30.08 ± 1.66
−6.75
15.25 ± 1.15
−2.33
descriptions appended to captions
33.58 ± 0.80
−3.25
12.50 ± 1.09
−5.08
MAGIC-Video-Qwen3.5-35B
30.92 ± 0.80
−5.91
10.08 ± 0.52
−7.50
WorldMM-Qwen3.5-35B
31.42 ± 2.27
−5.41
8.50 ± 2.63
−9.08
Appendix
Table 9: Withholding the biography text or appending the descriptions to the captions lowers the score on both splits. MM-Lifelong accuracy (%) under the benchmark’s GPT-5 judge, mean ± std over three runs. The ablation rows are defined as in Table 4 ; the two memory frameworks run under our protocol repeat the means of Table 1 with their std. Δ : difference to the full memory.
Model
# Frames
Test@Week ↑
Test@Day ↑
Human ∗
Full
95.6
99.2
General MLLMs
GPT-5 ∗
50
15.00
15.25
Qwen3-VL-235B ∗
1536
15.63
12.44
Qwen3-VL-30B ∗
1536
11.07
11.48
Qwen3.5-35B
1536
13.75
7.50
Appendix
Table 10: MM-Lifelong answer accuracy (%) with every published row of Chen et al. (2026) (their Tables 4 and 15) and every system we ran; entries marked with ∗ are taken from the original paper; MAGIC-Video, WorldMM and GEB report the mean over three runs; the rows without ∗ are run by us: the long-video models read 256 frames sampled uniformly at 1 fps with their released scripts. Bold marks the best number within each group.
Temporal grounding
Answering
Memory
mIoP ↑
mIoG ↑
IoU@0.3 ↑
mIoU ↑
Sim. ↑
Score ↑
MAGIC-Video-Qwen3.5-35B
29.6
30.7
19.0
17.6
61.3
2.58 ± 0.07
MAGIC-Video-Qwen3.5-35B, 6 units/round
30.2
33.2
20.7
18.5
63.1
2.75 ± 0.05
WorldMM-Qwen3.5-35B
30.6
34.2
21.5
18.9
62.2
2.74 ± 0.06
GEB-Qwen3.5-35B (ours)
32.0
36.7
25.9
20.5
65.2
3.11 ± 0.03
Ours with one decision removed
Appendix
Table 11: Under the benchmark’s released metrics, the three baseline rows and both variants lie below the full memory on mIoG, IoU@0.3, mIoU and Sim. Columns as in Table 2 , Score as the mean ± std over three runs. The upper block holds the memory frameworks of Table 2 and the six-unit volume control of Section 4.3 . The lower block removes one of the two decisions the memory rests on, as in Table 3 .
Score ↑
Category
n
MAGIC-Video- Qwen3.5-35B
MAGIC-Video- Qwen3.5-35B, 6 units/round
WorldMM- Qwen3.5-35B
GEB-Qwen3.5-35B (ours)
Qwen3.5-35B, all 60 frames
A: Repeated activities
73
2.91
3.00
2.97
3.24
4.10
B: Multiple actions
241
2.53
2.73
2.75
2.91
3.22
C: Multiple objects
193
2.40
2.51
2.43
2.90
3.47
D: Locations/people
72
2.51
2.79
2.59
3.25
3.85
E: Event composition
111
2.32
2.46
2.57
2.90
3.34
Appendix
Table 12: MultiHop-EgoQA judge score by question category (mean over three runs; n = judged questions). Categories follow the benchmark’s annotation; the smaller categories hold 34 to 73 questions.
All ↑ after r searches
Memory
r=1
r=2
r=3
r=4
r=5
MAGIC-Video-Qwen3.5-35B
0.171
0.251
0.296
0.331
0.354
MAGIC-Video-Qwen3.5-35B, 6 units/round
0.235
0.326
0.386
0.420
0.449
WorldMM-Qwen3.5-35B
0.294
0.363
0.436
0.489
0.526
GEB-Qwen3.5-35B (ours)
0.287
0.387
0.455
0.497
0.519
Ours with one decision removed
Appendix
Table 13: Evidence reached after each search round on MultiHop-EgoQA : the fraction of questions whose every evidence interval had been retrieved after the controller’s first r searches (mean over three runs, read off the controller’s search log as All in Table 3 is read off the answering context). A question that stops early keeps its final value.
Benchmark ( n )
Strongest baseline
GEB
Gap
95% CI
EgoLifeQA (500)
MAGIC-Video-Qwen3.5-35B ∗ (67.6)
72.0
+4.4
[+0.4, +8.4]
Ego-R1-Bench (50 × 3)
MAGIC-Video-Qwen3.5-35B ∗ (64.7 ± 1.15)
71.3 ± 3.1
+6.6
[+2.9, +10.3]
Test@Week (200 × 3)
WorldMM-Qwen3.5-35B (31.42)
36.83 ± 0.88
+5.41
[+0.16, +10.92]
Test@Day (200 × 3)
ReMA (GPT-5 agent) ∗ (16.75)
17.58 ± 0.14
+0.83
[ − 3.42, +5.58]
Appendix
Table 14: 95% confidence intervals on the headline gains of GEB over the strongest baseline of each column of Table 1 . EgoLifeQA and MM-Lifelong: one-sample bootstrap (2,000 resamples) of GEB’s per-question scores, the anchor subtracted as a constant; on MM-Lifelong a question’s score is its mean over our three runs. Where GEB has three runs its cell is the mean ± standard deviation over them; on Ego-R1 the interval adds the anchor’s reported per-seed variance to ours in quadrature. ∗ : the baseline’s published number; without a marker, its mean over three runs under Section 4.1 .
Memory
Acc. ↑
Δ
GEB (full)
72.0
–
w/o same-instance edges
69.0
−3.0
w/o same-instance and observation → timeline edges
67.0
−5.0
Appendix
Table 15: Removing the relevance propagation along a biography costs three points, and removing both edge families five. Further ablations on EgoLifeQA (accuracy, %), under the setting of Table 4 . w/o same-instance edges keeps the entity grouping and removes only the propagation of relevance between an entity’s observations; the combined row removes both edge families of Section 3.2 .
EgoLifeQA ↑
Ego-R1 ↑
Model
# Frames
Modality
EL
ER
HI
RM
TM
Avg.
Manual
Gemini
Avg.
125
126
61
125
63
500
25
25
50
General MLLMs
Qwen3.5-9B ∗
64
V
32.8
30.2
45.9
28.8
20.6
31.2
24.0
42.7
33.3
512
T
25.6
28.6
47.5
35.2
41.3
33.4
33.3
46.7
40.0
64
V+T
28.8
31.7
45.9
33.6
36.5
33.8
22.7
64.0
43.3
Appendix
Table 16: Full comparison on EgoLifeQA and Ego-R1-Bench (accuracy, %), with the frame budget and modality of every system. A mark on a model name gives the paper its numbers are taken from, each with its own answer model: ∗ Li et al. (2026a) , † Yeo et al. (2026) , which ran the marked systems itself, ‡ Rege et al. (2026) , § Yang et al. (2025) . Unmarked rows are our runs, with frame budgets and input modalities shown in the table. Ego-R1-Bench results are averaged over three runs (Appendix C ). Numbers reported on other question sets are omitted. Bold / underline : best/second-best in each column. Families as in Table 1 .
Figure 4: Grounded biographies connect object identity to event context. Two blue caps appear together at 15:09 on Day 3, so they are distinct instances despite sharing a description. GEB keeps a biography for each cap and retrieves the surrounding episodes to identify its wearer: Lucia and Tasha (option C). Cards summarize the retrieved biography and episodic evidence; colored boxes mark the two caps.