Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
Organizations: University of Illinois Urbana-Champaign · Amazon.com, Inc.
Abstract
Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.
Figures & tables
| EgoLifeQA | Ego-R1 | MM-Lifelong | |||||||||
| Model | EL | ER | HI | RM | TM | Avg. | Manual | Gemini | Avg. | Week | Day |
| General MLLMs | |||||||||||
| GPT-5 † | 47.2 | 42.1 | 47.5 | 53.6 | 55.6 | 48.6 | – | – | – | 15.00 ∗ | 15.25 ∗ |
| Qwen3-VL-235B | 44.0 | 44.4 | 50.8 | 44.8 | 50.8 | 46.0 | 40.0 | 62.7 | 51.3 | 15.63 ∗ | 12.44 ∗ |
| Qwen3.5-35B | 43.2 | 46.8 | 49.2 | 50.4 | 61.9 | 49.0 | 38.7 | 65.3 | 52.0 | 13.75 | 7.50 |
| Long Video MLLMs | |||||||||||
| Temporal grounding | Answering | ||||||
| Model | # Frames | mIoP | mIoG | IoU@0.3 | mIoU | Sim. | Score |
| Human ‡ | Full | 71.8 | 81.0 | 87.0 | 61.8 | 74.3 | – |
| Whole-clip MLLMs | |||||||
| GPT-4o ‡ | 60 | 18.9 | 24.4 | 12.0 | 12.2 | 73.7 | – |
| InternVL2-8B ‡ | 30 | 11.8 | 24.0 | 6.3 | 6.6 | 71.9 | 3.33 |
| LLaVA-NeXT-Video-7B ‡ | 32 | – | – | – | – | 62.1 | 2.91 |
| Evidence reached | Read | Outcome | vs. ours | |||
| Memory | All | All, 3 | Clip | Score | mIoU | All [ CI] |
| MAGIC-Video-Qwen3.5-35B | 0.354 | 0.116 | 0.230 | 2.58 0.07 | 17.6 | |
| 6 units/round | 0.449 | 0.182 | 0.319 | 2.75 0.05 | 18.5 | |
| WorldMM-Qwen3.5-35B | 0.526 | 0.244 | 0.399 | 2.74 0.06 | 18.9 | |
| GEB-Qwen3.5-35B (ours) | 0.521 | 0.235 | 0.343 | 3.11 0.03 | 20.5 | |
| Ours with one decision removed | ||||||
| Memory | Acc. | Memory | Acc. | ||
|---|---|---|---|---|---|
| GEB (full) | 72.0 | – | |||
| The index | The reading | ||||
| w/o association | 68.6 | w/o biography text (index only) | 68.0 | ||
| identity keyed by name | 69.2 | w/o biography text for the answer model | 70.8 | ||
| descriptions appended to captions | 68.2 | w/o identity notes | 68.2 | ||
| w/o observation timeline edges | 68.2 | w/o unsearched-observation line | 70.2 | ||
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Count |
|---|---|
| Thirty-second video clips | 6,266 |
| Approximate video duration | 52 hours |
| Tracked object observations | 308,244 |
| Entities after association | 77,716 |
| of which join two or more observations | 27,446 |
| of which span more than one day | 15,365 |
| WorldMM- Qwen3.5-35B | MAGIC-Video- Qwen3.5-35B | GEB-Qwen3.5-35B (ours) | |
| Nodes | |||
| Episode nodes (captions at 4 granularities) | 7,624 | 7,625 | 7,625 |
| Visual clip nodes | 6,223 † | 6,223 | 6,223 |
| Named-entity nodes extracted from text | 55,468 | 2,669 | – |
| Semantic triple nodes | 3,821 † | 3,821 | – |
| Observation nodes | – | – | 308,244 |
| Input | Test@Week | Test@Day |
|---|---|---|
| 1536 frames, longer side 512 pixels (the Qwen3.5-35B row of Table 1 ) | 13.75 | 7.50 |
| 1536 frames at the processor’s default budget, 160 96 per frame | 9.25 | 6.50 |
| 256 frames as images at 151,200 pixels each, greedy | 12.50 | 7.25 |
| 64 frames, longest edge 512 | 11.50 | 7.00 |
| 64 frames and their 64 captions | 9.00 | 4.75 |
| 512 captions, no frames | 9.00 | 3.25 |
| Memory | Tok | Fr |
|---|---|---|
| GEB-Qwen3.5-35B (ours) | 130.0k | 62.9 |
| w/o visual frames | 8.2k | 0.0 |
| MM-Lifelong | ||||
|---|---|---|---|---|
| Memory | Week | Day | ||
| GEB-Qwen3.5-35B (ours) | 36.83 0.88 | – | 17.58 0.14 | – |
| w/o biography text (index only) | 30.08 1.66 | 15.25 1.15 | ||
| descriptions appended to captions | 33.58 0.80 | 12.50 1.09 | ||
| MAGIC-Video-Qwen3.5-35B | 30.92 0.80 | 10.08 0.52 | ||
| WorldMM-Qwen3.5-35B | 31.42 2.27 | 8.50 2.63 | ||
| Model | # Frames | Test@Week | Test@Day |
|---|---|---|---|
| Human ∗ | Full | 95.6 | 99.2 |
| General MLLMs | |||
| GPT-5 ∗ | 50 | 15.00 | 15.25 |
| Qwen3-VL-235B ∗ | 1536 | 15.63 | 12.44 |
| Qwen3-VL-30B ∗ | 1536 | 11.07 | 11.48 |
| Qwen3.5-35B | 1536 | 13.75 | 7.50 |
| Temporal grounding | Answering | |||||
| Memory | mIoP | mIoG | IoU@0.3 | mIoU | Sim. | Score |
| MAGIC-Video-Qwen3.5-35B | 29.6 | 30.7 | 19.0 | 17.6 | 61.3 | 2.58 0.07 |
| MAGIC-Video-Qwen3.5-35B, 6 units/round | 30.2 | 33.2 | 20.7 | 18.5 | 63.1 | 2.75 0.05 |
| WorldMM-Qwen3.5-35B | 30.6 | 34.2 | 21.5 | 18.9 | 62.2 | 2.74 0.06 |
| GEB-Qwen3.5-35B (ours) | 32.0 | 36.7 | 25.9 | 20.5 | 65.2 | 3.11 0.03 |
| Ours with one decision removed | ||||||
| Score | ||||||
|---|---|---|---|---|---|---|
| Category | MAGIC-Video- Qwen3.5-35B | MAGIC-Video- Qwen3.5-35B, 6 units/round | WorldMM- Qwen3.5-35B | GEB-Qwen3.5-35B (ours) | Qwen3.5-35B, all 60 frames | |
| A: Repeated activities | 73 | 2.91 | 3.00 | 2.97 | 3.24 | 4.10 |
| B: Multiple actions | 241 | 2.53 | 2.73 | 2.75 | 2.91 | 3.22 |
| C: Multiple objects | 193 | 2.40 | 2.51 | 2.43 | 2.90 | 3.47 |
| D: Locations/people | 72 | 2.51 | 2.79 | 2.59 | 3.25 | 3.85 |
| E: Event composition | 111 | 2.32 | 2.46 | 2.57 | 2.90 | 3.34 |
| All after searches | |||||
| Memory | |||||
| MAGIC-Video-Qwen3.5-35B | 0.171 | 0.251 | 0.296 | 0.331 | 0.354 |
| MAGIC-Video-Qwen3.5-35B, 6 units/round | 0.235 | 0.326 | 0.386 | 0.420 | 0.449 |
| WorldMM-Qwen3.5-35B | 0.294 | 0.363 | 0.436 | 0.489 | 0.526 |
| GEB-Qwen3.5-35B (ours) | 0.287 | 0.387 | 0.455 | 0.497 | 0.519 |
| Ours with one decision removed | |||||
| Benchmark ( ) | Strongest baseline | GEB | Gap | 95% CI |
|---|---|---|---|---|
| EgoLifeQA (500) | MAGIC-Video-Qwen3.5-35B ∗ (67.6) | 72.0 | +4.4 | [+0.4, +8.4] |
| Ego-R1-Bench (50 3) | MAGIC-Video-Qwen3.5-35B ∗ (64.7 1.15) | 71.3 3.1 | +6.6 | [+2.9, +10.3] |
| Test@Week (200 3) | WorldMM-Qwen3.5-35B (31.42) | 36.83 0.88 | +5.41 | [+0.16, +10.92] |
| Test@Day (200 3) | ReMA (GPT-5 agent) ∗ (16.75) | 17.58 0.14 | +0.83 | [ 3.42, +5.58] |
| Memory | Acc. | |
|---|---|---|
| GEB (full) | 72.0 | – |
| w/o same-instance edges | 69.0 | |
| w/o same-instance and observation timeline edges | 67.0 |
| EgoLifeQA | Ego-R1 | ||||||||||
| Model | # Frames | Modality | EL | ER | HI | RM | TM | Avg. | Manual | Gemini | Avg. |
| 125 | 126 | 61 | 125 | 63 | 500 | 25 | 25 | 50 | |||
| General MLLMs | |||||||||||
| Qwen3.5-9B ∗ | 64 | V | 32.8 | 30.2 | 45.9 | 28.8 | 20.6 | 31.2 | 24.0 | 42.7 | 33.3 |
| 512 | T | 25.6 | 28.6 | 47.5 | 35.2 | 41.3 | 33.4 | 33.3 | 46.7 | 40.0 | |
| 64 | V+T | 28.8 | 31.7 | 45.9 | 33.6 | 36.5 | 33.8 | 22.7 | 64.0 | 43.3 | |