MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
Organizations: Meta Reality Labs · University of Virginia · Work done at Meta · University of Illinois Urbana-Champaign
Abstract
Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever fails to locate relevant entries due to retrieval competition in growing search spaces. To address these challenges, we introduce MemLife, a multimodal memory system that constructs entity-grounded, first-person text episodes and retrieves them via a time-indexed agentic reader. Without training or query-time video access, MemLife improves over the strongest training-free baseline by 4.6--12.0% across four long-horizon benchmarks. To further improve memory quality, we propose MemOpt, a reinforcement learning framework that optimizes the memory writer to produce faithful, informative, and retrievable memories. MemOpt consistently improves MemLife by 2.7--5.0% across different video and question distributions, with gains that generalize across writer and reader backbones and memory systems.
Figures & tables
| Method | SuperMemory-VQA | EgoLifeQA | SuperMemory-LVQA | EgoLife-EQA | ||||
| Accuracy | Recall | Accuracy | Recall | Accuracy | Recall | Accuracy | Recall | |
| Reference | ||||||||
| Oracle Context | 67.58 | 100.00 | 66.20 | 100.00 | 67.58 | 100.00 | 61.00 | 100.00 |
| Without Memory-Writer Training | ||||||||
| Video ReCap | 36.28 | 54.01 | 35.20 | 27.40 | 39.17 | 16.79 | 38.00 | 27.00 |
| EgoRAG | 49.28 | 66.79 | 48.20 | 29.60 | 47.51 | 47.90 | 38.00 | 13.00 |
| Frames | ||
| Question | Q: […] Did I add the milk before or after the eggs when making the batter? | Q: […] Did I set a reminder for a backup dinner plan? |
| A: You did not add milk or eggs to the batter; you only added lentils, salt, and spices. | A: No […] However, you did mention earlier that you have baked chicken available. | |
| MemLife | I am in a kitchen […] transfer yellow corn kernels from a small food processor bowl […] move toward the sink area […] | […] preparing food […] they respond to my question about being hungry by saying they can wait for the food […] |
| MemLife + MemOpt | I am in a kitchen […] yellowish-orange granular material, which appears to be cooked lentils […] to the sink area […] | […] I am preparing a meal, specifically baked chicken, and offering it to B , who indicates they are not hungry and can wait. |
| Writer | Training | Qwen3.5-9B Reader | Qwen3.6-27B Reader | ||
| SuperMemory-VQA | EgoLifeQA | SuperMemory-VQA | EgoLifeQA | ||
| Qwen3.5-9B | Zero-shot | 56.50 | 52.80 | 66.93 | 59.60 |
| MemOpt | 60.35 | 56.60 | 67.90 | 60.20 | |
| Qwen3.6-27B | Zero-shot | 57.14 | 55.20 | 66.77 | 59.80 |
| MemOpt | 60.67 | 56.60 | 68.06 | 60.60 | |
| Writer Ablation | ||||||||
| Multimodal Fusion | Entity Grounding | First-person Narration | SuperMemory-VQA | EgoLifeQA | EgoLifeQA Recall | |||
| Zero-shot | MemOpt | Zero-shot | MemOpt | Zero-shot | MemOpt | |||
| ✔ | ✗ | ✗ | 52.33 | 56.98 | 52.80 | 52.60 | 48.60 | 47.00 |
| ✔ | ✔ | ✗ | 54.09 | 59.87 | 52.00 | 54.40 | 45.60 | 47.60 |
| ✔ | ✔ | ✔ | 56.50 | 60.35 | 52.80 | 56.60 | 48.40 | 50.60 |
| Reader Ablation | ||||||||
| Writer Supervision | SuperMemory-VQA | EgoLifeQA | ||||||
| Teacher | Accuracy | Retrievability | Informativeness | Faithfulness | Accuracy | Recall | Accuracy | Recall |
| ✗ | ✗ | ✗ | ✗ | ✗ | 56.50 | 80.53 | 52.80 | 48.40 |
| ✔ | ✗ | ✗ | ✗ | ✗ | 57.78 | 83.59 | 53.00 | 49.00 |
| ✗ | ✔ | ✗ | ✗ | ✗ | 57.46 | 79.01 | 52.60 | 47.80 |
| ✗ | ✗ | ✔ | ✗ | ✗ | 53.45 | 82.63 | 54.40 | 53.20 |
| ✗ | ✗ | ✔ | ✔ | ✗ | 60.03 | 84.35 | 51.40 | 47.60 |
| Training Design | SuperMemory-VQA | EgoLifeQA | |||
| Faithfulness | Aggregation | Accuracy | Recall | Accuracy | Recall |
| Sequence-level | Additive | 58.75 | 83.78 | 55.00 | 46.40 |
| Token-level | Additive | 58.91 | 81.68 | 56.00 | 49.40 |
| Token-level | Multiplicative | 60.35 | 81.49 | 56.60 | 50.60 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Type | # Questions | # Options | Avg. Correct Options | Avg. Evidence Segments | Avg. Distinct Days |
| Frequency counting | 45 | 8 | 2.33 | 2.31 | 2.31 |
| Routine | 31 | 2–4 | 1.00 | 2.97 | 2.97 |
| Comparison | 24 | 3 | 1.00 | 8.04 | 5.54 |
| All | 100 | – | 1.60 | 3.89 | 3.29 |
| Dataset | # Questions | Avg. Video Hours | Avg. Time Span (days) | Avg. Evidence Segments | Avg. Evidence Duration (s) |
| SuperMemory-VQA | 4771 | 3.97 | 12.50 | 1.34 | 57.1 |
| EgoLifeQA | 500 | 22.65 | 2.80 | 1.10 | 32.3 |
| SuperMemory-LVQA | 623 | 47.90 | 127.36 | 1.24 | 56.3 |
| EgoLife-EQA | 100 | 43.06 | 6.33 | 3.89 | 116.7 |
| Group | Parameter | Value |
| Sampling | Group size (rollouts per segment) | 5 |
| Rollout temperature / top- / top- | 1.0 / 1.0 / disabled | |
| Maximum prompt length (tokens) | 12,288 | |
| Maximum response length (tokens) | 512 | |
| Optimization | Learning rate | |
| Weight decay | 0.01 |
| Method | SuperMemory-VQA | EgoLifeQA | ||
| Accuracy | Recall | Accuracy | Recall | |
| VST | 56.50 | 65.08 | 29.40 | 7.60 |
| TaskMem | 48.96 | 51.15 | 48.20 | 35.00 |
| M3-Agent | 53.93 | 54.20 | 35.40 | 9.60 |
| MemLife + MemOpt | 60.35 | 81.49 | 56.60 | 50.60 |
| Method | SuperMemory-VQA | EgoLifeQA | ||||||
| Accuracy | Recall | Accuracy | Recall | |||||
| EgoRAG | 49.28 | 49.18 0.95 | 66.79 | 66.79 0.00 | 48.20 | 46.48 0.95 | 29.60 | 28.80 0.00 |
| MemLife | 56.50 | 55.31 1.53 | 80.53 | 80.69 0.32 | 52.80 | 50.16 1.83 | 48.40 | 47.48 1.36 |
| MemLife -V | 57.78 | 58.78 0.41 | 84.35 | 84.58 0.64 | 53.20 | 52.44 1.72 | 50.20 | 49.72 0.30 |
| Method | SuperMemory-LVQA | EgoLife-EQA | ||||||
| Method | Memory Storage (MB/h) | Reading Speed (s/question) | ||
| SuperMemory-VQA | EgoLifeQA | SuperMemory-VQA | EgoLifeQA | |
| Video ReCap | 3.86 | 3.66 | 20.4 | 54.2 |
| EgoRAG | 4.98 | 5.49 | 33.2 | 43.5 |
| Video-RAG | 1.03 | 3.44 | 60.9 | 76.9 |
| VideoARM | – | – | 47.5 | 110.4 |
| EGAgent | 21.50 | 22.20 | 200.3 | 230.7 |
| Training-time proxy | SuperMemory-VQA | EgoLifeQA |
| None (zero-shot) | 56.50 | 52.80 |
| Original question | 57.62 | 55.40 |
| Agent actions | 60.35 | 56.60 |
| Writer trained | Reader trained | SuperMemory-VQA | EgoLifeQA | ||
| Accuracy | Recall | Accuracy | Recall | ||
| ✗ | ✗ | 56.50 | 80.53 | 52.80 | 48.40 |
| ✔ | ✗ | 60.35 | 81.49 | 56.60 | 50.60 |
| ✗ | ✔ | 61.00 | 89.50 | 51.40 | 59.00 |
| ✔ | ✔ | 63.08 | 87.98 | 54.20 | 56.80 |
| Predecessor Context | Conversational Memory | Intent Recall | In-Context Retrieval | Timeline Reconstruction | Object-Location Memory | Visual Recall | Overall |
| None | 63.41 | 77.42 | 43.18 | 57.41 | 52.29 | 44.12 | 56.50 |
| Text | 62.60 | 86.02 | 37.50 | 48.15 | 45.87 | 45.10 | 54.25 |
| Multimodal | 69.92 | 82.80 | 37.50 | 47.22 | 46.79 | 43.14 | 54.90 |