Madeleine: Learning Involuntary Recall for Conversational Memory from Simulated Lives
Organizations: Nanyang Technological University, Singapore
Abstract
A long-term conversational assistant must recall the right memory at the right moment, yet the memory that matters most is often not similar to what the user says now. Current systems recover such associations by letting an LLM reason at write or read time, at a cost of hundreds to over a thousand LLM calls per memory bank and up to several thousand context tokens per query. We argue that association is a learnable relevance: the pointwise mutual information of memories under how human lives unfold. We introduce Madeleine, which learns amortized association: offline, an LLM life simulator writes simulated lives, whose cue-trigger pairs teach a query encoder a residual association on top of frozen similarity; online, it calls no LLM and plugs into any vector memory by replacing only the query encoder. On LoCoMo-Plus under the official protocol, Madeleine (I) reaches 66.6 when plugged into HyperMem, the highest among all systems evaluated under this protocol; (II) used alone, reaches the score of HyperMem as released (52.4 vs. 52.9) with zero LLM calls and about 1/21 of its answer context; and (III) lifts T-Mem by 26.2 points, significantly outperforms the same untrained backbone inside both systems, and leaves ordinary QA intact on the 4B backbone.
Figures & tables
| Group | Method | Write | Read | Tokens | Causal | State | Goal | Value | All |
| gemini-embedding | 0 | 0 | 387 | 55.4 | 52.0 | 26.0 | 33.0 | 41.6 | |
| text-embedding-3-large | 0 | 0 | 396 | 58.4 | 45.0 | 30.0 | 31.0 | 41.1 | |
| Qwen3-Embedding-4B | 0 | 0 | 378 | 54.5 | 37.0 | 29.0 | 38.0 | 39.7 | |
| Qwen3-Embedding-8B | 0 | 0 | 386 | 50.5 | 36.0 | 28.0 | 35.0 | 37.4 | |
| (a) Embedding models | bge-m3 | 0 | 0 | 385 | 31.7 | 28.0 | 9.0 | 19.0 | 21.9 |
| (b) Full context | Full dialogue | 0 | 0 | 18,776 | 49.5 | 49.0 | 41.0 | 40.0 | 44.9 |
| System | Original | Untrained | Madeleine |
|---|---|---|---|
| Standalone | – | 39.7 | 52.4 {}^{**}_{{\color[rgb]{0.8516,0.2813,0.1094}\uparrow 12.7}} |
| T-Mem, 3 scenes | 33.7 | 53.9 | 59.9 {}^{*}_{{\color[rgb]{0.8516,0.2813,0.1094}\uparrow 6.0}} |
| HyperMem | 52.9 | 60.6 | 66.6 {}^{*}_{{\color[rgb]{0.8516,0.2813,0.1094}\uparrow 6.0}} |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | Generator | Calls | Cost ($) | Statements | Pair cos | Rows kept (removed) |
|---|---|---|---|---|---|---|
| v2 (static facts) | gpt-4.1-mini | 800 | 1.61 | 8,000 | 0.422 | 31,697 ( 1%) |
| v3 (trajectories) | DeepSeek-V4.1-Flash | 800 | 0.79 | 7,975 | 0.398 | 32,419 (49) |
| v2d (v2 prompt) | DeepSeek-V4.1-Flash | 800 | 0.73 | 8,136 | – | 31,697 (1,092) |
| Item | Value |
|---|---|
| Backbone | Qwen3-Embedding-4B (also 0.6B, 8B) ( Zhang et al., 2025 ) |
| Adapter | LoRA ( Hu et al., 2022 ) , , , dropout 0.05 |
| LoRA targets | q, k, v, o, gate, up, down projections |
| Trainable parameters (4B) | 66.1M |
| Loss | InfoNCE on the summed score, |
| Negatives | all same-source training facts |
| Retriever | =3 | =5 | =10 |
| Madeleine (4B), seed 0 | 52.4 | 53.1 | 54.1 |
| Madeleine (4B), seed 1 | 51.6 | 53.9 | 51.9 |
| Madeleine (4B), seed 2 | 53.1 | 54.4 | 53.9 |
| Madeleine (4B), mean | 52.4 | 53.8 | 53.3 |
| Untrained Qwen3-Embedding-4B | 39.7 | 47.4 | 49.9 |
| gemini-embedding | 41.6 | 45.9 | 46.4 |
| System | Query encoder | Score |
| HyperMem | original | 52.9 |
| HyperMem | untrained 4B (with instruction) | 60.6 |
| HyperMem | Madeleine | 66.6 |
| T-Mem, 3 scenes | original (bge-m3) | 33.7 |
| T-Mem, 3 scenes | untrained 4B | 53.9 |
| T-Mem, 3 scenes | Madeleine | 59.9 |
| Retriever in T-Mem | All | Causal | Goal | State | Value | Cue@10 | 3 scenes (Cue@3) |
|---|---|---|---|---|---|---|---|
| bge-m3 (original) | 51.9 | 64.4 | 44.0 | 59.0 | 40.0 | 70.1 | 41.4 (40.4) |
| Untrained Qwen3-Embedding-4B | 67.6 | 78.2 | 62.0 | 67.0 | 63.0 | 89.3 | 64.1 (72.6) |
| Madeleine | 70.6 | 78.2 | 67.0 | 72.0 | 65.0 | 94.5 | 72.6 ( 84.3 ) |
| Benchmark | Metric | Untrained 4B | Madeleine | text-emb-3-large |
|---|---|---|---|---|
| LoCoMo-Conv (1,540) | Recall@10 (paper metric) | 46.2 | 48.1 | 49.6 |
| LoCoMo-Conv | answer score, =10 | 0.440 | 0.456 | 0.453 |
| LoCoMo-Conv, user turns only | answer score, =10 | 0.456 | 0.482 | 0.489 |
| InMind single fact (125) | indirect query | 5.6 | 27.7 | 0.8 |
| InMind single fact | direct query | 96.8 | 93.1 | 80.8 |
| Retriever | Causal | State | Goal | Value | All |
| Untrained 0.6B | 50.5 | 41.0 | 33.0 | 19.0 | 35.9 |
| 0.6B, v2 only | 72.3 ∗ | 61.0 ∗ | 53.0 ∗ | 51.0 ∗ | 59.4 ∗ |
| Untrained 4B | 89.1 | 65.0 | 67.0 | 63.0 | 71.1 |
| 4B, v2 only | 90.1 | 82.0 ∗ | 72.0 | 71.0 ∗ | 78.8 ∗ |
| 4B, full recipe | 90.1 | 85.0 ∗ | 79.0 ∗ | 77.0 ∗ | 82.8 ∗ |
| Variant | LP | LoCoMo | LoCoMo QA |
|---|---|---|---|
| (a) Which side to train, and which score to train on [0.6B, v2 data] | |||
| Untrained 0.6B | 35.9 | 67.3 | – |
| Both sides, full fine-tuning | 49.6 | 55.1 | – |
| Both sides, LoRA | 50.1 | 56.1 | – |
| Query side, association term trained alone | 54.1 | 64.5 | – |
| Query side, residual (summed) score | 59.4 | 60.8 | – |
| Method | Read | Score | ||
| Madeleine (seed 0) | 0 | 52.4 | – | – |
| Untrained Qwen3-Embedding-4B | 0 | 39.7 | 12.7 | |
| (i) Read-time LLM rewriting or guessing | ||||
| QueryLink, original config. | 2 | 23.7 | 28.7 | |
| QueryLink, native union | 2 | 35.9 | 16.5 | |
| QueryLink, rewriting + Qwen3-4B | 2 | 34.4 | 18.0 | |