Learning to Retrieve Missing Evidence for Long-Term Memory QA
Organizations: Tsinghua University
Abstract
Long-term memory enables language models to use past interactions in future conversations. However, evidence needed to answer a question may be scattered across distant turns, while the question itself omits clues needed to locate it. Retrieved facts can reveal these clues, motivating retrieval decisions conditioned on evidence already found. We introduce MERA (Missing-Evidence Retrieval Augmentation), which separates globally searchable memory from a question-specific evidence state. Verified evidence guides subsequent retrieval without restricting access to the global memory. We train a lightweight planner through reinforcement learning, rewarding queries that recover previously missing evidence. MERA achieves strong answer accuracy across Qwen3-30B and GPT-4o-mini backbones. With Qwen3-30B for evidence processing and answer generation, the trained 0.6B planner achieves 77.40% accuracy on LoCoMo and 71.29% on LongMemEval-S, exceeding a 30B planner without retrieval-grounded training by 4.10% and 3.96%, respectively. On LoCoMo, later retrieval rounds increase cumulative evidence recall from 55.5% to 80.5%.
Figures & tables
| Backbone | Method | ACC (%) | Sum. In (k) | Sum. Out (k) | Upd. In (k) | Upd. Out (k) | Total (k) |
|---|---|---|---|---|---|---|---|
| LongMemEval-S | |||||||
| Qwen3-30B | FullText | 54.80 | – | – | – | – | 105.07 |
| NaiveRAG | 60.80 | – | – | – | – | – | |
| IRCoT ∗ | 55.45 | – | – | – | – | – | |
| GraphReader ∗ | 46.53 | 155.52 | 99.61 | – | – | 255.13 | |
| LangMem | 50.80 | – | – | 1,311.96 | 118.06 | 1,430.02 | |
| Planner | LoCoMo | LongMemEval-S |
|---|---|---|
| MERA-Small | 68.40 | 59.80 |
| MERA-Large | 73.30 | 67.33 |
| MERA-RL | 77.40 | 71.29 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Histories | Generated QAs | Collected states | Rollout input | Retained |
|---|---|---|---|---|---|
| LoCoMo | 170 | 24,787 | 74,742 | 74,742 | 2,996 (4.0%) |
| LongMemEval | 277 | 11,017 | 17,752 | 14,475 | 2,312 (16.0%) |
| Method | In (k) | Out (k) | Total (k) |
|---|---|---|---|
| IRCoT | 4.29 | 0.13 | 4.43 |
| GraphReader | 24.50 | 0.54 | 25.04 |
| MERA | 13.47 | 1.80 | 15.27 |
| Method | In (k) | Out (k) | Total (k) |
|---|---|---|---|
| IRCoT | 19.02 | 0.24 | 19.26 |
| GraphReader | 32.82 | 0.52 | 33.34 |
| MERA | 14.97 | 2.43 | 17.40 |
| Configuration | LLM calls/Q | Runtime/Q (s) | Relative runtime |
|---|---|---|---|
| MERA-Small | 6.8 | 43.2 | 1.00 |
| MERA-RL | 6.6 | 48.3 | 1.12 |
| MERA-Large | 10.9 | 155.5 | 3.60 |
| Metric | MERA-Small | MERA-RL | Change |
|---|---|---|---|
| Average query word count | 3.30 | 3.50 | +0.20 |
| Intents per planner call | 2.75 | 2.98 | +0.23 |
| Normalized query repetition | 17.1% | 16.0% | pp |
| JSON parsing failures | 0.0% | 0.0% | 0.0 pp |
| First-round evidence recall | 0.623 | 0.631 | +0.8 pp |
| Final evidence recall | 0.761 | 0.776 | +1.5 pp |