Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression. However, when future information needs are unknown, compressing an entire interaction in one pass can overlook locally important details that may matter later. To this end, we introduce RIME, a retrieval-induced memory framework that shifts memory construction from monolithic compression toward evidence-centered integration. RIME uses generic self-questions to retrieve focused dialogue evidence and grounds memory formation in both the retrieved evidence and relevant historical memories, which are jointly reconciled into an evolving memory bank with temporal and provenance information. At inference time, compressed memory serves as the primary rather than the sole source of evidence: when it cannot support an answer, RIME retrieves relevant source dialogue together with its local context to recover information omitted during memory formation, without resorting to full-history processing. Extensive experiments on LoCoMo with Qwen3-235B-A22B and GPT-5.6 Sol show that RIME consistently achieves the best performance across all three quality metrics among the compared methods, while requiring substantially fewer query-time LLM tokens.
Figures & tables
Figure 1: Overview of RIME. (a) During memory formation, generic formation questions are used as retrieval cues to identify focused dialogue evidence, which is expanded with local context and combined with relevant historical memories for joint memory formation and reconciliation. Unlike conventional direct compression, RIME retrieves relevant evidence before consolidation. (b) During memory use, RIME first answers from retrieved semantic memories and selectively returns to the source dialogue only when the memory-based answer is insufficient.
Method
Judge ↑
F1 ↑
BLEU-1 ↑
Tokens ↓
Qwen3-235B-A22B
RAG
31.36
20.87
17.46
3.64K
Mem0
53.12
35.76
30.49
2.20K
A-MEM
68.77
44.12
37.86
63.19K
Nemori
77.34
46.01
39.88
20.92K
D-Mem †
78.60
51.00
42.60
15.57K
Table 1: Main results on the 1,540 non-adversarial LoCoMo questions. Higher is better for Judge, F1, and BLEU-1; lower is better for inference tokens. † D-Mem results are directly reported from the original paper using Qwen3-235B-Instruct, as its implementation is not publicly available.
Formation Strategy
Judge ↑
F1 ↑
BLEU-1 ↑
Direct Compression
71.62
49.26
42.32
Question-Prompted Compression
72.66
48.87
42.11
RIME
81.82
51.46
44.22
Table 2: Effect of memory formation strategies on LoCoMo using Qwen3-235B-A22B ( n=1,540 ). All variants use the same selective memory-use procedure.
Method
Judge ↑
F1 ↑
BLEU-1 ↑
Qwen3-235B-A22B
RIME w/o selective retrieval
71.23
44.73
38.41
RIME
81.82
51.46
44.22
GPT-5.6 Sol
RIME w/o selective retrieval
80.97
54.27
46.96
RIME
84.29
56.41
48.85
Table 3: Ablation of selective source retrieval on LoCoMo.