Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression. However, when future information needs are unknown, compressing an entire interaction in one pass can overlook locally important details that may matter later. To this end, we introduce RIME, a retrieval-induced memory framework that shifts memory construction from monolithic compression toward evidence-centered integration. RIME uses generic self-questions to retrieve focused dialogue evidence and grounds memory formation in both the retrieved evidence and relevant historical memories, which are jointly reconciled into an evolving memory bank with temporal and provenance information. At inference time, compressed memory serves as the primary rather than the sole source of evidence: when it cannot support an answer, RIME retrieves relevant source dialogue together with its local context to recover information omitted during memory formation, without resorting to full-history processing. Extensive experiments on LoCoMo with Qwen3-235B-A22B and GPT-5.6 Sol show that RIME consistently achieves the best performance across all three quality metrics among the compared methods, while requiring substantially fewer query-time LLM tokens.
Figures & tables
Figure 1: Overview of RIME. (a) During memory formation, generic formation questions are used as retrieval cues to identify focused dialogue evidence, which is expanded with local context and combined with relevant historical memories for joint memory formation and reconciliation. Unlike conventional direct compression, RIME retrieves relevant evidence before consolidation. (b) During memory use, RIME first answers from retrieved semantic memories and selectively returns to the source dialogue only when the memory-based answer is insufficient.
Method
Judge ↑
F1 ↑
BLEU-1 ↑
Tokens ↓
Qwen3-235B-A22B
RAG
31.36
20.87
17.46
3.64K
Mem0
53.12
35.76
30.49
2.20K
A-MEM
68.77
44.12
37.86
63.19K
Nemori
77.34
46.01
39.88
20.92K
D-Mem †
78.60
51.00
42.60
15.57K
Table 1: Main results on the 1,540 non-adversarial LoCoMo questions. Higher is better for Judge, F1, and BLEU-1; lower is better for inference tokens. † D-Mem results are directly reported from the original paper using Qwen3-235B-Instruct, as its implementation is not publicly available.
Formation Strategy
Judge ↑
F1 ↑
BLEU-1 ↑
Direct Compression
71.62
49.26
42.32
Question-Prompted Compression
72.66
48.87
42.11
RIME
81.82
51.46
44.22
Table 2: Effect of memory formation strategies on LoCoMo using Qwen3-235B-A22B ( n=1,540 ). All variants use the same selective memory-use procedure.
Method
Judge ↑
F1 ↑
BLEU-1 ↑
Qwen3-235B-A22B
RIME w/o selective retrieval
71.23
44.73
38.41
RIME
81.82
51.46
44.22
GPT-5.6 Sol
RIME w/o selective retrieval
80.97
54.27
46.96
RIME
84.29
56.41
48.85
Table 3: Ablation of selective source retrieval on LoCoMo.
Long-term memory enables LLM agents to leverage past interactions, but dialogue histories quickly exceed the context window, forcing agents to retrieve relevant subsets at query time. Because useful evidence is sparse and scattered across verbose conversations, retrieval faces a fundamental tension: broadening recall improves coverage but floods downstream reasoning with noise, while compressing memories at write time eases retrieval but irreversibly discards details that future queries may need. We introduce LazyMem, which resolves this tension by deferring all memory construction to query time. Given a retrieved candidate pool, a lightweight model processes it in overlapping parallel windows, selectively retaining and compressing only query-relevant content. The model is trained with supervised fine-tuning followed by reinforcement learning, using a reward that jointly encourages the identification of relevant messages and the generation of compressions that are faithful to the source and useful for answering the query. On LongMemEval, LazyMem-4B achieves an LLM-judge accuracy of 0.85, outperforming the strongest non-oracle baseline while using only 213 answer-context memory tokens, 21.0 times fewer than the baseline. It further generalizes to LoCoMo without target-domain training and reduces mean latency relative to the prior query-time baseline. Code is available at https://github.com/allacnobug/LazyMem.
Jing Yu, Yibo Zhao, Jiaming Zhang +1
School of Data Science and Engineering, East China Normal University
Long-term memory is essential for LLM-based agents to sustain interactions and reliably leverage distant history. However, existing memory systems typically process heterogeneous dialogue content through a uniform summarization and retrieval pipeline, leading to either excessive token consumption or irreversible loss of fine-grained evidence. We argue that historical dialogue content should be handled differently according to its compressibility, temporal dynamics, and fidelity requirements. Based on this insight, we propose LeanMem, a lightweight long-term memory framework. LeanMem first filters out low-value content, then stores informative segments as compact profile memory, temporally structured event memory, or source-grounded record memory, depending on the nature of the information. During maintenance, only dynamically evolving event memories are selectively updated, avoiding redundant consolidation of stable profiles and immutable records. During inference, LeanMem dynamically selects memory types and allocates retrieval budgets according to query-specific evidence demands, assembling relevant evidence on demand. On LoCoMo and LongMemEval-S with GPT-4.1-mini and Qwen3-8B, LeanMem improves accuracy over the strongest memory-based baseline in every setting, by up to 15.1 points, at the lowest or near-lowest construction cost, inference tokens, and latency. The code and datasets are included in the supplementary materials.
To enable reliable long-term interaction, LLM agents require a memory system that can faithfully store, efficiently retrieve, and deeply reason over accumulated dialogue history. Most existing methods adopt an extracted fact based paradigm: handcrafted static prompts compress raw dialogues into atomic facts, which are then stored, matched, and injected into downstream reasoning. Nevertheless, such fact-centric designs inevitably discard fine-grained details in original dialogues and fail to support deep reasoning over scattered isolated facts. Moreover, static prompts cannot maintain consistent extraction granularity across diverse dialogue styles. To address these limitations, we propose TriMem, which maintains three coexisting representation granularities, including raw dialogue segments anchored by source identifiers for storage fidelity, extracted atomic facts for efficient memory retrieval, synthesized profiles that aggregate dispersed facts into holistic semantic understanding for deep reasoning. We further adopt TextGrad-based prompt optimization, which iteratively refines extraction and profiling prompts via response quality feedback, achieving lifelong evolution without any parameter updating. Extensive experiments on LoCoMo and PerLTQA across multiple LLM backbones demonstrate that TriMem consistently outperforms strong memory baselines. The code is available at https://TMLR-TriMem.github.io .
Jingwei Sun, Jianing Zhu, Jiangchao Yao +2
TMLR Group, Hong Kong Baptist University · The University of Texas at Austin · Shanghai Jiao Tong University +1