Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This reactive paradigm means every retrieval failure corresponds to a real interaction in which the cost has already been paid. We propose MemDream, a framework that enables self-probing memory evolution for LLM agents. Our framework periodically enters offline dream cycles where three specialized agents (Dreamer, Analyst, Consolidator) collaboratively probe, diagnose, and repair the memory graph before failures occur. A policy trained via Group Relative Policy Optimization learns which repair operations produce durable retrieval improvements, while a soft decay mechanism provides reversible forgetting driven by the same anticipatory signal. Experiments on LoCoMo and MemoryAgentBench demonstrate that MemDream improves answer F1 by 4.5 points on LoCoMo and achieves a 9.1-point higher overall score on MAB over the strongest reactive-evolution baselines.
Figures & tables
Figure 1: (a) Existing reactive systems store a user’s nut allergy and baking preferences as disconnected nodes, missing the dependency between them. When queried for a recipe, the system recommends a nut-containing dish, a harmful answer the user must catch. (b) Through asynchronous background consolidation, MemDream ’s Dream Cycle finds and repairs this missing link, preventing the unsafe recommendation before the query arrives.
Figure 2: Overview of MemDream . Three LoRA-specialized agents share one Llama-3.1-8B backbone: the Dreamer generates probing queries aimed at weak regions, the Analyst diagnoses failures along four axes (buried, scattered, redundant, noisy), and the Consolidator applies repairs (extract, merge, synthesize, prune). The cycle repeats until retrieval F1 converges. We first train the agents independently with GRPO, then jointly fine-tune the Dreamer and Consolidator (with the Analyst frozen) using end-to-end memory quality as the reward.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Base model
Llama-3.1-8B-Instruct
Quantization
8-bit (bitsandbytes)
LoRA rank r
16
LoRA alpha α
32
Target modules
q_proj, k_proj, v_proj, o_proj
LoRA dropout
0.05
Appendix
Table 7: LoRA adapter configuration.
Parameter
Value
Group size N
4
Learning rate
1×10−3
Exploration rate ϵ
0.2 (decaying)
Training data
MSC dialogue (5 sessions)
Dream queries per conversation
8–12
Training steps (independent)
500 per agent
Appendix
Table 8: GRPO training hyperparameters.
Rounds
F1
B-1
R-L
0 (no dream)
48.72
38.15
46.83
1
51.36
40.62
49.58
3
53.41
42.87
51.92
5
53.89
43.21
52.35
8
53.52
42.94
51.87
Appendix
Table 9: Effect of dream cycle rounds on LoCoMo.
Dreamer Variant
Disc.
Cov.
Align.
Rule-Based
0.36
5.0
0.37
LLM Dreamer
0.15
4.8
0.35
Appendix
Table 10: Dream quality comparison.
Signal Source
Mean Reward
Std
Baseline (no dream)
0.198
0.031
Dream-guided (top-50%)
0.232
0.035
Dream-guided (top-30%)
0.241
0.033
Oracle (real queries)
0.262
0.038
Appendix
Table 11: Training signal comparison .
Component
Time
GPU Mem.
Load base model (8-bit)
45s
9.2 GB
Load 3 LoRA adapters
8s
+1.8 GB
Dreamer: generate queries
2.1 min
12.4 GB
Retrieval (BM25 + dense)
0.8 min
3.2 GB
Analyst: diagnose
1.5 min
12.4 GB
Consolidator: execute
3.2 min
12.4 GB
Appendix
Table 14: Approximate offline timing and GPU memory profile. Per-stage times describe one round; the final row reports the full-cycle estimate.
Quantity
Mean per conversation
Per initial node
LLM calls
60.9
0.199
Input tokens
33,588.4
109.69
Output tokens
13,004.5
42.47
Total tokens
46,592.9
152.16
Model generation time (s)
486.41
1.59
Cycle wall time (s)
2,208.84
7.21
Appendix
Table 15: Recorded overhead for the rule-probe configuration. Per-node values divide totals by 3,062 initial nodes. Model loading, initial memory construction, and LLM Dreamer generation are excluded; these measurements use a different configuration from Table 14 .
Method
Signal or criterion
Memory update
Auto-Dreamer
Task performance and counterfactual memory utility during training
Rewrite a selected region into a compact replacement set
TrustMem
Transition-level coverage, preservation, and faithfulness
Learn trustworthy updates from preferences over candidate transitions
RecMem
Recurrence of semantically similar interactions
Trigger consolidation and refine fine-grained semantic facts
MemDream
Failures exposed by self-generated retrieval probes
Diagnose failures, apply targeted repairs, and re-evaluate retrieval
Appendix
Table 16: Signals and update mechanisms in recent memory consolidation methods.
Long-term memory is essential for LLM agents that operate across multiple sessions, yet existing memory systems treat retrieval infrastructure as fixed: stored content evolves while scoring functions, fusion strategies, and answer-generation policies remain frozen at deployment. We argue that truly adaptive memory requires co-evolution at two levels: the stored knowledge and the retrieval mechanism that queries it. We present EvolveMem, a self-evolving memory architecture that exposes its full retrieval configuration as a structured action space optimized by an LLM-powered diagnosis module. In each evolution round, the module reads per-question failure logs, identifies root causes, and proposes targeted configuration adjustments; a guarded meta-analyzer applies them with automatic revert-on-regression and explore-on-stagnation safeguards. This closed-loop self-evolution realizes an AutoResearch process: the system autonomously conducts iterative research cycles on its own architecture, replacing manual configuration tuning. Starting from a minimal baseline, the process converges autonomously, discovering effective retrieval strategies including entirely new configuration dimensions not present in the original action space. On LoCoMo, EvolveMem outperforms the strongest baseline by 25.7% relative and achieves a 78.0% relative improvement over the minimal baseline. On MemBench, EvolveMem exceeds the strongest baseline by 18.9% relative. Evolved configurations transfer across benchmarks with positive rather than catastrophic transfer, indicating that the self-evolution process captures universal retrieval principles rather than benchmark-specific heuristics. Code is available at https://github.com/aiming-lab/SimpleMem.
Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression. However, when future information needs are unknown, compressing an entire interaction in one pass can overlook locally important details that may matter later. To this end, we introduce RIME, a retrieval-induced memory framework that shifts memory construction from monolithic compression toward evidence-centered integration. RIME uses generic self-questions to retrieve focused dialogue evidence and grounds memory formation in both the retrieved evidence and relevant historical memories, which are jointly reconciled into an evolving memory bank with temporal and provenance information. At inference time, compressed memory serves as the primary rather than the sole source of evidence: when it cannot support an answer, RIME retrieves relevant source dialogue together with its local context to recover information omitted during memory formation, without resorting to full-history processing. Extensive experiments on LoCoMo with Qwen3-235B-A22B and GPT-5.6 Sol show that RIME consistently achieves the best performance across all three quality metrics among the compared methods, while requiring substantially fewer query-time LLM tokens.
Memory is an indispensable capability for long-horizon LLM agents, enabling them to preserve and utilize information accumulated across extended interactions. Existing memory-agent approaches are typically trained end-to-end with reinforcement learning on downstream tasks. However, collecting high-quality annotated problems for memory-intensive scenarios is costly, and the resulting training data often lack sufficient diversity to cover general memory behaviors. In this work, we propose MemTrain, a self-supervised training framework for generally enhancing the context-memory capability of LLM agents for more effective downstream post-training. MemTrain introduces two coupled proxy tasks over unlabeled Wikipedia corpora: (1) an end-to-end masked reconstruction objective, which requires the model to recover masked entities after multiple rounds of memory updates, thereby encouraging memory maintenance from the final outcome perspective; and (2) an intermediate memory recall objective, which requires the model to reconstruct masked historical information using intermediate memory states, encouraging faithful compression and memory completeness throughout the interaction process. The two objectives are jointly optimized using GRPO. Extensive experiments on long-text QA and search-based QA benchmarks demonstrate that MemTrain consistently improves downstream memory-intensive reasoning performance across different models, achieving gains of up to 17.67 points over direct task-specific post-training.
Ziheng Li, Xingrun Xing, Haoqing Wang +2
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · Samsung Research, Beijing, China