Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This reactive paradigm means every retrieval failure corresponds to a real interaction in which the cost has already been paid. We propose MemDream, a framework that enables self-probing memory evolution for LLM agents. Our framework periodically enters offline dream cycles where three specialized agents (Dreamer, Analyst, Consolidator) collaboratively probe, diagnose, and repair the memory graph before failures occur. A policy trained via Group Relative Policy Optimization learns which repair operations produce durable retrieval improvements, while a soft decay mechanism provides reversible forgetting driven by the same anticipatory signal. Experiments on LoCoMo and MemoryAgentBench demonstrate that MemDream improves answer F1 by 4.5 points on LoCoMo and achieves a 9.1-point higher overall score on MAB over the strongest reactive-evolution baselines.
Figures & tables
Figure 1: (a) Existing reactive systems store a user’s nut allergy and baking preferences as disconnected nodes, missing the dependency between them. When queried for a recipe, the system recommends a nut-containing dish, a harmful answer the user must catch. (b) Through asynchronous background consolidation, MemDream ’s Dream Cycle finds and repairs this missing link, preventing the unsafe recommendation before the query arrives.
Figure 2: Overview of MemDream . Three LoRA-specialized agents share one Llama-3.1-8B backbone: the Dreamer generates probing queries aimed at weak regions, the Analyst diagnoses failures along four axes (buried, scattered, redundant, noisy), and the Consolidator applies repairs (extract, merge, synthesize, prune). The cycle repeats until retrieval F1 converges. We first train the agents independently with GRPO, then jointly fine-tune the Dreamer and Consolidator (with the Analyst frozen) using end-to-end memory quality as the reward.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Base model
Llama-3.1-8B-Instruct
Quantization
8-bit (bitsandbytes)
LoRA rank r
16
LoRA alpha α
32
Target modules
q_proj, k_proj, v_proj, o_proj
LoRA dropout
0.05
Appendix
Table 7: LoRA adapter configuration.
Parameter
Value
Group size N
4
Learning rate
1×10−3
Exploration rate ϵ
0.2 (decaying)
Training data
MSC dialogue (5 sessions)
Dream queries per conversation
8–12
Training steps (independent)
500 per agent
Appendix
Table 8: GRPO training hyperparameters.
Rounds
F1
B-1
R-L
0 (no dream)
48.72
38.15
46.83
1
51.36
40.62
49.58
3
53.41
42.87
51.92
5
53.89
43.21
52.35
8
53.52
42.94
51.87
Appendix
Table 9: Effect of dream cycle rounds on LoCoMo.
Dreamer Variant
Disc.
Cov.
Align.
Rule-Based
0.36
5.0
0.37
LLM Dreamer
0.15
4.8
0.35
Appendix
Table 10: Dream quality comparison.
Signal Source
Mean Reward
Std
Baseline (no dream)
0.198
0.031
Dream-guided (top-50%)
0.232
0.035
Dream-guided (top-30%)
0.241
0.033
Oracle (real queries)
0.262
0.038
Appendix
Table 11: Training signal comparison .
Component
Time
GPU Mem.
Load base model (8-bit)
45s
9.2 GB
Load 3 LoRA adapters
8s
+1.8 GB
Dreamer: generate queries
2.1 min
12.4 GB
Retrieval (BM25 + dense)
0.8 min
3.2 GB
Analyst: diagnose
1.5 min
12.4 GB
Consolidator: execute
3.2 min
12.4 GB
Appendix
Table 14: Approximate offline timing and GPU memory profile. Per-stage times describe one round; the final row reports the full-cycle estimate.
Quantity
Mean per conversation
Per initial node
LLM calls
60.9
0.199
Input tokens
33,588.4
109.69
Output tokens
13,004.5
42.47
Total tokens
46,592.9
152.16
Model generation time (s)
486.41
1.59
Cycle wall time (s)
2,208.84
7.21
Appendix
Table 15: Recorded overhead for the rule-probe configuration. Per-node values divide totals by 3,062 initial nodes. Model loading, initial memory construction, and LLM Dreamer generation are excluded; these measurements use a different configuration from Table 14 .
Method
Signal or criterion
Memory update
Auto-Dreamer
Task performance and counterfactual memory utility during training
Rewrite a selected region into a compact replacement set
TrustMem
Transition-level coverage, preservation, and faithfulness
Learn trustworthy updates from preferences over candidate transitions
RecMem
Recurrence of semantically similar interactions
Trigger consolidation and refine fine-grained semantic facts
MemDream
Failures exposed by self-generated retrieval probes
Diagnose failures, apply targeted repairs, and re-evaluate retrieval
Appendix
Table 16: Signals and update mechanisms in recent memory consolidation methods.
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · Samsung Research, Beijing, China