GenMem: Generative Symbolic Memory for Self-Evolving Harness
Authors: Xinke Jiang, Tao Feng, Weixuan Xu, Zhixin Zhang, Zhibang Yang, Wentao Zhang, Runchuan Zhu, Xu Chu, +2 more
Organizations: National Engineering Research Center of Software Engineering, Peking University, Beijing, China · School of Computer Science, Peking University, Beijing, China · Key Laboratory of High Confidence Software Technologies, Ministry of Education, Beijing, China · Center on Frontiers of Computing Studies, Peking University, Beijing, China · Peking University Information Technology Institute (Tianjin Binhai), Tianjin, China
Long-term memory supports the self-evolution of LLM agents by retaining experience and skills across tasks and enabling their retrieval, reuse, and revision in subsequent long-horizon decision-making. Yet existing memory management approaches remain limited to discriminative retrieval and to address the sparse, hierarchical, and highly redundant structure of reusable experience: only a small, task-dependent subset of trajectories and memories warrants retention, retrieval, or revision. Learning these operations is further complicated by sparse, delayed, and indirect task-level feedback, with weak supervision across the memory lifecycle. Moreover, continual memory evolution introduces an architectural tension as addressing invariance: stored experience is perpetually revised, yet the addressing interface consumed by learned retrieval policies must remain stable. To address, we present GenMem, which reformulates memory management as generative symbolic addressing. Its core mechanism is the Symbolic Identifier (SID), a multi-level discrete token tuple drawn from a Cartesian-product address space that factorizes a million-scale sparse memory space using fewer than one hundred discrete symbols. Instead of generating ever-changing raw content, the memory agent learns to generate SIDs, while memory evolution rewrites the payload at a fixed address without shifting the address itself. Architecturally, GenMem couples a MemRetriever and a MemEvolver within a multi-agent harness, trained via GRPO with dense process and outcome rewards with two-channels optimization. Under offline memory evolution, experiments spanning ALFWorld, WebShop, multi-hop QA, medical reasoning, and deep research evaluate GenMem against strong memory-augmented baselines...
Figures & tables
Figure 1: Two challenges in evolving experience memory. Left: Reusable experience is sparse and organized around shared sub-skills, while delayed task rewards hinder credit assignment. Right: Content updates can shift embedding-based retrieval, motivating stable addresses that remain invariant as experience evolves.
Figure 2: Overview of GenMem .
ALFWorld Success Rate (%)
WebShop
Method
Pick
Look
Clean
Heat
Cool
Pick2
All
Score
Succ.
Closed-source LLMs (absolute scores)
GPT-4o
75.3
60.8
31.2
56.7
21.6
49.8
48.0
31.8
23.7
Gemini-2.5-Pro
92.8
63.3
62.1
69.0
26.6
58.7
60.3
42.5
35.9
Prompt-based Agentic or Memory-based Methods ( Δ vs. ReAct)
ReAct
48.5 Δ 0.0
35.4 Δ 0.0
34.3 Δ 0.0
13.2 Δ 0.0
18.2 Δ 0.0
17.6 Δ 0.0
31.2 Δ 0.0
46.2 Δ 0.0
19.5 Δ 0.0
Table 1: Main results on ALFWorld and WebShop. ALFWorld reports per-subtask and overall success rates (%); WebShop reports task score and success rate (%).
In-Domain F1 (%)
Out-of-Domain F1 (%)
Overall
Method
2Wiki
HotpotQA
Bamboogle
MuSiQue
NQ
TriviaQA
Avg.
Prompt-based Agentic or Memory-based Methods ( Δ vs. ReAct)
Base
25.41 Δ -2.10
26.63 Δ -16.18
17.86 Δ -9.77
12.15 Δ -7.19
19.72 Δ -10.29
49.08 Δ -5.47
25.1 Δ -8.5
CoT
23.55 Δ -3.96
29.10 Δ -13.71
37.56 Δ +9.93
14.35 Δ -4.99
22.47 Δ -7.54
49.33 Δ -5.22
29.4 Δ -4.2
FS-RAG
17.71 Δ -9.80
29.21 Δ -13.60
16.86 Δ -10.77
10.74 Δ -8.60
16.82 Δ -13.19
35.02 Δ -19.53
21.1 Δ -12.5
FL-RAG
19.78 Δ -7.73
34.42 Δ -8.39
24.10 Δ -3.53
12.46 Δ -6.88
19.72 Δ -10.29
42.66 Δ -11.89
25.5 Δ -8.1
Table 2: Search-based QA results (Qwen2.5-7B; F1, %). 2Wiki and HotpotQA are in-domain; the remaining benchmarks are out-of-domain.
StackPlanner
TCRAG
OpenHands
Method
Search
Research
SQL
Code
Search
Research
SQL
Code
Search
Research
SQL
Code
Skill-free
58.0
49.96
78.0
74.0
56.0
48.88
80.0
62.0
57.0
51.48
79.0
72.0
+ GenMem Skills
58.0 Δ 0.0
52.03 Δ +2.07
83.0 Δ +5.0
75.0 Δ +1.0
55.0 Δ -1.0
50.92 Δ +2.04
81.0 Δ +1.0
63.0 Δ +1.0
58.0 Δ +1.0
53.60 Δ +2.12
81.0 Δ +2.0
74.0 Δ +2.0
Table 3: Skill transfer across held-out benchmarks and agent harnesses under DeepSeek-V4-Flash. Scores and parenthesized changes over the skill-free settings are in %; bold marks the better result within each harness.
Method
Metric
t=0
t=1
t=2
t=3
t=4
t=5
t=6
t=7
t=8
t=9
Qwen3-Emb
Cand. Hit
8.41
8.79
8.32
9.07
8.60
8.04
8.51
7.85
8.13
7.66
Hit@1
0.19
0.28
0.19
0.37
0.28
0.19
0.28
0.19
0.28
0.19
Hit@5
2.89
3.17
2.80
3.36
3.08
2.71
2.99
2.61
2.80
2.52
Task ACC
28.53
28.91
28.62
29.10
28.74
28.31
28.65
28.12
28.46
27.98
TF-IDF
Cand. Hit
18.47
19.03
20.06
19.31
19.78
20.34
19.59
19.97
19.41
19.69
Hit@1
2.36
2.64
3.01
2.73
2.92
3.20
2.83
3.01
2.73
2.92
Table 4: Retrieval hit rates (%) and downstream task accuracy (%) across evolution checkpoints. Dense methods rebuild indices at each step; SID addresses are frozen by construction.
Figure 3: Memory bank evolution dynamics over the query stream: SID lineage, SID update concentration, operation-type distribution, and payload edit distance.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Discrete token sequences as object representations. In GenMem , a text encoder and RQ-KMeans map seed experiences to SID addresses (Section 2 ). Other object types illustrate the general concept and are not evaluated in this work.
Figure 6: Conceptual illustration of SID-addressed memory and experience evolution. The illustrated versions represent successive contents at the same address. Semantic categories and the three-level SID are illustrative; the implemented system uses four residual-quantization levels without predefined semantic roles.
Benchmark
Evaluation role
Reported metric
ALFWorld
interactive task
Subtask and overall success rate
WebShop
interactive task
Task score and success rate
2WikiMultiHopQA (2Wiki)
In-domain QA
Answer F1
HotpotQA
In-domain QA
Answer F1
Bamboogle
Out-of-domain QA
Answer F1
MuSiQue
Out-of-domain QA
Answer F1
Appendix
Table 10: Evaluation benchmarks and metric correspondence. In-domain and out-of-domain designations follow the evaluation protocol in the main text. GPQA is used separately for the symbolic reasoning analysis.
Hyperparameter
Value
Backbone model
Qwen2.5-7B-Instruct
SID codebook ( N1×N2×N3×N4 )
48×16×8×8
Learning rate (alignment)
2×10−5
Learning rate (GRPO)
3×10−5
Global Batch size
256
GRPO group size G
8
Appendix
Table 11: Hyperparameters for reinforcement learning
Metric
Prefix
No prefix
Δ
Test reconstruction loss ↓
0.000332
0.000392
−15.27%
Single-code utilization (L1/L2/L3) ↑
100/100/100%
100/100/100%
0
Cumulative L1+L2 utilization ↑
84.72%
83.25%
+1.48 pp
Leaf-combination utilization ↑
23.75%
23.93%
−0.18 pp
Used leaves
6,567
6,617
−50
DUI uniformity ↑
0.8929
0.8730
+0.0199
Appendix
Table 12: Prefix-prompt comparison with a three-level 48×24×24 codebook and the reported Weighted setting ( w=0.5 ). The encoder and seed bank are held fixed. Values and rounded differences are transcribed from the supplied experiment summary; raw-precision outputs remain to be checked.
Table 13: Codebook search space. Each tuple lists the codebook sizes (N1,…,NL) ; counts in parentheses give the number of configurations at each depth.
Distance
Recon. Loss ( ×10−4 )
Leaf Util. (%)
Uniqueness Ratio
Avg. DUI (%)
Euclidean
3.38±0.09
14.92±3.68
0.0765±0.0241
56.53±3.71
Spherical
5.98±1.60
5.14±6.28
0.0270±0.0375
81.90±1.93
Weighted ( w=0.5 )
3.32±0.10
49.15±8.89
0.2523±0.0719
88.36±2.04
Appendix
Table 14: Distance-setting comparison, with mean ± standard deviation across 40 architectures. Loss is reported in ×10−4 .
Codebook
L
Capacity
Vocab
Recon.
Leaf Util.
Uniq.
DUI
Size
( ×10−4 )
(%)
Ratio
(%)
48 × 16 × 8 × 8
4
49,152
80
3.48
62.22
0.2212
90.70
128 × 24 × 16
3
49,152
168
3.31
64.25
0.2284
89.48
96 × 24 × 24
3
55,296
144
3.34
56.71
0.2268
88.74
72 × 32 × 24
3
55,296
128
3.36
51.26
0.2051
88.44
96 × 32 × 16
3
49,152
144
3.34
57.49
0.2044
87.53
Appendix
Table 15: Configuration comparison within the 40K–60K capacity band. The selected configuration ( bold ) uses the fewest SID tokens among the configurations shown. Loss is reported in ×10−4 .
Component
L1
L1 – L2
L1 – L3
L1 – L4
MemRetriever (query → SID)
40.90
13.00
6.30
3.75
MemEvolver (traj → SID)
40.41
11.29
3.27
1.36
SID → experience rewrite
ROUGE-L : 0.53
LLM-Judge : 0.57
Appendix
Table 16: SID alignment accuracy (%) at each cumulative hierarchical level, measured at Top-5 after reranking. For SID → experience reconstruction, ROUGE-L and LLM-Judge scores are reported separately.
Property
3-Level
4-Level
Number of levels
3
4
L1 size
48
48
L2 size
24
16
L3 size
24
8
L4 size
—
8
SID tokens per address
3
4
Appendix
Table 17: Codebook configuration: 3-level vs. 4-level SID.
Model
N
Avg Cand.
R@1
R@5
R@10
R@20
R@50
HR@1
HR@5
PT-3L (48 × 24 × 24)
2000
49.98
1.15
4.30
7.30
11.25
20.80
2.25
4.85
PT-4L (48 × 16 × 8 × 8)
2000
49.57
1.40
4.95
7.90
13.40
20.25
7.00
13.15
MT-3L MemR (48 × 24 × 24)
2000
29.43
0.75
1.75
2.90
5.05
10.85
2.95
5.35
MT-4L COT MemR (Full Beam)
2000
33.72
0.40
1.05
1.50
2.20
4.90
1.85
3.10
MT-4L COT MemR (SID-only)
2000
50.00
0.35
1.15
1.80
3.35
6.80
2.25
3.75
MT-4L COT MemE
735
50.00
0.14
0.54
—
—
5.31
0.68
1.36
Appendix
Table 18: Full SID retrieval results across PT and MT configurations. R@ k : beam recall; HR@ k : reranking hit rate.
Decoding
Beam
Top- k
Temp.
Top- p
Top-1 SID
Recall@5
Latency
Greedy
1
0
1.0
1.0
1.87
1.87
35.7
Beam deterministic
5
0
1.0
1.0
2.93
7.93
121.0
Beam deterministic
10
0
1.0
1.0
2.93
8.27
221.1
Beam deterministic
20
0
1.0
1.0
3.13
8.53
406.1
Beam sample
5
0
0.7
0.9
2.53
6.53
121.5
Beam sample
5
5
0.7
0.9
2.73
7.27
121.8
Appendix
Table 19: SID decoding configurations and retrieval results. Top-1 SID and Recall@5 are percentages; latency is reported in milliseconds. Greedy decoding returns one candidate, so its two retrieval scores coincide. Temperature and top- p entries are transcribed as reported and do not imply sampling in the deterministic rows. Bold denotes the best value in each result column, including ties.
Per-Level Independent (%)
Cumulative Prefix (%)
Model
Cutoff
L1
L2
L3
L4
L1
L1+L2
L1…L3
L1…L4
PT-3L
Top-1
31.85
24.80
14.00
—
31.85
7.85
1.15
—
Top-5
53.85
35.40
26.60
—
53.85
18.50
4.30
—
Top-50
88.20
69.55
67.05
—
88.20
49.95
20.80
—
PT-4L
Top-1
22.05
26.75
31.65
33.60
22.05
7.10
2.95
1.40
Top-5
44.25
42.30
47.50
52.05
44.25
18.80
9.05
4.95
Appendix
Table 20: Beam search hit rates (%) across PT and MT configurations. Left: per-level independent accuracy. Right: cumulative prefix accuracy (L1 through L ℓ must all match). L4 is N/A for 3-level models.
Per-Level Independent (%)
Cumulative Prefix (%)
Model
Cutoff
L1
L2
L3
L4
L1
L1+L2
L1…L3
L1…L4
PT-3L
Top-1
31.80
22.05
12.60
—
31.80
7.20
2.25
—
Top-5
59.55
39.25
28.80
—
59.55
19.35
4.85
—
PT-4L
Top-1
26.45
28.45
32.60
32.60
26.45
11.25
8.25
7.00
Top-5
54.45
50.25
57.35
59.85
54.45
25.70
16.40
13.15
MT-3L
Top-1
20.05
24.55
15.85
—
20.05
6.50
2.95
—
Appendix
Table 21: Rerank hit rates (%) (Beam@50 rescored by dense reranker). Layout matches Table 20 . L4 is N/A for 3-level models.
Type
N
R@1
R@5
R@50
Avg Cand.
merge_existing
590
0.169
0.339
5.763
50.00
insert_new
145
0.000
1.379
3.448
50.00
Total
735
0.136
0.544
5.306
50.00
Appendix
Table 22: MemE retrieval performance by update-operation type.
Figure 7: Alignment pretraining loss: 3-level vs. 4-level SID (500-step moving average, batch size 48). The 4-level SID converges faster and to a lower plateau.
Figure 8: SID codebook utilization across four levels (138,243 pre-merge entries). Bars show the fraction assigned to each token index; dashed line = uniform baseline. Normalized entropy and effective code counts annotated per level.
Figure 9: Mid-training loss for MemRetriever (left) and MemEvolver (right). Both converge smoothly on ReAct-style traces with embedded SID generation.
Figure 10: MemRetriever GRPO post-training curves (MA7-smoothed), tracking three eligibility rates and four reward signals over training steps. All metrics show steady improvement, confirming that GRPO effectively optimizes both SID targeting accuracy and downstream task utility.
Figure 11: MemEvolver GRPO post-training curves (MA7-smoothed). The first two panels report answer eligibility and retrieval hits, respectively, each divided by the total number of rollouts. The remaining panels track SID matching, SID and semantic rewards, outcome reward, and total reward.
Figure 12: StackPlanner RL training diagnostics. Search and WebShop use oracle-assisted central-agent training, while ALFWorld trains the Interactor with a 1:1 mixture of oracle-assisted and no-memory rollouts. Top panels show mean training reward and bottom panels show logged GRPO surrogate loss. The dashed line marks the Search restart after checkpoint 20.
Large Language Models (LLMs) show promise as tool-using agents but remain limited in long-horizon tasks that require remembering, organizing, and reusing knowledge. Prior memory approaches aim to resolve the situation, but mainly focus on storing factual information. Recent work on procedural memory improves task reuse, yet often reduces to replaying past successes without addressing failure cases or online scalability. We introduce a unified and automatic memory framework that integrates semantic, episodic, and procedural memory in a bi-level design combining short-term and long-term stores. A multi-agent architecture with actor, memory, and critic agents enables automatic memory generation, reward annotation, and adaptive retrieval. Long-term memory is managed through reward-based evaluation, merging, and pruning, ensuring scalability and continual improvement. Experiments across various environments show that our approach improves robustness and success on long multi-turn tasks compared to existing baselines. This work highlights the importance of comprehensive, adaptive memory for advancing LLM-based agents.
Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core bottleneck for agentic learning because optimal memory behavior is fundamentally context-dependent. The early stages of the tasks, benefit from minimal retrieval because memory is sparse; recurring goal types benefit from plan reuse rather than generic nearest-neighbor lookup; stuck agents benefit from re-retrieval with alternative queries; and across long task streams, the memory store itself must be consolidated and pruned to remain useful. We present Memory as a Controlled Process (MemCon), a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget. MemCon is backend-agnostic: it wraps any existing memory implementation, learns from task-by-task binary feedback with no pretraining and no additional LLM calls, and uses a lightweight tabular contextual bandit with UCB exploration that converges within tens of tasks. Across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones, MemCon consistently outperforms multiple memory baselines by up to 15.2 points in task success while reducing token consumption by 5--20%.
Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu +11
University of California Los Angeles · University of Washington · Northwestern University
Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This reactive paradigm means every retrieval failure corresponds to a real interaction in which the cost has already been paid. We propose MemDream, a framework that enables self-probing memory evolution for LLM agents. Our framework periodically enters offline dream cycles where three specialized agents (Dreamer, Analyst, Consolidator) collaboratively probe, diagnose, and repair the memory graph before failures occur. A policy trained via Group Relative Policy Optimization learns which repair operations produce durable retrieval improvements, while a soft decay mechanism provides reversible forgetting driven by the same anticipatory signal. Experiments on LoCoMo and MemoryAgentBench demonstrate that MemDream improves answer F1 by 4.5 points on LoCoMo and achieves a 9.1-point higher overall score on MAB over the strongest reactive-evolution baselines.
Mingfei Lu, Mengjia Wu, Runsong Jia +2
Australian Artificial Intelligence Institute (AAII) University of Technology Sydney