Memory self-evolution uses task feedback to iteratively improve executable memory programs that store and retrieve information from past interactions. Existing approaches typically adopt holistic evolution, deriving revision directions from mixed feedback and judging progress by overall performance. This can obscure optimization directions and hide capability-specific gains offset by regressions elsewhere, leaving promising directions underexplored. We introduce capability-driven evolution, which extends search guidance from overall performance to individual capability dimensions, preserving promising revisions and expanding exploration beyond the boundaries of holistic evolution. We propose PrisMem, which uses dependency-aware capability selection to prioritize targets with potential cross-capability benefits and history-guided diagnosis to refine capability specialists. Trace-guided integration compares evaluated programs on paired differential cases, using their behavioral differences to consolidate complementary gains into a unified memory program. Experiments show that PrisMem outperforms the strongest baselines by 10.54 and 7.83 percentage points on BEAM-1M and LongMemEval-M, respectively, demonstrating its effectiveness on million-token histories.
Figures & tables
Figure 1: Motivation for capability-driven evolution. Analysis of M ⋆ and EvolveMem evolution traces on BEAM. (a) Overall and maximum capability-level score changes for each revision. Many revisions improve at least one capability despite lower overall scores (upper-left region). (b) Among revisions that improve one capability (row), the percentage for which another capability’s score decreases (column). (c) Holistic evolution confines exploration to a limited region. (d) Capability-driven evolution enables broader exploration through explicit capability directions.
Capability
Memory requirement
Factual retrieval ( F )
Preserve and retrieve explicit facts, values, and other localized details.
Temporal tracking ( T )
Order events, determine recency, track states and knowledge updates.
Preference extraction ( P )
Recover and apply user preferences, instructions and constraints.
Multi-session synthesis ( M )
Aggregate and synthesize evidence across multiple sessions or events.
Adversarial ( A )
Abstain when memory lacks supporting evidence, or cautiously identify and resolve contradictory or misleading evidence.
Table 1: Five memory capabilities used in our experiments.
Figure 2: Overview of PrisMem . Stage I: Cold start quickly establishes a balanced program and initial evolution history. Stage II: Capability refinement uses dependency-aware capability selection and history-guided case selection to target capabilities and diagnose informative failures, respectively, refining their specialists through reflection and code revision. Stage III: Trace-guided integration compares the base program ( ⋆ ) and other capability specialists on paired differential cases, using their behavioral differences to guide integration into a unified memory program.
Method
Score (%) ↑
Avg. #Tokens (K) ↓
Factual
Temporal
Preference
Multi-session
Adversarial
Overall
retrieval
tracking
extraction
synthesis
BEAM-1M
Mem0
54.74 ± 3.06
46.51 ± 0.79
60.48 ± 2.39
38.02 ± 0.63
49.11 ± 1.24
48.95 ± 0.61
190.28 ± 0.20
A-MEM
56.21 ± 1.83
46.22 ± 0.62
65.39 ± 0.52
53.87 ± 0.86
41.13 ± 0.51
51.56 ± 0.62
428.95 ± 3.58
HippoRAG2
65.24 ± 1.26
47.06 ± 1.85
66.43 ± 1.57
44.83 ± 0.60
44.82 ± 2.28
51.86 ± 0.87
207.70 ± 0.19
Table 2: Performance on BEAM-1M and LongMemEval-M with Qwen3.8-27B. “Score” denotes the benchmark-provided LLM judge score, and “Avg. #Tokens” denotes the task LLM’s average token usage per question. All results are reported as mean ± std. Best results are in bold .
Method
BEAM
LongMemEval
M ⋆
76.5 ± 7.1
202.3 ± 10.3
EvolveMem
69.2 ± 7.3
102.7 ± 8.4
PrisMem
73.7 ± 6.4
143.3 ± 9.6
Table 3: Evolution tokens (millions, ↓ ) with Qwen3.8-27B on two benchmarks.
Variant
Overall score
Δ
(%) ↑
(pp)
Full PrisMem
63.39 ± 1.02
–
w/o capability evolution
53.24 ± 2.38
− 10.15
w/o trace-guided integration
60.23 ± 1.86
− 3.16
Round-robin target selection
60.84 ± 1.13
− 2.55
Random case selection
58.13 ± 1.25
− 5.26
Table 4: Ablations with Qwen3.8-27B on BEAM-1M. Δ is relative to full PrisMem .
Figure 3: Evolution trace analysis. (a) Gains of each Stage II iteration relative to program 05 obtained after cold start. Cell numbers are the creation iteration of the retained specialist for each capability; dots mark specialist updates. Scores improve intermittently across capabilities as refinement progresses. (b) Capability-wise scores of M ⋆ , EvolveMem, and PrisMem . Solid lines show the final selected programs, while dashed lines show the best capability-wise score reached across all iterations. PrisMem outperforms holistic evolution across capabilities, expanding the performance boundary, while its final integrated program closely matches its historical capability-wise best.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Question type
Primary memory capabilities
LoCoMo ( Maharana et al., 2024 )
Single-hop; Open-domain knowledge 1
Factual retrieval
Multi-hop
Multi-session synthesis
Temporal reasoning
Temporal tracking
Adversarial
Adversarial
LongMemEval ( Wu et al., 2025 )
Appendix
Table 5: Correspondence between benchmark question types and the memory capabilities used in our study. The entries describe primary requirements; individual questions may involve additional capabilities.
Changed component
Re-execution begins at
Reusable work
Memory-unit schema or extraction stage
Extraction
Segmentation
Constructor, index, or helpers
Indexing
Segmentation and extracted units
Query schema, planning, or retrieval
Query planning
Segmentation, extraction, and indexed memory
Answer instruction only
Answer generation
Built memory, query plan, and retrieved context
Appendix
Table 6: Re-execution boundaries for a candidate relative to its parent. For changes spanning several components, the earliest affected stage determines what can be reused.
Dataset
Split
F
T
P
M
A
Total
Avg. tokens (K)
BEAM-100K
Train
20
43
32
33
32
160
125.11
BEAM-100K
Validation
18
32
24
22
24
120
146.83
BEAM-1M
Test
70
210
140
140
140
700
1,089.57
LongMemEval-S
Train
19
17
11
9
4
60
109.40
LongMemEval-S
Validation
20
17
8
8
7
60
109.40
LongMemEval-M
Test
64
64
30
32
10
200
1,079.04
Appendix
Table 7: Sizes, primary capability composition, and average history length of each dataset split. F, T, P, M, and A denote factual retrieval, temporal tracking, preference extraction, multi-session synthesis, and adversarial, respectively. Average history length is measured with the Qwen3.8-27B tokenizer.
Parameter
Value
Stage II: Capability refinement
Case score threshold δ
0.8
Refinement case budget K
5
Stage III: Integration
Capability-wise score degradation ratio θ
0.1
Integration case budget
1 per capability
Appendix
Table 8: Parameter settings of PrisMem we used in experiments.
Method
Overall Score
Avg. #Tokens
(%) ↑
(K) ↓
Mem0
57.77 ± 0.17
180.21 ± 0.12
A-MEM
60.92 ± 0.41
352.11 ± 0.25
HippoRAG2
57.73 ± 0.50
193.52 ± 0.18
SimpleMem
52.36 ± 0.55
121.39 ± 0.69
LightMem
46.68 ± 0.05
94.05 ± 0.19
Appendix
Table 9: Performance on BEAM with GPT-5.5.
Method
Evolved with Qwen3.8-27B (transfer)
Evolved with GPT-5.5 (target model)
PrisMem
67.25 ± 1.35
67.09 ± 1.79
Appendix
Table 10: Model transfer evaluated with GPT-5.5 on BEAM-1M. All programs are evolved on BEAM-100K. Overall scores (%) are reported as “mean ± std” over three runs.
Method
Evolved on BEAM (transfer)
Evolved on LongMemEval (target dataset)
PrisMem
76.50 ± 1.67
75.50 ± 1.80
Appendix
Table 11: Dataset transfer evaluated with Qwen3.8-27B on LongMemEval-M. All programs are evolved with Qwen3.8-27B. Overall scores (%) are reported as “mean ± std” over three runs.
Strategy
Overall Score (%)
Question-only router
57.74 ± 1.21
Oracle-label router
62.27 ± 1.14
PrisMem
63.39 ± 1.02
Appendix
Table 12: Routing versus integration on BEAM-1M with Qwen3.8-27B.
Evolution Trace
Failure diagnosis
Revision summary
Current best for
Stage I: Cold-start revisions
n05
A single flat query could not capture multiple semantic aspects of synthesis questions, causing retrieval to over-concentrate on one topic while missing complementary evidence scattered across sessions.
Decompose multi-aspect questions into focused sub-queries and retrieve each aspect independently with proportional budget allocation and cross-query deduplication. schema planning retrieval answering
M T A P F
Stage II: Capability-directed refinement
n05 → n06 Synthesis
Over-atomized memory units fragmented each logical item into multiple small facts, causing related evidence to appear as an unstructured list and hindering distinct-item counting and cross-session progression synthesis.
Add a topic label to each memory unit and group retrieved evidence by topic, so related facts are presented as coherent blocks for counting and multi-session synthesis. schema extraction indexing retrieval answering
F
n05 → n07 Temporal
Session-level temporal provenance was lost during indexing, so retrieved memories lacked discussion-time context and the model could not reliably distinguish statement order from event dates.
Add session date to each memory unit, recover it from the source session during indexing, and expose explicit date markers in retrieved evidence while preserving relevance-based ranking. schema extraction indexing retrieval answering helpers
F T P
n05 → n08 Adversarial
Memory units lacked speaker provenance, causing user-confirmed facts to be conflated with assistant recommendations or hypotheticals and obscuring genuine contradictions between user statements.
Add explicit speaker labels to memory units and propagate them through retrieval, then constrain answering to treat user-sourced units as direct evidence and surface conflicting user statements without resolution. schema extraction indexing retrieval answering
A
Appendix
Table 13: BEAM evolution traces with failure diagnoses, revision summaries, and current specialist. F : factual retrieval, T : temporal tracking, P : preference extraction, M : multi-session synthesis, A : adversarial.
Long-horizon autonomous agents require memory systems to retain historical information, track evolving states, and reuse relevant knowledge beyond finite context windows. Existing agentic memory systems typically follow a memory construction-retrieval (MCR) pipeline, but often adapt mainly the memory bank while keeping the surrounding pipeline fixed after deployment. This fixed-pipeline design struggles to handle heterogeneous task-specific failure modes and can become misaligned with memory banks that evolve in scale and structure over time. To address these limitations, we propose MemPro, a system-level evolution framework that treats the entire MCR pipeline as an evolvable program rather than adapting only the memory bank or prompt text. MemPro maintains a version tree of runnable memory-system implementations, where an Evolving Agent iteratively selects promising versions, diagnoses recurring failures, and creates improved child versions through failure-mode-guided edit-debug refinement. Experiments on LongMemEval, LoCoMo, HotpotQA, and NarrativeQA show that MemPro consistently outperforms strong static and prompt-level evolving baselines within a few iterations, continues to improve with evolution, and achieves a favorable performance-cost trade-off. Code is available at https://github.com/wanghai673/MemPro.
Long-term memory is essential for LLM agents that operate across multiple sessions, yet existing memory systems treat retrieval infrastructure as fixed: stored content evolves while scoring functions, fusion strategies, and answer-generation policies remain frozen at deployment. We argue that truly adaptive memory requires co-evolution at two levels: the stored knowledge and the retrieval mechanism that queries it. We present EvolveMem, a self-evolving memory architecture that exposes its full retrieval configuration as a structured action space optimized by an LLM-powered diagnosis module. In each evolution round, the module reads per-question failure logs, identifies root causes, and proposes targeted configuration adjustments; a guarded meta-analyzer applies them with automatic revert-on-regression and explore-on-stagnation safeguards. This closed-loop self-evolution realizes an AutoResearch process: the system autonomously conducts iterative research cycles on its own architecture, replacing manual configuration tuning. Starting from a minimal baseline, the process converges autonomously, discovering effective retrieval strategies including entirely new configuration dimensions not present in the original action space. On LoCoMo, EvolveMem outperforms the strongest baseline by 25.7% relative and achieves a 78.0% relative improvement over the minimal baseline. On MemBench, EvolveMem exceeds the strongest baseline by 18.9% relative. Evolved configurations transfer across benchmarks with positive rather than catastrophic transfer, indicating that the self-evolution process captures universal retrieval principles rather than benchmark-specific heuristics. Code is available at https://github.com/aiming-lab/SimpleMem.
Existing self-evolving memory systems mainly improve agent memory based on textual outputs, such as task trajectories and reflections. However, this text-based paradigm rarely incorporates internal mechanistic signals, leaving how retrieved memory is actually utilized during task execution underexplored. This limitation can lead to unreliable error attribution and hallucinated memory modifications. In this work, we show that retrieval-head attention provides a mechanistic signal for revealing segment-level memory utilization. By aggregating attention over memory segments and decision steps, we construct a context utilization matrix that exposes recurring memory-use patterns and indicates corresponding refinement strategies. Building on this observation, we propose Attention-Guided Memory Refinement (AGMR), a framework that uses utilization patterns revealed by attention to guide targeted segment-level memory updates. AGMR corrects or enhances memory for failed executions, simplifies memory for successful executions, and verifies each update through re-execution. Experiments on interactive decision-making benchmarks show that AGMR improves both task performance and memory efficiency over text-only memory refinement baselines. Code is available at https://anonymous.4open.science/r/AGMR_code-3262/
Yechao Hong, Haiquan Qiu, Yaqing Wang +1
Department of Electronic Engineering, Tsinghua University · Beijing Institute of Mathematical Sciences and Applications