Large language model (LLM) agents reuse external memory to guide new tasks, but effective retrieval requires learning which memory sets improve execution. Such learning relies on costly outcome feedback: ordinary retrieval observes only executed sets, while evaluating alternatives requires additional rollouts. We introduce \textsc{UpliftMem}, which learns memory retrieval from set-level execution uplift relative to the same executor without memory. A theoretical analysis of how retrieval preferences restrict feedback coverage motivates targeted probing of alternative memory sets. Probe selection follows an expected value of sample information (EVSI) criterion, derived in closed form under a correlated Gaussian model, to allocate limited training rollouts according to their expected improvement in local retrieval decisions. The shared scorer is trained with a frozen executor and selects memory sets without test-time probes. Across ALFWorld, WebShop, and BigCodeBench, \textsc{UpliftMem} achieves the best success rates among evaluated baselines on the main evaluation sets. Controlled fixed-store and matched probe budget evaluations further demonstrate improved memory-use decisions and more effective use of execution feedback.
Figures & tables
Figure 1: Learning memory retrieval from set-level execution feedback. (a) A relevant memory set may reduce utility relative to no-memory execution. (b) Only executed sets receive direct uplift labels; evaluating alternatives requires additional rollouts.
Figure 2: Overview of UpliftMem . A shared scorer Gθ guides retrieval and EVSI-based probing, trained with value and policy supervision while the executor remains frozen. Validated writing supplies additional labels and expands the memory pool.
Figure 3: Retrieval gains on MemSyco-Bench with fixed memory stores and K=5 . Each cell shows the change from native retrieval in accuracy (top) and the indicated behavioral metric (bottom), in percentage points. Changes in sycophancy and outdated-memory use are sign-reversed so that positive values indicate improvement.
Metric
Native
Base
Warm-up
Random
No policy
No value
UpliftMem
Decision ↑
41.7/44.4
43.1/44.5
49.6/48.8
55.9/54.9
56.7/56.1
55.8/53.1
66.5/68.8
Syco. ↓
49.0/40.0
47.2/36.7
44.2/34.7
38.7/31.3
39.2/31.9
41.6/35.3
29.5/16.5
Select ↑
54.3/61.1
55.1/62.2
53.7/57.9
58.3/60.3
56.9/60.2
58.0/62.8
61.6/70.7
Old ↓
47.0/40.8
46.5/39.6
47.6/42.4
42.5/41.3
43.8/40.3
43.6/38.9
39.3/31.7
Table 3: MemSyco-Bench ablations at K=5 , averaged over five memory stores. Each entry reports Qwen3-4B/Qwen3-8B.
Figure 4: Cumulative success rate and mean reward across probe budgets on WebShop.
Figure 5: Effect of the maximum retrieval budget K on MemSyco-Bench, averaged across five memory stores. Higher Decision and Select and lower Syco. and Old indicate better performance.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Execution-validated memory writing. Completed trajectories yield memory candidates that are validated on their source tasks and revised when needed. The selected version supplies value supervision, while validated uplift governs memory-pool admission.
Benchmark
Seed
Runtime
Evaluation
ALFWorld
500
1,200
140 valid-seen / 134 valid-unseen
WebShop
500
2,300
500 official test goals
BigCodeBench
—
798
342 held-out manifest tasks
Appendix
Table 4: Data partitions for the agent benchmarks. Seed and runtime counts denote execution samples or trajectories.
Scenario
Total
Train
Test
Personalized Memory Use
300
210
90
Valid Memory Selection
350
245
105
Memory Evidence Conflict
300
210
90
Contextual Scope Control
300
210
90
Objective Fact Judgment
300
210
90
Total
1,550
1,085
465
Appendix
Table 5: Dialogue-grouped MemSyco-Bench partition with seed 42. Scenario names follow the benchmark.
Setting
Value
LoRA rank / scaling / dropout
8 / 16 / 0.1
Learning rate
10−5
Training epochs
3
Processed tasks per acquisition event
400
Maximum new probe executions per event
300
Recalled memories per task
50
Appendix
Table 6: Scorer training, probe acquisition, and retrieval settings.
Table 7: Prompt interfaces and their principal inputs and outputs.
Environment
Revision output
ALFWorld
WHY IT FAILED , followed by a procedure of three to seven abstract steps.
WebShop
WHY IT FAILED , followed by one Type / Trigger / Action rule.
BigCodeBench, passing candidate
SUCCESS_PROCEDURE with Contract , Implementation , and Boundary .
BigCodeBench, failing candidate
FAILURE_REFLECTION identifying one unresolved bug class.
Appendix
Table 8: Environment-specific output contracts appended to the revision wrapper.
Figure 7: Probe-budget curves on ALFWorld (seen: top; unseen: bottom).
Figure 8: Probe-budget curves on the BigCodeBench validation split.
Memory System
Variant
When to Use Memory
How to Use Memory
Objective Fact
Scope Control
Evidence Conflict
Personalized Use
Valid Selection
Acc. ↑
Syco. ↓
Acc. ↑
Syco. ↓
Acc. ↑
Syco. ↓
Acc. ↑
Cor.Mem. ↑
Acc. ↑
Outd.Mem. ↓
Qwen3-4B
RAG
Base
35.3 ± 3.5
74.4 ± 4.4
24.4 ± 2.6
32.2 ± 1.1
57.0 ± 1.7
43.0 ± 1.7
85.1 ± 1.7
90.7 ± 1.7
53.7 ± 2.3
47.0 ± 2.5
UpliftMem
50.9 ± 2.4
49.1 ± 1.6
87.1 ± 1.3
8.7 ± 2.0
84.9 ± 0.6
14.9 ± 1.0
85.9 ± 2.8
91.1 ± 1.1
63.8 ± 2.1
37.1 ± 2.2
A-MEM
Base
36.7 ± 1.1
66.3 ± 0.6
48.4 ± 1.3
34.0 ± 2.7
46.2 ± 2.2
53.8 ± 2.2
82.2 ± 3.8
92.2 ± 2.2
55.2 ± 1.9
47.0 ± 3.1
Appendix
Table 9: Per-system MemSyco results at K=5 . Native retrieval (Base) and UpliftMem use identical memory stores. Entries are means over five inference runs, with ± denoting the sample standard deviation. Arrows indicate the preferred metric direction, and bold marks the better value within each pair.
Backbone
K
When to Use Memory
How to Use Memory
Overall
Obj. Fact
Scope Ctrl.
Evid. Conf.
Person. Use
Valid Sel.
Qwen3-4B
3
13.3
33.7
19.9
4.3
7.2
15.7
5
12.5
30.8
23.3
0.8
7.5
15.0
10
14.0
15.8
24.3
2.5
8.8
13.1
Qwen3-8B
3
13.6
12.4
25.0
2.8
9.7
12.7
5
14.6
12.3
45.1
1.4
9.4
16.5
Appendix
Table 10: Retrieval-budget sensitivity on MemSyco-Bench. Entries report direction-aligned UpliftMem gains over native retrieval in percentage points. Each task score averages its two metrics and five memory stores, and Overall averages the five task scores. The cross-model average weights Qwen3-4B and Qwen3-8B equally. Bold marks the largest gain within each backbone and column across budgets.
Memory System
Variant
When to Use Memory
How to Use Memory
Avg. Δ
Objective Fact
Scope Control
Evidence Conflict
Personalized Use
Valid Selection
Acc. ↑
Syco. ↓
Acc. ↑
Syco. ↓
Acc. ↑
Syco. ↓
Acc. ↑
Cor.Mem. ↑
Acc. ↑
Outd.Mem. ↓
RAG
Full UpliftMem
50.9
49.1
87.1
8.7
84.9
14.9
85.9
91.1
63.8
37.1
–
Base reranker
41.5
65.2
12.6
48.1
65.6
34.4
69.6
79.3
54.3
47.0
22.6
Warm-up only
41.5
64.1
59.3
24.8
61.5
38.5
69.6
82.2
54.0
47.3
16.1
Random probing
43.7
58.1
77.8
8.1
61.5
38.5
76.3
83.0
63.5
36.8
0 9.0
Appendix
Table 11: Per-system ablations at K=5 (Qwen3-4B). Full UpliftMem is averaged over five inference runs, while the ablation variants are averaged over three. Avg. Δ is the mean direction-aligned difference between Full and each variant across the ten metrics, in percentage points, with positive values favoring Full. Bold marks the best value within each memory system and metric.
Memory System
Variant
When to Use Memory
How to Use Memory
Avg. Δ
Objective Fact
Scope Control
Evidence Conflict
Personalized Use
Valid Selection
Acc. ↑
Syco. ↓
Acc. ↑
Syco. ↓
Acc. ↑
Syco. ↓
Acc. ↑
Cor.Mem. ↑
Acc. ↑
Outd.Mem. ↓
RAG
Full UpliftMem
51.8
41.2
82.4
0.4
94.2
5.8
82.6
78.9
73.9
26.1
–
Base reranker
47.0
51.5
45.2
1.9
40.0
60.0
74.1
70.0
61.9
38.4
20.4
Warm-up only
48.5
44.1
61.1
1.5
28.5
71.5
77.4
74.1
61.0
40.3
19.7
Random probing
48.9
42.2
81.2
0.4
31.5
68.5
78.5
75.9
64.4
38.7
16.0
Appendix
Table 12: Per-system ablations at K=5 (Qwen3-8B). Metrics and reporting conventions follow Table 11 .
Recent benchmarks for Large Language Model (LLM) agents mainly evaluate reasoning, planning, and execution. However, memory is also essential for agents, as it enables them to store, update, and retrieve information over time. This ability remains under-evaluated, largely because existing benchmarks do not provide a systematic way to assess memory mechanisms. In this paper, we study agent memory from a self-evolving perspective and introduce EvoMemBench, a unified benchmark organized along two axes: memory scope (in-episode vs. cross-episode) and memory content (knowledge-oriented vs. execution-oriented). We compare 15 representative memory methods with strong long-context baselines under a standardized protocol. Results show that current memory systems are still far from a general solution: long-context baselines remain highly competitive, memory helps most when the current context is insufficient or tasks are difficult, and no single memory form works consistently across all settings. Retrieval-based methods remain strong for knowledge-intensive settings, whereas procedural and long-term memory methods are more effective for execution-oriented tasks when their stored experience matches the task structure. We hope EvoMemBench facilitates future research on more effective memory systems for LLM-based agents. Our code is available at https://github.com/DSAIL-Memory/EvoMemBench.
Yuyao Wang, Zhongjian Zhang, Mo Chi +7
Hong Kong University of Science and Technology (Guangzhou) · Beijing University of Posts and Telecommunications · Beijing Institute of Technology +1
Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core bottleneck for agentic learning because optimal memory behavior is fundamentally context-dependent. The early stages of the tasks, benefit from minimal retrieval because memory is sparse; recurring goal types benefit from plan reuse rather than generic nearest-neighbor lookup; stuck agents benefit from re-retrieval with alternative queries; and across long task streams, the memory store itself must be consolidated and pruned to remain useful. We present Memory as a Controlled Process (MemCon), a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget. MemCon is backend-agnostic: it wraps any existing memory implementation, learns from task-by-task binary feedback with no pretraining and no additional LLM calls, and uses a lightweight tabular contextual bandit with UCB exploration that converges within tens of tasks. Across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones, MemCon consistently outperforms multiple memory baselines by up to 15.2 points in task success while reducing token consumption by 5--20%.
Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu +11
University of California Los Angeles · University of Washington · 3Northwestern University
Long-horizon large language model (LLM) agents accumulate interaction trajectories that quickly exceed any practical prompt budget, and existing memory methods either truncate aggressively and lose non-local evidence or retain boilerplate that degrades decision quality. We ask a mechanism question rather than claiming a better general-purpose memory system: when does organizing trajectory memory into overlapping semantic units (OSUs) -- groups of related steps in which one step may belong to several units -- help retrieval over flat or disjoint alternatives? We instantiate this in OSU-Mem, which retrieves from an overlapping OSU pool via budgeted coarse-to-fine expansion, and show its benefit is conditional: overlapping memory helps when the evidence steps a query needs share tool calls or entities, but hurts when those steps are fully heterogeneous and share neither. On a synthetic benchmark where evidence carries such shared structure by construction, OSU-Mem improves over the strongest baseline as the theory predicts; yet on a concatenated, constructed unaugmented τ-bench setting its aggregate advantage over flat retrieval vanishes. Splitting queries by whether their evidence shares tools and entities shows this near-tie to be an artifact of mixing query types rather than a property of either method, and ToolBench, a controlled probe built to carry shared structure by design, corroborates the same mechanism via an overlap-vs.-disjoint construction contrast (under a coverage-guided variant), isolating the construction principle rather than validating the full default system. Because the relevant sharing is cheaply estimable from metadata, the analysis yields a metadata-based heuristic for predicting when overlap is likely to improve retrieval. We deliberately isolate the retrieval layer, assessed by retrieval quality and an LLM-mediated evidence-selection stage.
Mellow Baixuan Chen, Xiangguo Sun
Courant Institute of Mathematical Sciences, New York University, New York, USA · Southeast University, Nanjing, China