Large language model (LLM) agents reuse external memory to guide new tasks, but effective retrieval requires learning which memory sets improve execution. Such learning relies on costly outcome feedback: ordinary retrieval observes only executed sets, while evaluating alternatives requires additional rollouts. We introduce \textsc{UpliftMem}, which learns memory retrieval from set-level execution uplift relative to the same executor without memory. A theoretical analysis of how retrieval preferences restrict feedback coverage motivates targeted probing of alternative memory sets. Probe selection follows an expected value of sample information (EVSI) criterion, derived in closed form under a correlated Gaussian model, to allocate limited training rollouts according to their expected improvement in local retrieval decisions. The shared scorer is trained with a frozen executor and selects memory sets without test-time probes. Across ALFWorld, WebShop, and BigCodeBench, \textsc{UpliftMem} achieves the best success rates among evaluated baselines on the main evaluation sets. Controlled fixed-store and matched probe budget evaluations further demonstrate improved memory-use decisions and more effective use of execution feedback.
Figures & tables
Figure 1: Learning memory retrieval from set-level execution feedback. (a) A relevant memory set may reduce utility relative to no-memory execution. (b) Only executed sets receive direct uplift labels; evaluating alternatives requires additional rollouts.
Figure 2: Overview of UpliftMem . A shared scorer Gθ guides retrieval and EVSI-based probing, trained with value and policy supervision while the executor remains frozen. Validated writing supplies additional labels and expands the memory pool.
Figure 3: Retrieval gains on MemSyco-Bench with fixed memory stores and K=5 . Each cell shows the change from native retrieval in accuracy (top) and the indicated behavioral metric (bottom), in percentage points. Changes in sycophancy and outdated-memory use are sign-reversed so that positive values indicate improvement.
Metric
Native
Base
Warm-up
Random
No policy
No value
UpliftMem
Decision ↑
41.7/44.4
43.1/44.5
49.6/48.8
55.9/54.9
56.7/56.1
55.8/53.1
66.5/68.8
Syco. ↓
49.0/40.0
47.2/36.7
44.2/34.7
38.7/31.3
39.2/31.9
41.6/35.3
29.5/16.5
Select ↑
54.3/61.1
55.1/62.2
53.7/57.9
58.3/60.3
56.9/60.2
58.0/62.8
61.6/70.7
Old ↓
47.0/40.8
46.5/39.6
47.6/42.4
42.5/41.3
43.8/40.3
43.6/38.9
39.3/31.7
Table 3: MemSyco-Bench ablations at K=5 , averaged over five memory stores. Each entry reports Qwen3-4B/Qwen3-8B.
Figure 4: Cumulative success rate and mean reward across probe budgets on WebShop.
Figure 5: Effect of the maximum retrieval budget K on MemSyco-Bench, averaged across five memory stores. Higher Decision and Select and lower Syco. and Old indicate better performance.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Execution-validated memory writing. Completed trajectories yield memory candidates that are validated on their source tasks and revised when needed. The selected version supplies value supervision, while validated uplift governs memory-pool admission.
Benchmark
Seed
Runtime
Evaluation
ALFWorld
500
1,200
140 valid-seen / 134 valid-unseen
WebShop
500
2,300
500 official test goals
BigCodeBench
—
798
342 held-out manifest tasks
Appendix
Table 4: Data partitions for the agent benchmarks. Seed and runtime counts denote execution samples or trajectories.
Scenario
Total
Train
Test
Personalized Memory Use
300
210
90
Valid Memory Selection
350
245
105
Memory Evidence Conflict
300
210
90
Contextual Scope Control
300
210
90
Objective Fact Judgment
300
210
90
Total
1,550
1,085
465
Appendix
Table 5: Dialogue-grouped MemSyco-Bench partition with seed 42. Scenario names follow the benchmark.
Setting
Value
LoRA rank / scaling / dropout
8 / 16 / 0.1
Learning rate
10−5
Training epochs
3
Processed tasks per acquisition event
400
Maximum new probe executions per event
300
Recalled memories per task
50
Appendix
Table 6: Scorer training, probe acquisition, and retrieval settings.
Table 7: Prompt interfaces and their principal inputs and outputs.
Environment
Revision output
ALFWorld
WHY IT FAILED , followed by a procedure of three to seven abstract steps.
WebShop
WHY IT FAILED , followed by one Type / Trigger / Action rule.
BigCodeBench, passing candidate
SUCCESS_PROCEDURE with Contract , Implementation , and Boundary .
BigCodeBench, failing candidate
FAILURE_REFLECTION identifying one unresolved bug class.
Appendix
Table 8: Environment-specific output contracts appended to the revision wrapper.
Figure 7: Probe-budget curves on ALFWorld (seen: top; unseen: bottom).
Figure 8: Probe-budget curves on the BigCodeBench validation split.
Memory System
Variant
When to Use Memory
How to Use Memory
Objective Fact
Scope Control
Evidence Conflict
Personalized Use
Valid Selection
Acc. ↑
Syco. ↓
Acc. ↑
Syco. ↓
Acc. ↑
Syco. ↓
Acc. ↑
Cor.Mem. ↑
Acc. ↑
Outd.Mem. ↓
Qwen3-4B
RAG
Base
35.3 ± 3.5
74.4 ± 4.4
24.4 ± 2.6
32.2 ± 1.1
57.0 ± 1.7
43.0 ± 1.7
85.1 ± 1.7
90.7 ± 1.7
53.7 ± 2.3
47.0 ± 2.5
UpliftMem
50.9 ± 2.4
49.1 ± 1.6
87.1 ± 1.3
8.7 ± 2.0
84.9 ± 0.6
14.9 ± 1.0
85.9 ± 2.8
91.1 ± 1.1
63.8 ± 2.1
37.1 ± 2.2
A-MEM
Base
36.7 ± 1.1
66.3 ± 0.6
48.4 ± 1.3
34.0 ± 2.7
46.2 ± 2.2
53.8 ± 2.2
82.2 ± 3.8
92.2 ± 2.2
55.2 ± 1.9
47.0 ± 3.1
Appendix
Table 9: Per-system MemSyco results at K=5 . Native retrieval (Base) and UpliftMem use identical memory stores. Entries are means over five inference runs, with ± denoting the sample standard deviation. Arrows indicate the preferred metric direction, and bold marks the better value within each pair.
Backbone
K
When to Use Memory
How to Use Memory
Overall
Obj. Fact
Scope Ctrl.
Evid. Conf.
Person. Use
Valid Sel.
Qwen3-4B
3
13.3
33.7
19.9
4.3
7.2
15.7
5
12.5
30.8
23.3
0.8
7.5
15.0
10
14.0
15.8
24.3
2.5
8.8
13.1
Qwen3-8B
3
13.6
12.4
25.0
2.8
9.7
12.7
5
14.6
12.3
45.1
1.4
9.4
16.5
Appendix
Table 10: Retrieval-budget sensitivity on MemSyco-Bench. Entries report direction-aligned UpliftMem gains over native retrieval in percentage points. Each task score averages its two metrics and five memory stores, and Overall averages the five task scores. The cross-model average weights Qwen3-4B and Qwen3-8B equally. Bold marks the largest gain within each backbone and column across budgets.
Memory System
Variant
When to Use Memory
How to Use Memory
Avg. Δ
Objective Fact
Scope Control
Evidence Conflict
Personalized Use
Valid Selection
Acc. ↑
Syco. ↓
Acc. ↑
Syco. ↓
Acc. ↑
Syco. ↓
Acc. ↑
Cor.Mem. ↑
Acc. ↑
Outd.Mem. ↓
RAG
Full UpliftMem
50.9
49.1
87.1
8.7
84.9
14.9
85.9
91.1
63.8
37.1
–
Base reranker
41.5
65.2
12.6
48.1
65.6
34.4
69.6
79.3
54.3
47.0
22.6
Warm-up only
41.5
64.1
59.3
24.8
61.5
38.5
69.6
82.2
54.0
47.3
16.1
Random probing
43.7
58.1
77.8
8.1
61.5
38.5
76.3
83.0
63.5
36.8
0 9.0
Appendix
Table 11: Per-system ablations at K=5 (Qwen3-4B). Full UpliftMem is averaged over five inference runs, while the ablation variants are averaged over three. Avg. Δ is the mean direction-aligned difference between Full and each variant across the ten metrics, in percentage points, with positive values favoring Full. Bold marks the best value within each memory system and metric.
Memory System
Variant
When to Use Memory
How to Use Memory
Avg. Δ
Objective Fact
Scope Control
Evidence Conflict
Personalized Use
Valid Selection
Acc. ↑
Syco. ↓
Acc. ↑
Syco. ↓
Acc. ↑
Syco. ↓
Acc. ↑
Cor.Mem. ↑
Acc. ↑
Outd.Mem. ↓
RAG
Full UpliftMem
51.8
41.2
82.4
0.4
94.2
5.8
82.6
78.9
73.9
26.1
–
Base reranker
47.0
51.5
45.2
1.9
40.0
60.0
74.1
70.0
61.9
38.4
20.4
Warm-up only
48.5
44.1
61.1
1.5
28.5
71.5
77.4
74.1
61.0
40.3
19.7
Random probing
48.9
42.2
81.2
0.4
31.5
68.5
78.5
75.9
64.4
38.7
16.0
Appendix
Table 12: Per-system ablations at K=5 (Qwen3-8B). Metrics and reporting conventions follow Table 11 .