Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and discard details that later prove essential. Recent studies have begun shifting memory processing toward runtime adaptation, but typically specialize in particular operations or fixed processing schemes, leaving flexible control over performance, cost, and latency largely underexplored. To address this challenge, we present \textbf{MemPilot}, a flexible framework that orchestrates on-demand memory curation under different performance--cost--latency preferences. Specifically, we optimize a multi-step LLM policy via reinforcement learning to iteratively choose between retrieving from query-agnostic memory and delegating query-specific curation of raw multimodal history to heterogeneous LLMs and VLMs. The policy jointly controls evidence amount, curation instructions, model selection, and visual access, enabling fine-grained allocation of runtime computation. To optimize this policy under competing objectives, we adapt objective-wise advantage decoupling by separately estimating each objective's advantage before aggregation. Moreover, we introduce prefix-based marginal utility estimation for fine-grained credit assignment across multi-step rollouts. Experiments on five multimodal agent-memory benchmarks demonstrate favorable performance--cost--latency trade-offs across optimization preferences, with preference sweeps yielding broader frontiers than existing trade-off-aware baselines.
Figures & tables
Figure 1 : Overview of MemPilot. (a) The orchestrator iteratively retrieves compressed memories or delegates query-specific curation of raw multimodal histories to heterogeneous LLMs/VLMs. Factorized controls govern evidence retrieval, curation strategy, visual access, and model selection, with accumulated evidence guiding subsequent decisions and final answering. (b) Objective-wise advantage decoupling enables controllable trade-offs among quality, cost, and latency, while prefix-based answer probing estimates each step’s marginal utility for fine-grained credit assignment.
Method
Mem-Gallery
WorldMemArena
H2HMem
MemEye †
MemLens †
F1 ↑
L-J ↑
Cost ↓
F1 ↑
L-J ↑
Cost ↓
F1 ↑
L-J ↑
Cost ↓
F1 ↑
L-J ↑
Cost ↓
F1 ↑
L-J ↑
Cost ↓
Answer Model: Qwen3-VL-4B-Instruct
A-Mem
54.24
51.45
2.8e-2
21.38
47.73
1.8e-1
23.27
23.88
3.6e-2
15.20
20.28
1.3e-2
22.99
21.15
2.6e-2
SimpleMem
27.32
32.00
1.9e-2
15.58
26.36
4.5e-2
16.67
20.38
2.2e-2
14.69
16.87
4.6e-3
18.46
20.23
2.3e-3
MIRIX
36.70
49.45
3.3e-1
19.23
54.09
7.8e-1
19.81
25.56
3.8e-1
12.92
21.23
2.4e-1
15.91
13.73
8.4e-2
M2A
39.05
50.73
8.3e-1
20.14
54.49
1.1e+0
27.21
33.19
9.1e-1
12.26
17.86
8.5e-1
15.23
17.20
3.0e-1
Table 1 : Overall comparison on Mem-Gallery, WorldMemArena, H2HMem, MemEye, and MemLens . We report F1 score ( F1 ), LLM-judge score ( L-J ), and inference cost ( Cost ), with latency analyzed separately in Section 4.4 . Within each column, the best result is boldfaced and highlighted with a darker shade, while the second-best result is highlighted with a lighter shade.
Figure 2 : Performance–cost–latency frontiers. Results are macro-averaged over five benchmarks. Lines connect Pareto-optimal points, while faint markers show other evaluated settings. Operating points are obtained by varying optimization preferences or budget settings.
Figure 3 : Allocation of runtime computation under different optimization preferences. Panels (a,b) start with the same quality-oriented policy, with cost and latency proxy decreasing from left to right, respectively. Panel (c) shows model-selection shares across model groups. Values are averaged across five benchmarks; darker cells indicate larger values within each row.
Variant
L-J ↑
Cost ↓
MemPilot-Bal
37.40
1.1e-2
w/o LLM/VLM Delegation
14.58
5.6e-3
w/o Marginal Utility
14.82
5.5e-3
w/o GDPO
21.94
5.9e-3
w/o Retrieval–Instruction Decoupling
15.35
5.4e-3
Table 2: Ablation study. Results are averaged across five benchmarks.
Figure 4 : Memory-bank compatibility. Performance–cost frontiers with different memory-bank initializations.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
pmin
pmout
Latency profile ϕm
Text-only LLMs
Llama-3.2-3B-Instruct
0.05
0.33
(1.231,4.8×10−5,5.54×10−3)
Llama-3.1-8B-Instruct
0.05
0.08
(0.439,3.6×10−5,4.99×10−3)
Llama-3.3-70B-Instruct
0.13
0.40
(0.475,9.2×10−5,2.194×10−2)
Qwen3-30B-A3B-Instruct-2507
0.10
0.30
(0.524,4.5×10−5,5.33×10−3)
Qwen3-Next-80B-A3B-Instruct
0.10
1.10
(0.517,5.2×10−5,1.126×10−2)
Appendix
Table 3 : Heterogeneous model pool used in our experiments . Input and output prices, denoted by pmin and pmout , are in USD per million tokens. The latency profile is ordered as (bm,amin,amout) for LLMs and (bm,amin,amout,amimg) for VLMs.
Train
Validation
Test
Benchmark
Setting
#Inst.
#QA
#Inst.
#QA
#Inst.
#QA
Mem-Gallery
IID
14
1,062
2
190
4
275
WorldMemArena-Lifelong
IID
27
1,485
3
165
8
440
H2HMem-Dyadic
IID
13
1,122
2
174
4
327
H2HMem-Multiparty
IID
3
97
1
33
1
33
IID total
57
3,766
8
562
17
1,075
Appendix
Table 4 : Benchmark statistics after applying our data-integrity and answer-refusal filters . An instance denotes an independently partitioned memory history; for MEMLENS, each question-specific 32K history constitutes one instance. IID datasets are split at the instance level, whereas OOD benchmarks are used exclusively for testing.
Hyperparameter
Value
Policy model
Qwen3-4B-Instruct-2507
Embedding model
Qwen3-Embedding-0.6B
Maximum retrieved items per call
5
Rollouts per query G
4
Rollout temperature
1.0
Training batch size
32
Appendix
Table 5 : Hyperparameters used in the main experiments .
Figure 5 : Per-benchmark runtime computation allocation. Runtime curation behaviors under varying cost and latency preferences on Mem-Gallery, WorldMemArena, H2HMem, MemEye, and MemLens, respectively from top to bottom.
Variant
Mem-Gallery
WorldMemArena
H2HMem
MemLens †
MemEye †
L-J ↑
Cost ↓
L-J ↑
Cost ↓
L-J ↑
Cost ↓
L-J ↑
Cost ↓
L-J ↑
Cost ↓
MemPilot-Bal
64.27
1.4e-2
50.00
1.3e-2
28.96
1.6e-2
21.39
1.8e-4
22.37
1.0e-2
w/o LLM/VLM Delegation
32.00
8.1e-3
9.09
6.6e-3
12.08
8.2e-3
6.94
1.1e-4
12.80
5.2e-3
w/o Marginal Utility
31.18
7.6e-3
9.77
6.5e-3
17.57
7.9e-3
3.47
1.1e-4
12.13
5.1e-3
w/o GDPO
41.91
7.8e-3
21.36
7.3e-3
20.69
8.3e-3
9.83
1.1e-4
15.90
5.6e-3
w/o Retrieval–Instruction Decoupling
34.45
7.5e-3
10.91
6.4e-3
16.04
7.8e-3
3.47
1.1e-4
11.86
5.0e-3
Appendix
Table 6: Per-benchmark ablation results. Absolute Judge scores and inference costs are reported for each variant.
Figure 6 : Per-benchmark memory-bank compatibility. Performance–cost frontiers of MemPilot with different memory-bank initializations without retraining across five benchmarks.