Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only become useful many steps later, while much of the observed utility may be inherited from information already stored before the rewrite. We introduce Memory Gain Policy Optimization (MGPO), which isolates the incremental value of each memory rewrite by crediting it for its marginal contribution to current and future downstream utility. This turns delayed memory utility into a direct learning signal for optimizing what information should persist. We study MGPO on document-level information extraction, where structured supervision makes the effects of individual memory updates directly measurable. MGPO improves extraction while reducing average memory length by nearly 80% relative to the initial memory policy before optimization. The learned memory policy also supports reuse and transfer across domains, downstream models without further training. These results show that effective memory learning depends not only on preserving useful information, but on identifying which memory updates create lasting incremental value.
Figures & tables
Figure 1: Overview of our methodology. A trainable writer sequentially updates a persistent memory state, while a frozen reader evaluates its downstream utility. A counterfactual comparison evaluates the pre- and post-rewrite memories on the same current and future targets. Summing these utility differences yields the Memory Gain , which serves as the learning signal for optimizing the writer.
Figure 2: Counterfactual evaluation matrix. Ft,j measures the utility of memory state mt on downstream target j , with the diagonal corresponding to the streaming trajectory during evaluation. The difference Ft,j−Ft−1,j attributes the marginal utility on target j to rewrite t , and MGt aggregates this contribution over all current and future targets ( j≥t ).
SciREX (in-domain)
AIPAN-10K (OOD)
Method
Memory length
Entity cluster
Binary relation
Memory length
Entity cluster
Binary relation
P
R
F1
P
R
F1
P
R
F1
P
R
F1
Direct Readout
Direct Readout
–
29.5
58.4
39.2
10.8
35.0
16.5
–
38.0
31.6
34.5
24.2
9.4
13.5
R1-RE
–
43.9
59.8
50.6
19.2
24.9
21.7
–
28.7
4.6
7.9
6.0
0.7
1.2
External Memory
Table 1: Document-level IE results on SciREX and AIPAN-10K. All memory policies are trained only on SciREX. SciREX reports in-domain performance, while AIPAN-10K evaluates out-of-domain (OOD) transfer across domain and extraction schema without further training or adaptation. Memory length is the average number of tokens in the memory state provided to the frozen reader for each chunk. Results are averaged over 4 evaluation runs. The 14B subscript denotes evaluation with the frozen Qwen3-14B reader in place of the method’s learned reader.
Figure 3: Inherited utility and gradient-estimator variance under the untrained memory policy on SciREX test. Left: Mean absolute utility before and after each memory rewrite, showing the utility already present in the preceding memory state. Mid: Mean absolute factual and MGPO returns. Right: Paired gradient-variance estimates at 261 fixed states from 66 documents. Variance is measured over 12,288 selected normalization parameters.
Figure 4: Effect of credit assignment on document-level extraction. Left: overall F1. Middle and right: precision and recall versus chunk distances, defined as the number of chunks separating the evidence required for a prediction. Error bars are standard deviations for 4 evaluation runs.
Figure 5: How credit assignment shapes memory rewrites. Top : Immediate versus future extraction gain for each rewrite. Bottom : Counterfactual contribution of each rewrite to subsequent chunks.
Setting
Reader
No Memory
Memory
Cluster
Relation
Cluster
Relation
Coupled
Self (8B)
44.7 ± 0.2
17.5 ± 0.8
44.9 ± 1.0
20.1 ± 1.4
Decoupled
Qwen3-8B ( Yang et al., 2025 )
34.3 ± 0.2
12.8 ± 0.6
44.5 ± 1.0
20.4 ± 1.1
Qwen3-14B
40.0 ± 0.8
18.0 ± 0.6
51.3 ± 0.9
26.5 ± 1.3
Qwen3-32B
45.2 ± 0.2
21.5 ± 0.5
50.7 ± 1.3
27.0 ± 1.4
Llama3.1-8B ( Grattafiori et al., 2024 )
34.5 ± 1.4
12.9 ± 0.9
41.5 ± 0.9
18.3 ± 0.6
Table 2: Pairing the learned memory writer with different readers. All rows use MGPO-trained memory writers. In the coupled variant, the writer and reader are trained jointly based on the same model. Results are mean ± standard deviation over 4 runs.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
# Docs
Avg. length
Max. length
Avg. entities
Avg. relations
SciREX
66
6,947
16,908
5.12
8.67
AIPAN-10K
50
15,375
59,344
28.00
65.54
Appendix
Table 3: IE Test-set statistics. Entity and relation counts are per-document averages over entity clusters and unique binary relations.
Method
Step reward rt
Return Gt
Terminal
–
F1doc
Factual
Ft,t
∑k=tNFk,k
Myopic
Δt,t
∑k=tNΔk,k
Full
MGt
∑k=tNMGk
Appendix
Table 4: Credit-assignment variants.
Type
Method
Trainable component
Readout
Chunk
Memory budget
Direct
Direct-Readout
–
Qwen3-14B
Whole doc.
None
R1-RE
Qwen3-8B readout
Self
Whole doc.
None
External
LightRAG
–
Qwen3-14B
1,024 / 4,096
256 / 1,024
Mem0
–
Qwen3-14B
1,024 / 4,096
256 / 1,024
Learned
MemAgent
Qwen3-8B writer
Self / Qwen3-14B
1,024 / 4,096
256 / 1,024
HiMPO
Qwen3-8B writer
Self / Qwen3-14B
1,024 / 4,096
256 / 1,024
Appendix
Table 5: Implementation settings of baselines.
Dataset
Cluster matching
SciREX
Greedy one-to-one matching by gold-span coverage ( ≥0.5 ), without an entity-type constraint.
AIPAN-10K
Same-type matching using exact normalized aliases for data types and token-F1 ( ≥0.5 ) for other entity types. Many-to-one matches are allowed.
Appendix
Table 6: Dataset-specific cluster matching for document-level evaluation.
Structured IE
SciREX
Typed Mention
Salient cluster
Binary relation
Binary relation
Method
P
R
F1
P
R
F1
P
R
F1
P
R
F1
Prior work
SciREX-P
50.3
90.2
64.6
17.6
78.4
28.8
2.8
58.2
5.4
6.5
41.1
9.6
TempGen
79.0
7.9
14.4
24.2
23.7
23.9
8.3
6.3
7.1
17.1
13.6
14.5
Direct Readout
Appendix
Table 7: Complete SciREX test results. Structured IE uses our document-level scorer, while SciREX reports binary relations using the unmodified TempGen evaluator ( Huang et al., 2021 ) . All non-prior-work results are mean@4.
Figure 6: Accuracy on BABILong QA1–QA5 across context lengths.
Figure 7: The linearized reward closely tracks exact F1 gain. Left: Exact document-level F1 gain versus summed Memory Gain on SciREX, computed from the same history-masked counterfactual counts. MGPO (blue; n=261 ) achieves Spearman ρ=0.988 and matching signs whenever both gains are nonzero; credit ablations and the base policy are shown in grey. Right: OLS slopes of exact on linearized gain with document-bootstrap 95% CIs ( B=2000 ). The surrogate slightly underestimates MGPO gains (1.16, 95% CI [1.09, 1.24]) and overestimates ablation loss magnitudes (0.77–0.86).
Figure 8: Noise in single-sample utility estimation. Empirical CDFs of ∣MGt∣ for unchanged-memory controls (grey) and actual rewrites (blue). The dashed line is the 95th percentile of the control distribution.
Memory
Chunk
Salient cluster
Binary relation
Memory budget, chunk fixed at 1024
64
1024
50.8 ± 0.4
26.1 ± 0.7
128
1024
51.3 ± 1.2
26.4 ± 1.6
256 †
1024 †
51.3 ± 0.9
26.5 ± 1.3
512
1024
51.5 ± 0.7
26.7 ± 0.8
Chunk size, memory fixed at 256
Appendix
Table 8: Sensitivity to memory budget and chunk size. We conduct experiments using the model trained with a configuration of a 1024 tokens chunk size and a 256 tokens memory budget. Performance is largely insensitive to the memory budget, while varying the chunk size shows a minor effect, peaking at 2048. † denotes the training configuration.
Method
G
Mem. tokens
Policy rollouts
Reader evaluations
GPU- hours
Peak mem. (GiB)
Cluster F1
Relation F1
MGPO-base
–
202
–
–
–
–
37.8 ± 0.4
15.5 ± 0.4
MGPO-GR
8
72
8
50.7
142.0
198.4
45.5 ± 0.5
20.1 ± 0.3
MGPO-EMA
1
43
1
8.3
26.9
192.5
51.3 ± 0.9
26.5 ± 1.3
Appendix
Table 9: Position EMA versus group-relative baseline on SciREX. Both variants use the same Memory Gain objective. MGPO-EMA uses position EMA ( G=1 ), whereas MGPO-GR uses GRPO ( G=8 ). Policy rollouts and reader evaluations are reported per training document. GPU-hours cover the complete training run, and peak memory is the analytical per-step training footprint.
Figure 9: Analytical FLOPs under the SciREX instrument, every call charged at its configured maximum width. Left: Inference: a full-context 14B reader scales quadratically in document length while chunk-wide MGPO scales linearly. Right: Training: Memory Gain (G=1) pays a triangular counterfactual reward matrix, GRPO pays G trajectories with diagonal reader cells. At the mean SciREX (train) document (dashed line, 6,346 tokens) Memory Gain costs half of GRPO-8, and only overtakes it past 16K tokens.
Figure 10: Training-time GPU memory breakdown on SciREX. All methods train the same Qwen3-8B policy under their respective training configurations. Policy memory is therefore identical, while rollout KV cache and activations vary with each method’s sequence and batch geometry. MGPO additionally maintains a frozen Qwen3-14B reader and its KV cache. Nevertheless, its chunk-wise training setup substantially reduces rollout and activation memory, yielding a total footprint comparable to MemAgent and lower than R1-RE.
Large language model (LLM) agents require long-term user memory for consistent personalization, but limited context windows hinder tracking evolving preferences over long interactions. Existing memory systems mainly rely on static, hand-crafted update rules; although reinforcement learning (RL)-based agents learn memory updates, sparse outcome rewards provide weak supervision, resulting in unstable long-horizon optimization. Drawing on memory schema theory and the functional division between prefrontal regions and hippocampus regions, we introduce MemCoE, a cognition-inspired two-stage optimization framework that learns how memory should be organized and what information to update. In the first stage, we propose Memory Guideline Induction to optimize a global guideline via contrastive feedback interpreted as textual gradients; in the second stage, Guideline-Aligned Memory Policy Optimization uses the induced guideline to define structured process rewards and performs multi-turn RL to learn a guideline-following memory evolution policy. We evaluate on three personalization memory benchmarks, covering explicit/implicit preference and different sizes and noise, and observe consistent improvements over strong baselines with favorable robustness, transferability, and efficiency.
Derong Xu, Shuochen Liu, Pengfei Luo +8
University of Science and Technology of China & State Key Laboratory of Cognitive Intelligence · City University of Hong Kong · Dalian University of Technology +1
Existing large language model (LLM) based memory systems apply universal, static policies that overlook a fundamental reality: the contexts that are worth storing in memory are different across users. This misalignment wastes limited memory budget on transient interactions while failing to preserve critical context for long horizon tasks. To address this gap, we investigate an underexplored question: can LLM based memory systems learn personalized memory policies? We introduce PerMemBench, the first benchmark for evaluating personalized memory systems, featuring multi year, multi domain interaction histories across diverse user personas. We further present the first empirical study of memory personalization, proposing session level storage gating, a lightweight framework that selectively bypasses memory operations for transient sessions. Our study confirms that personalization yields substantial retention gains under perfect gating, yet reveals that accurate gating remains an open and critical challenge.
Long-horizon agents rely on memory mechanisms to compress interaction history, but optimizing memory writing faces a distinct credit assignment challenge: a memory update may be rewarded or penalized due to downstream tool failures, noisy observations, or reasoning errors rather than its own contribution. We propose HiMPO, a Hindsight-Informed Memory Policy Optimization framework for assigning less-entangled credit to memory-writing actions in long-horizon agents. HiMPO first estimates the local utility of a memory update by comparing the task-relevant information recoverable from the previous and updated memories under the same pre-write state. It then uses hindsight relevance as a bounded retrospective filter that attenuates memory credit when local utility is not supported by the target outcome. The resulting memory-specific advantage is applied only to memory tokens, while trajectory-level rewards optimize the rest of the agent's behavior. Across judge-based open-domain tasks and objective compressive-memory QA, HiMPO improves over strong memory-based and RL-based baselines while preserving compressed-context efficiency. Controlled interventions and live replay studies further show that HiMPO reduces blame leakage from tool-induced errors, assigns memory credit that aligns with the functional impact of memory writes, and remains robust to noisy training targets.
Jiangze Yan, Yi Shen, Wenjing Zhang +5
Unicom Data Intelligence, China Unicom · Data Science & Artificial Intelligence Research Institute, China Unicom