Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only become useful many steps later, while much of the observed utility may be inherited from information already stored before the rewrite. We introduce Memory Gain Policy Optimization (MGPO), which isolates the incremental value of each memory rewrite by crediting it for its marginal contribution to current and future downstream utility. This turns delayed memory utility into a direct learning signal for optimizing what information should persist. We study MGPO on document-level information extraction, where structured supervision makes the effects of individual memory updates directly measurable. MGPO improves extraction while reducing average memory length by nearly 80% relative to the initial memory policy before optimization. The learned memory policy also supports reuse and transfer across domains, downstream models without further training. These results show that effective memory learning depends not only on preserving useful information, but on identifying which memory updates create lasting incremental value.
Figures & tables
Figure 1: Overview of our methodology. A trainable writer sequentially updates a persistent memory state, while a frozen reader evaluates its downstream utility. A counterfactual comparison evaluates the pre- and post-rewrite memories on the same current and future targets. Summing these utility differences yields the Memory Gain , which serves as the learning signal for optimizing the writer.
Figure 2: Counterfactual evaluation matrix. Ft,j measures the utility of memory state mt on downstream target j , with the diagonal corresponding to the streaming trajectory during evaluation. The difference Ft,j−Ft−1,j attributes the marginal utility on target j to rewrite t , and MGt aggregates this contribution over all current and future targets ( j≥t ).
SciREX (in-domain)
AIPAN-10K (OOD)
Method
Memory length
Entity cluster
Binary relation
Memory length
Entity cluster
Binary relation
P
R
F1
P
R
F1
P
R
F1
P
R
F1
Direct Readout
Direct Readout
–
29.5
58.4
39.2
10.8
35.0
16.5
–
38.0
31.6
34.5
24.2
9.4
13.5
R1-RE
–
43.9
59.8
50.6
19.2
24.9
21.7
–
28.7
4.6
7.9
6.0
0.7
1.2
External Memory
Table 1: Document-level IE results on SciREX and AIPAN-10K. All memory policies are trained only on SciREX. SciREX reports in-domain performance, while AIPAN-10K evaluates out-of-domain (OOD) transfer across domain and extraction schema without further training or adaptation. Memory length is the average number of tokens in the memory state provided to the frozen reader for each chunk. Results are averaged over 4 evaluation runs. The 14B subscript denotes evaluation with the frozen Qwen3-14B reader in place of the method’s learned reader.
Figure 3: Inherited utility and gradient-estimator variance under the untrained memory policy on SciREX test. Left: Mean absolute utility before and after each memory rewrite, showing the utility already present in the preceding memory state. Mid: Mean absolute factual and MGPO returns. Right: Paired gradient-variance estimates at 261 fixed states from 66 documents. Variance is measured over 12,288 selected normalization parameters.
Figure 4: Effect of credit assignment on document-level extraction. Left: overall F1. Middle and right: precision and recall versus chunk distances, defined as the number of chunks separating the evidence required for a prediction. Error bars are standard deviations for 4 evaluation runs.
Figure 5: How credit assignment shapes memory rewrites. Top : Immediate versus future extraction gain for each rewrite. Bottom : Counterfactual contribution of each rewrite to subsequent chunks.
Setting
Reader
No Memory
Memory
Cluster
Relation
Cluster
Relation
Coupled
Self (8B)
44.7 ± 0.2
17.5 ± 0.8
44.9 ± 1.0
20.1 ± 1.4
Decoupled
Qwen3-8B ( Yang et al., 2025 )
34.3 ± 0.2
12.8 ± 0.6
44.5 ± 1.0
20.4 ± 1.1
Qwen3-14B
40.0 ± 0.8
18.0 ± 0.6
51.3 ± 0.9
26.5 ± 1.3
Qwen3-32B
45.2 ± 0.2
21.5 ± 0.5
50.7 ± 1.3
27.0 ± 1.4
Llama3.1-8B ( Grattafiori et al., 2024 )
34.5 ± 1.4
12.9 ± 0.9
41.5 ± 0.9
18.3 ± 0.6
Table 2: Pairing the learned memory writer with different readers. All rows use MGPO-trained memory writers. In the coupled variant, the writer and reader are trained jointly based on the same model. Results are mean ± standard deviation over 4 runs.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
# Docs
Avg. length
Max. length
Avg. entities
Avg. relations
SciREX
66
6,947
16,908
5.12
8.67
AIPAN-10K
50
15,375
59,344
28.00
65.54
Appendix
Table 3: IE Test-set statistics. Entity and relation counts are per-document averages over entity clusters and unique binary relations.
Method
Step reward rt
Return Gt
Terminal
–
F1doc
Factual
Ft,t
∑k=tNFk,k
Myopic
Δt,t
∑k=tNΔk,k
Full
MGt
∑k=tNMGk
Appendix
Table 4: Credit-assignment variants.
Type
Method
Trainable component
Readout
Chunk
Memory budget
Direct
Direct-Readout
–
Qwen3-14B
Whole doc.
None
R1-RE
Qwen3-8B readout
Self
Whole doc.
None
External
LightRAG
–
Qwen3-14B
1,024 / 4,096
256 / 1,024
Mem0
–
Qwen3-14B
1,024 / 4,096
256 / 1,024
Learned
MemAgent
Qwen3-8B writer
Self / Qwen3-14B
1,024 / 4,096
256 / 1,024
HiMPO
Qwen3-8B writer
Self / Qwen3-14B
1,024 / 4,096
256 / 1,024
Appendix
Table 5: Implementation settings of baselines.
Dataset
Cluster matching
SciREX
Greedy one-to-one matching by gold-span coverage ( ≥0.5 ), without an entity-type constraint.
AIPAN-10K
Same-type matching using exact normalized aliases for data types and token-F1 ( ≥0.5 ) for other entity types. Many-to-one matches are allowed.
Appendix
Table 6: Dataset-specific cluster matching for document-level evaluation.
Structured IE
SciREX
Typed Mention
Salient cluster
Binary relation
Binary relation
Method
P
R
F1
P
R
F1
P
R
F1
P
R
F1
Prior work
SciREX-P
50.3
90.2
64.6
17.6
78.4
28.8
2.8
58.2
5.4
6.5
41.1
9.6
TempGen
79.0
7.9
14.4
24.2
23.7
23.9
8.3
6.3
7.1
17.1
13.6
14.5
Direct Readout
Appendix
Table 7: Complete SciREX test results. Structured IE uses our document-level scorer, while SciREX reports binary relations using the unmodified TempGen evaluator ( Huang et al., 2021 ) . All non-prior-work results are mean@4.
Figure 6: Accuracy on BABILong QA1–QA5 across context lengths.
Figure 7: The linearized reward closely tracks exact F1 gain. Left: Exact document-level F1 gain versus summed Memory Gain on SciREX, computed from the same history-masked counterfactual counts. MGPO (blue; n=261 ) achieves Spearman ρ=0.988 and matching signs whenever both gains are nonzero; credit ablations and the base policy are shown in grey. Right: OLS slopes of exact on linearized gain with document-bootstrap 95% CIs ( B=2000 ). The surrogate slightly underestimates MGPO gains (1.16, 95% CI [1.09, 1.24]) and overestimates ablation loss magnitudes (0.77–0.86).
Figure 8: Noise in single-sample utility estimation. Empirical CDFs of ∣MGt∣ for unchanged-memory controls (grey) and actual rewrites (blue). The dashed line is the 95th percentile of the control distribution.
Memory
Chunk
Salient cluster
Binary relation
Memory budget, chunk fixed at 1024
64
1024
50.8 ± 0.4
26.1 ± 0.7
128
1024
51.3 ± 1.2
26.4 ± 1.6
256 †
1024 †
51.3 ± 0.9
26.5 ± 1.3
512
1024
51.5 ± 0.7
26.7 ± 0.8
Chunk size, memory fixed at 256
Appendix
Table 8: Sensitivity to memory budget and chunk size. We conduct experiments using the model trained with a configuration of a 1024 tokens chunk size and a 256 tokens memory budget. Performance is largely insensitive to the memory budget, while varying the chunk size shows a minor effect, peaking at 2048. † denotes the training configuration.
Method
G
Mem. tokens
Policy rollouts
Reader evaluations
GPU- hours
Peak mem. (GiB)
Cluster F1
Relation F1
MGPO-base
–
202
–
–
–
–
37.8 ± 0.4
15.5 ± 0.4
MGPO-GR
8
72
8
50.7
142.0
198.4
45.5 ± 0.5
20.1 ± 0.3
MGPO-EMA
1
43
1
8.3
26.9
192.5
51.3 ± 0.9
26.5 ± 1.3
Appendix
Table 9: Position EMA versus group-relative baseline on SciREX. Both variants use the same Memory Gain objective. MGPO-EMA uses position EMA ( G=1 ), whereas MGPO-GR uses GRPO ( G=8 ). Policy rollouts and reader evaluations are reported per training document. GPU-hours cover the complete training run, and peak memory is the analytical per-step training footprint.
Figure 9: Analytical FLOPs under the SciREX instrument, every call charged at its configured maximum width. Left: Inference: a full-context 14B reader scales quadratically in document length while chunk-wide MGPO scales linearly. Right: Training: Memory Gain (G=1) pays a triangular counterfactual reward matrix, GRPO pays G trajectories with diagonal reader cells. At the mean SciREX (train) document (dashed line, 6,346 tokens) Memory Gain costs half of GRPO-8, and only overtakes it past 16K tokens.
Figure 10: Training-time GPU memory breakdown on SciREX. All methods train the same Qwen3-8B policy under their respective training configurations. Policy memory is therefore identical, while rollout KV cache and activations vary with each method’s sequence and batch geometry. MGPO additionally maintains a frozen Qwen3-14B reader and its KV cache. Nevertheless, its chunk-wise training setup substantially reduces rollout and activation memory, yielding a total footprint comparable to MemAgent and lower than R1-RE.
May 1, 2026·Derong Xu, Shuochen Liu, Pengfei Luo +8Personalization
University of Science and Technology of China & State Key Laboratory of Cognitive Intelligence · City University of Hong Kong · Dalian University of Technology +1