Long-context reasoning is essential for complex and long-horizon tasks, yet the performance of large language models (LLMs) degrades as context length increases. Recent approaches address this by processing input chunk by chunk while maintaining a bounded textual memory in model context. However, premature information compression can discard critical details essential for subsequent reasoning. In this paper, we introduce Commit-on-Evidence Memory (CoEM), which learns when to convert source evidence into compact memory facts. Specifically, under a fixed context-memory budget, CoEM preserves potentially useful source excerpts verbatim in a pending set, allowing subsequent context to clarify their relevance before irreversible compression. As new context arrives, a learned policy revisits each pending excerpt and decides whether to promote it to the committed memory, retain it for further consideration, or discard it. A frozen verifier ensures proposed facts are accepted only if supported by retained excerpts and current context. To further guide effective memory management, we train this policy using reinforcement learning by combining fine-grained, step-level evidence rewards with final answer rewards. Extensive experiments demonstrate that CoEM consistently improves long-context reasoning. When evaluated on 6,400 documents long-context input, CoEM outperforms the strongest memory baseline by 10.4-11.4 F1 points on Qwen3.5-9B. Code repository: https://github.com/benmagnifico/CoEM.
Figures & tables
Figure 1: Methodology comparison of CoEM and two baselines on a HotpotQA question ( Yang et al., 2018 ) . CoEM does not make a decision about the unclear source evidence until its value is clarified by the later evidence.
Figure 2: Memory occupancy of different methods throughout reading procedure on HotpotQA.
Figure 3: Overview of CoEM . At intermediate steps t<T , the policy promotes, keeps, or drops evidence. A frozen verifier checks proposed facts before commitment. At the final step T , remaining pending evidence is resolved before the model answers from committed memory.
Figure 4: RL Training with step-level evidence rewards. Grouped rollouts combine final answer rewards with evidence return-to-go, carrying later source-verification feedback to earlier memory decisions. Group-centered advantages drive the clipped policy update.
Figure 5: Analysis of evidence retention, source verification, and inference efficiency with Qwen3.5-4B. In (a), Capacity events occur when admitting new evidence would exceed the pending-set budget. In (b), Forced drops discard the oldest pending evidence to free space. In (e), Self is the 4B policy self-check, Raw/Cal. are the raw/calibrated DeBERTaV3-small verifiers, and 9B is Qwen3.5-9B; the blue line uses the right axis. In (f), MA, GM, RR † , and RN denote MemAgent, GRU-Mem, budget-matched ReMemR1, and native ReMemR1.
Variant
HotpotQA
2Wiki.
F1 ↑
Faith. ↑
F1 ↑
Faith. ↑
No RL
67.5
93.2
55.1
92.7
Eager, no verifier
73.4
85.2
61.8
84.7
Eager + verifier
77.2
95.7
66.0
95.2
Pending, no verifier
81.5
86.7
71.3
86.2
Outcome-only RL
80.4
94.4
69.9
93.9
Table 2: Ablation analysis of CoEM . Faith. is the percentage of accepted facts supported by their cited sources.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Policy / verifier precision
bfloat16 / float32 logits for calibration
Chunk / total context memory
C=5,000 / B=1,024 tokens
Pending / committed allocations
BP=256 / BM=768 tokens
Question / instruction and feedback reserve
Lq=256 / 1,912 tokens
Input / controller output / final output
8,192 / Wπ=1,024 / Wans=256 tokens
Total serving window
9,216 tokens (input plus output)
Appendix
Table 3: Policy, controller, training, and evaluation settings.
Model
Dataset
ReMemR1 †
ReMemR1 (native)
CoEM
Qwen3.5-4B
HotpotQA
81.2
83.4
85.7
Qwen3.5-4B
2Wiki
70.9
73.0
76.4
Qwen3.5-4B
MuSiQue
61.4
63.3
67.0
Qwen3.5-9B
HotpotQA
84.2
86.4
89.6
Qwen3.5-9B
2Wiki
75.6
78.0
81.9
Qwen3.5-9B
MuSiQue
67.2
69.3
73.6
Appendix
Table 4: Length-averaged F1 (%). Native ReMemR1 retains historical snapshots outside the 1,024-token budget and is reported separately from the matched-storage comparison.
Configuration
ECE (%) ↓
NLL ↓
Brier ( ×100 ) ↓
Raw
7.10
0.52
8.00
Temperature scaling
3.20
0.49
7.40
Appendix
Table 5: Held-out calibration metrics. ECE uses ten equal-width bins; lower values are better.
Verifier
Faith. ↑
FAR ↓
FRR ↓
ms/pair ↓
No verification
64.0
100.0
0.0
0.0
Policy self-check (Qwen3.5-4B)
89.8
18.5
8.3
165.0
DeBERTaV3-small, raw
94.9
9.2
3.9
4.2
DeBERTaV3-small, calibrated
97.2
4.9
6.4
4.3
Qwen3.5-9B verifier
97.5
4.3
5.8
295.0
Appendix
Table 6: Independent audit of source-support verifiers on the same blinded set of 2,000 pre-verification promotion attempts. Faithfulness measures source support among accepted claims; FAR and FRR measure unsupported claims accepted and supported claims rejected, respectively. Latency is measured per source–claim pair on one H800 with batch size one.
Figure 6: Context-memory occupancy during the Qwen3.5-4B HotpotQA experiment with 6,400 documents and 128 chunks. The shaded band shows pending source text above committed memory; both components contribute to CoEM ’s total retained context memory. This enlarged view uses the same experimental trace as Figure 2 .
2Wiki condition
Mean / P95 tokens
Capacity events
Forced drops
Overflow
Ordinary ordering
142 / 223
2.8%
0.5%
0.0%
Gap 128
183 / 246
6.2%
1.8%
0.0%
Source-dense stress
224 / 254
19.7%
7.6%
0.0%
Appendix
Table 7: Pending-set stress diagnostics for Qwen3.5-4B on 2WikiMultiHopQA with 6,400 documents and BP=256 . Mean / P95 tokens report the mean and 95th-percentile occupancy of P . The capacity-event rate is the fraction of eligible admissions that require freeing space; the forced-drop rate counts evicted resident pending records per admitted record; overflow denotes token-budget violations.
Method
800 docs
1,600 docs
3,200 docs
6,400 docs
MemAgent
84
171
351
720
GRU-Mem
55
112
231
474
ReMemR1 †
102
207
425
870
ReMemR1 (native)
108
221
459
951
CoEM
63
130
267
548
Appendix
Table 8: End-to-end latency (s/question) for the Qwen3.5-4B policies on one H800. Full-scan timing includes all controller operations, verification, memory callback, and final-answer generation.
Figure 7: Training diagnostics for Qwen3.5-4B. (a) Training answer F1. (b) Development answer F1, evaluated every 25 updates; stars mark the maxima of the mean curves, rather than the checkpoints selected separately for each run. (c) Verifier-rejection rate among proposals passing all non-entailment checks. (d) Invalid-operation rate among all proposal attempts. Lines show three-seed means and shaded bands show one sample standard deviation across policy seeds.
Figure 8: Memory decisions during training. (a) Step-level evidence reward. (b) Fact proposals checked by the verifier. (c) Proposals passing verification. (d) Voluntary Drop actions. Panels (b)–(d) report counts per reading chunk. Lines show three-seed means and shaded bands show one sample standard deviation across policy seeds. Passing verification does not imply that the corresponding atomic transaction is committed. Evidence rewards for Outcome-only RL are logged but do not affect its updates.
Setting
Comparator
CoEM
Comparator
Δ [95% CI]
4B / HotpotQA / 800
Outcome-only RL
85.5±0.46
80.4±0.85
5.1 [3.3, 7.0]
4B / 2Wiki / 800
Outcome-only RL
76.1±0.56
69.9±0.85
6.2 [4.3, 8.2]
9B / HotpotQA / 6,400
ReMemR1 †
89.1±0.40
78.7±0.75
10.4 [7.6, 13.4]
9B / 2Wiki / 6,400
ReMemR1 †
81.0±0.56
69.6±0.75
11.4 [8.7, 14.3]
9B / MuSiQue / 6,400
ReMemR1 †
72.3±0.66
61.6±0.85
10.7 [8.2, 13.4]
Appendix
Table 9: Uncertainty summaries for key comparisons. Settings list backbone / dataset / number of documents. F1 entries are mean ± SD over policy seeds 17, 29, and 43; Δ is CoEM minus comparator F1. Each setting contains 128 paired questions, and the 95% interval for Δ uses 10,000 question-level bootstrap resamples that preserve method and seed pairing. † denotes budget-matched ReMemR1.
Figure 9: Supporting-evidence diagnostics for Qwen3.5-4B on HotpotQA. (a) Recoverable supporting relations at 32 normalized reading positions, before final resolution. (b) Pending-source survival immediately before bridge arrival under the existing gap intervention; labeled gap settings are equally spaced. (c) Commitment timing or removal for initially admitted records containing supporting evidence at gap 32. Within 1 step denotes commitment at the bridge step or the following step. Shaded bands in (a) and (b) show one sample standard deviation across policy seeds; panel (c) reports pooled record proportions.
Training objective
α
Evidence reward
HotpotQA F1
2Wiki F1
Outcome-only RL
1.0
None
80.4±0.85
69.9±0.85
Matched α , no evidence
0.8
None
81.0±0.70
70.5±0.80
Immediate evidence reward
0.8
Current chunk
83.1±0.60
73.2±0.70
Full CoEM
0.8
Return-to-go
85.5±0.46
76.1±0.56
Appendix
Table 10: Reward controls with Qwen3.5-4B at 800 documents (mean ± SD over three seeds). α is the answer-advantage weight; matched controls fix α=0.8 and the KL coefficient at 10−3 , and both evidence variants share the same 1/T scaling and penalty weights.
Policy
GPUs/run
Hours/run
GPU-h/run
GPU-h/3 seeds
Qwen3.5-4B
8
18.0
144
432
Qwen3.5-9B
8
42.0
336
1,008
Total for the six runs
—
—
—
1,440
Appendix
Table 11: Training resources for CoEM on nodes with eight H800 GPUs.
Resource or operation
Value
Regular / additional / answer policy calls
128 / 12 / 1
Total policy calls
141
Policy input / generated tokens
1,020,672 / 21,432
Final-answer tokens (included above)
52
Verifier pairs / processed pair tokens
360 / 64,800
Policy / verifier / controller time (s)
540.300 / 1.548 / 6.152
Appendix
Table 12: Inference resources. Additional calls cover capacity handling and final resolution; verifier memory is included in peak memory.
Figure 10: Source admission and retention before an alias relation makes the biography relevant to the question.
Figure 11: Verified commitment after the alias arrives. Each claim is checked against an available source, and the three entries are inserted as one atomic transaction.
Figure 12: Capacity pressure develops before the cast relation arrives. The new source passes its individual size check, but admitting it requires resolution of an existing pending record.
Figure 13: A promotion can fail the committed-memory budget check even when its source states the proposed fact. Forced removal then makes the date unavailable when the later cast relation identifies the relevant person.