Organizations: Shanghai Academy of AI for Science · Fudan University · Shanghai Jiao Tong University · University of Michigan · The Chinese University of Hong Kong · Alibaba Group · Nanjing University · Shanghai Innovation Institute
Long-horizon agents compact their history to continue within a finite context window, but a textual summary alone may not support every subsequent decision. We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens. Given the history and summary, the network learns to generate tokens that help a frozen LLM approximate the continuation it would produce with the full history. The tokens are conditioned on the summary and appended after it, forming an analogue of a residual connection along the sequence dimension. On SummHay, REMORY improves source attribution at nearly unchanged insight coverage and approaches the full-context joint score using only 5.2% of the input positions. Across long-horizon agent benchmarks, Qwen3.8-27B and GLM-5.3-Flash show consistent gains with residual memory. Both models also exhibit substantially fewer repeated tool outputs and tool errors on BrowseComp and Terminal-Bench 2.1.
Figures & tables
Figure 1: BrowseComp and Terminal-Bench 2.1 scores. Pale and solid bars compare the same local actor without and with residual memory. Gray bars are published frontier references, with their original harnesses and protocols ( Google DeepMind, 2026 ; Terminal-Bench Team, 2026 ; OpenAI, 2026b ) ; they are not controlled comparisons with our runs.
Figure 2: Context compaction and human memory. (a) Input length across successive requests in an agent run, showing repeated growth and sharp reductions. (b) Memory systems illustrated using the Harvard–Oxford atlas ( Harvard–Oxford Atlas Contributors, 2026 ) : prefrontal regions (blue), hippocampus (cyan), and association cortex (purple).
Figure 3: Residual compensation and training. (a) Codex CLI passes the /compact summary to the next actor call. (b) The Remory network uses that summary to condition memory construction from the active context; summary and continuous memory jointly support subsequent actions. Archive access remains an ordinary tool action in either setting. (c,d) Qwen3.8-27B teacher–student top-1 agreement during Stage I reconstruction and Stage II residual training. The residual panel compares reconstruction initialization with random initialization under the same residual objective and data order. Curves are optimization diagnostics from the training streams.
Representation
Coverage
Citation F1
Joint score
Raw context
72.46
56.60
43.41
Summary
67.55
56.98
39.84
Summary++
67.52
56.94
39.78
Summary + Remory
67.95
61.03
43.39
Table 1: SummHay on the same 92 queries. Summary++ adds text under the residual memory’s additional-position budget. Higher is better; all queries, including generation and formatting failures, are retained.
AutomationBench
JobBench
Configuration
Score
Cost
Score
Cost
w/o residual
35.5 ± 1.3
$0.59
33.4 ± 1.25
$3.09
w/ residual
45.3 ± 1.0
$0.53
41.0 ± 1.41
$3.21
Table 2: Qwen3.8-27B workflow results: mean ± standard deviation over five runs. Costs are mean API-equivalent USD per task.
Terminal-Bench 2.1
BrowseComp
Actor
Residual
Score
Cost
Repeats
Errors
Score
Cost
Repeats
Errors
Qwen3.8-27B
w/o
71.9
$4.29
459
933
74.0
$5.39
25,801
25,244
w/
76.4
$2.86
377
568
77.0
$4.69
18,105
12,698
GLM-5.3-Flash
w/o
84.3
$3.27
334
664
84.9
$5.65
8,316
7,498
w/
87.6
$3.05
240
496
89.0
$5.00
5,834
5,412
Table 3: Terminal-Bench 2.1 and BrowseComp: one run per configuration. Costs are mean API-equivalent USD per task. Repeats and errors are tool-output counts over the full task set.
Figure 4: Live Trading context use. (a) An illustrative September 22 trace. (b) Compaction intervals in the separate September 20–23 evaluation: 56/45 completed compactions and 51/41 within-session intervals without/with residual memory. Intervals do not cross restarts; two residual intervals exceed the display range but remain in the box statistics.
Figure 5: Cell segmentation after repair. The task image and the exact overlay viewed by the actor are reproduced without pixel changes.
Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each turn, a CCM agent emits an updated memory together with an environment action; its next prompt contains the original task, retained memory, and newest observation rather than the complete transcript. We first evaluate CCM without fine-tuning on TerminalBench-2 using Claude Sonnet 4.6, Claude Opus 4.6, GLM-5, and Kimi K3. CCM substantially reduces cumulative input usage and active-prompt size, although it lowers task success for most models while preserving performance for Kimi K3. We use GRPO with privileged full-history distillation to improve CCM in open-weight models. A frozen copy of the student's initial model scores each sampled student action under the complete history reconstructed from that student's rollout, providing dense action-token supervision without a separate teacher rollout or reference solution. On WebShop, this objective substantially improves CCM over GRPO at both evaluated model scales and surpasses full-history GRPO for Qwen3-4B-Instruct, though not for Qwen3-8B. On Endless Terminals, the augmented method provides a modest improvement over GRPO, with both CCM policies outperforming the untrained full-history baseline. These results demonstrate that CCM is a viable inference paradigm for agents operating with substantially reduced retained context and that its performance can be improved through reinforcement learning with privileged full-history distillation.
Long-context tasks require LLMs to identify and preserve answer-relevant information from large contexts. Chunk-wise memory agents address this issue by sequentially reading document chunks, updating a compact memory, and generating the final answer from the accumulated memory. However, existing RL-based chunk-wise agents either rely on sparse final-answer rewards or use lexical intermediate rewards for memory and retrieval actions. These signals supervise task success or local overlap, but do not directly evaluate whether the final memory supports the ground-truth answer. We propose InfoMem, a reward mechanism for training chunk-wise memory agents that evaluates final-memory utility using answer-conditioned information. InfoMem measures how much the final memory increases the model's per-token log-likelihood of the ground-truth answer. To stabilize RL optimization, InfoMem applies this signal only to successful trajectories and normalizes it before reward composition. Under the same GRPO framework and training budget, InfoMem improves long-context memory-agent performance over comparable memory-agent RL baselines. Analyses show that effective final-memory rewards should operate on successful trajectories, be normalized before reward composition, and be conditioned on the answer rather than the query. Our code is available at https://github.com/GenSouKa1/InfoMem.
Tiancheng Han, Yong Li, Wuzhou Yu +2
1Tongji University · 2Shanghai Innovation Institute · 3Shanghai AI Laboratory
Memory is an indispensable capability for long-horizon LLM agents, enabling them to preserve and utilize information accumulated across extended interactions. Existing memory-agent approaches are typically trained end-to-end with reinforcement learning on downstream tasks. However, collecting high-quality annotated problems for memory-intensive scenarios is costly, and the resulting training data often lack sufficient diversity to cover general memory behaviors. In this work, we propose MemTrain, a self-supervised training framework for generally enhancing the context-memory capability of LLM agents for more effective downstream post-training. MemTrain introduces two coupled proxy tasks over unlabeled Wikipedia corpora: (1) an end-to-end masked reconstruction objective, which requires the model to recover masked entities after multiple rounds of memory updates, thereby encouraging memory maintenance from the final outcome perspective; and (2) an intermediate memory recall objective, which requires the model to reconstruct masked historical information using intermediate memory states, encouraging faithful compression and memory completeness throughout the interaction process. The two objectives are jointly optimized using GRPO. Extensive experiments on long-text QA and search-based QA benchmarks demonstrate that MemTrain consistently improves downstream memory-intensive reasoning performance across different models, achieving gains of up to 17.67 points over direct task-specific post-training.
Ziheng Li, Xingrun Xing, Haoqing Wang +2
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · Samsung Research, Beijing, China