Organizations: Shanghai Academy of AI for Science · Fudan University · Shanghai Jiao Tong University · University of Michigan · The Chinese University of Hong Kong · Alibaba Group · Nanjing University · Shanghai Innovation Institute
Long-horizon agents compact their history to continue within a finite context window, but a textual summary alone may not support every subsequent decision. We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens. Given the history and summary, the network learns to generate tokens that help a frozen LLM approximate the continuation it would produce with the full history. The tokens are conditioned on the summary and appended after it, forming an analogue of a residual connection along the sequence dimension. On SummHay, REMORY improves source attribution at nearly unchanged insight coverage and approaches the full-context joint score using only 5.2% of the input positions. Across long-horizon agent benchmarks, Qwen3.8-27B and GLM-5.3-Flash show consistent gains with residual memory. Both models also exhibit substantially fewer repeated tool outputs and tool errors on BrowseComp and Terminal-Bench 2.1.
Figures & tables
Figure 1: BrowseComp and Terminal-Bench 2.1 scores. Pale and solid bars compare the same local actor without and with residual memory. Gray bars are published frontier references, with their original harnesses and protocols ( Google DeepMind, 2026 ; Terminal-Bench Team, 2026 ; OpenAI, 2026b ) ; they are not controlled comparisons with our runs.
Figure 2: Context compaction and human memory. (a) Input length across successive requests in an agent run, showing repeated growth and sharp reductions. (b) Memory systems illustrated using the Harvard–Oxford atlas ( Harvard–Oxford Atlas Contributors, 2026 ) : prefrontal regions (blue), hippocampus (cyan), and association cortex (purple).
Figure 3: Residual compensation and training. (a) Codex CLI passes the /compact summary to the next actor call. (b) The Remory network uses that summary to condition memory construction from the active context; summary and continuous memory jointly support subsequent actions. Archive access remains an ordinary tool action in either setting. (c,d) Qwen3.8-27B teacher–student top-1 agreement during Stage I reconstruction and Stage II residual training. The residual panel compares reconstruction initialization with random initialization under the same residual objective and data order. Curves are optimization diagnostics from the training streams.
Representation
Coverage
Citation F1
Joint score
Raw context
72.46
56.60
43.41
Summary
67.55
56.98
39.84
Summary++
67.52
56.94
39.78
Summary + Remory
67.95
61.03
43.39
Table 1: SummHay on the same 92 queries. Summary++ adds text under the residual memory’s additional-position budget. Higher is better; all queries, including generation and formatting failures, are retained.
AutomationBench
JobBench
Configuration
Score
Cost
Score
Cost
w/o residual
35.5 ± 1.3
$0.59
33.4 ± 1.25
$3.09
w/ residual
45.3 ± 1.0
$0.53
41.0 ± 1.41
$3.21
Table 2: Qwen3.8-27B workflow results: mean ± standard deviation over five runs. Costs are mean API-equivalent USD per task.
Terminal-Bench 2.1
BrowseComp
Actor
Residual
Score
Cost
Repeats
Errors
Score
Cost
Repeats
Errors
Qwen3.8-27B
w/o
71.9
$4.29
459
933
74.0
$5.39
25,801
25,244
w/
76.4
$2.86
377
568
77.0
$4.69
18,105
12,698
GLM-5.3-Flash
w/o
84.3
$3.27
334
664
84.9
$5.65
8,316
7,498
w/
87.6
$3.05
240
496
89.0
$5.00
5,834
5,412
Table 3: Terminal-Bench 2.1 and BrowseComp: one run per configuration. Costs are mean API-equivalent USD per task. Repeats and errors are tool-output counts over the full task set.
Figure 4: Live Trading context use. (a) An illustrative September 22 trace. (b) Compaction intervals in the separate September 20–23 evaluation: 56/45 completed compactions and 51/41 within-session intervals without/with residual memory. Intervals do not cross restarts; two residual intervals exceed the display range but remain in the box statistics.
Figure 5: Cell segmentation after repair. The task image and the exact overlay viewed by the actor are reproduced without pixel changes.
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · Samsung Research, Beijing, China