Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each turn, a CCM agent emits an updated memory together with an environment action; its next prompt contains the original task, retained memory, and newest observation rather than the complete transcript. We first evaluate CCM without fine-tuning on TerminalBench-2 using Claude Sonnet 4.6, Claude Opus 4.6, GLM-5, and Kimi K3. CCM substantially reduces cumulative input usage and active-prompt size, although it lowers task success for most models while preserving performance for Kimi K3. We use GRPO with privileged full-history distillation to improve CCM in open-weight models. A frozen copy of the student's initial model scores each sampled student action under the complete history reconstructed from that student's rollout, providing dense action-token supervision without a separate teacher rollout or reference solution. On WebShop, this objective substantially improves CCM over GRPO at both evaluated model scales and surpasses full-history GRPO for Qwen3-4B-Instruct, though not for Qwen3-8B. On Endless Terminals, the augmented method provides a modest improvement over GRPO, with both CCM policies outperforming the untrained full-history baseline. These results demonstrate that CCM is a viable inference paradigm for agents operating with substantially reduced retained context and that its performance can be improved through reinforcement learning with privileged full-history distillation.
Figures & tables
Peak prompt tokens
Model
Mode
Success, n (%)
Median
Maximum
Maximum retained memory
Estimated cost per task
Claude Sonnet 4.6
Baseline
38 (42.7)
16,673
131,157
–
$0.468
CCM
28 (31.5)
3,568
50,614
1,219
$0.472
Claude Opus 4.6
Baseline
52 (58.4)
15,405
103,512
–
$0.831
CCM
36 (40.4)
4,280
51,457
1,403
$1.087
GLM-5
Baseline
31 (34.8)
17,560
116,407
–
$0.515
Table 1: Terminal-Bench 2.0 pass@1 results over 89 tasks. Costs are estimated means per task and include prompt caching where available; prompt caching was unavailable for GLM-5 through Amazon Bedrock.
Best performance
Memory behavior
Memory size
Peak prompt tokens
Condition
Exact success (%)
Step
Update rate (%)
Prefix retention (%)
Mean
Maximum
Median
Maximum
Qwen3-4B-Instruct
Untrained full history
3.91
–
–
–
–
–
4,096
4,096
Full-history GRPO
81.25
70
–
–
–
–
1,721
3,076
Untrained CCM
0.78
–
57.6
69.6
225.7
861
1,232
1,908
CCM + GRPO
32.81
60
12.9
98.5
95.6
147
1,045
1,327
Table 2: Best observed task success and context behavior on the fixed 128-task WebShop evaluation. Prompts are truncated at the 4,096-token model-input limit.
Best performance
Memory behavior
Memory size (tokens)
Condition
Success (%)
Step
Update rate (%)
Prefix retention (%)
Mean
Maximum
Untrained full history
20.33
–
–
–
–
–
Full-history GRPO
35.33
30
–
–
–
–
Untrained CCM
21.00
–
34.8
81.2
63.9
346
CCM + GRPO
29.00
30
16.9
94.6
77.1
720
CCM + GRPO + distillation
32.00
50
36.9
72.9
144.0
906
Table 3: Best observed task success and memory behavior through training step 60 on the fixed 300-task Endless Terminals evaluation with Qwen3-8B. For each trained method, we report the evaluated checkpoint with the highest success, breaking ties in favor of the earlier checkpoint.
Method
MMLU-Pro
HellaSwag
IFEval strict
WebShop, Qwen3-4B-Instruct
Pretrained base
65.25±0.66
80.14±0.07
83.30±0.56
Full-history GRPO
62.48±0.49(−2.77)
67.92±0.42(−12.22)
80.35±0.65(−2.96)
CCM + GRPO
55.16±0.60(−10.09)
66.68±0.71(−13.46)
82.32±0.47(−0.99)
CCM + GRPO + distillation
64.23±0.25(−1.03)
74.53±0.20(−5.62)
83.18±0.49(−0.12)
WebShop, Qwen3-8B
Table 4: General-capability retention after agent post-training. Values are mean percentage scores over three decoding seeds, with sample standard deviations. Parentheses report absolute percentage-point changes from the corresponding pretrained base model, computed before rounding.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Claude Sonnet 4.6
Claude Opus 4.6
GLM-5
Kimi K3
Metric
Baseline
CCM
Baseline
CCM
Baseline
CCM
Baseline
CCM
Evaluation
Success, n (%)
38 (42.7)
28 (31.5)
52 (58.4)
36 (40.4)
31 (34.8)
26 (29.2)
56 (62.9)
57 (64.0)
Average turns
20.88
30.33
21.60
26.00
25.90
32.50
14.49
19.02
Tokens per task
Uncached input
14,955
48,056
23,914
86,052
480,325
75,481
102
27,039
Appendix
Table 5: Detailed Terminal-Bench 2.0 pass@1 and mean cumulative token usage over 89 tasks. Baseline denotes full-history prompting.
Hyperparameter
WebShop
Endless Terminals
Training steps
100
60
Tasks per step
32
32
Rollouts per task
8
8
Trajectories per step
256
256
PPO mini-batch size
32
32
Micro-batch size per GPU
1
1
Appendix
Table 6: Optimization hyperparameters used for WebShop and Endless Terminals.
Model
Success (%)
Solved
Partial score (%)
Turns
Qwen3-4B-Instruct
0.00
0/128
0.59
14.97
Qwen3-8B
0.78
1/128
5.27
14.11
Appendix
Table 7: Distillation-only ablation on the fixed 128-task WebShop evaluation. Both models are evaluated at checkpoint 100 under CCM. Partial score is the mean percentage of WebShop constraints satisfied.
Hyperparameter
WebShop
Endless Terminals
Maximum turns
15
16
Sampling temperature
1.0
0.6
Top- p
1.0
1.0
Top- k
Unrestricted
Unrestricted
Student prompt limit
4,096
16,384
Teacher prompt limit
32,768
16,384
Appendix
Table 8: Rollout and sequence-length configuration. WebShop uses 1,024 generation tokens for Qwen3-4B-Instruct and 2,048 for Qwen3-8B.
Model
Condition
Exact
Partial
Turns
Qwen3-4B-Instruct
Untrained full history
No
0.000
15
Untrained CCM
No
0.000
15
CCM + GRPO
No
0.857
15
CCM + GRPO + Distillation
Yes
1.000
5
Qwen3-8B
Untrained full history
No
0.000
12
Untrained CCM
No
0.000
15
Appendix
Table 9: Outcomes for the matched WebShop trajectory. Exact denotes exact task success; partial is the WebShop constraint-matching score.
Task
Condition
Success
Turns
4d41da7b
Untrained full history
No
2
Untrained CCM
No
2
CCM + GRPO
No
16
CCM + GRPO + Distillation
Yes
2
beff73f4
Untrained full history
Yes
4
Untrained CCM
No
16
Appendix
Table 10: Outcomes for the two matched Endless Terminals trajectories.
LLM agents increasingly face long-horizon tasks such as web search and deep research in real-world applications, where accumulated context can cause long-context degradation and reasoning failures. Prior work mitigates this through context management with agent-side context control or fixed strategies such as summarization, which require training the agent itself for adaptation - making it impractical for closed-source agents and ignoring that different agents may require different strategies. We introduce Adaptive Context Management (AdaCoM), which trains an external LLM to manage the context of a frozen agent through flexible modification actions and end-to-end reinforcement learning. Across diverse agents on web search and deep research benchmarks, AdaCoM substantially improves performance by preserving task constraints and progress while pruning stale content. The learned strategies reveal a Fidelity-Reliability Trade-off: agents with higher vanilla ReAct performance benefit from higher-fidelity context preservation, whereas lower-performing agents require more aggressive compression to stay within a reliable reasoning regime. Transfer experiments show that AdaCoM generalizes most effectively across agents with similar capability (measured by vanilla ReAct performance), suggesting a practical path toward reusable context managers for agent systems.
Lu Yi, Runlin Lei, Liuyi Yao +6
1Renmin University of China · Work done during internship at Tongyi Lab, Alibaba Group · 2Tongyi Lab, Alibaba Group +2
We present Context Window Lifecycle (CWL), a context-management scheme that gives long-horizon LLM agents an effectively unbounded working horizon. As a session accumulates history, CWL keeps the context within budget through graduated, semantically-aware eviction: the agent annotates its trajectory as typed, dependency-linked episodes as work proceeds, and a deterministic, LLM-free policy evicts content in priority order within that structure when a token budget is exceeded. CWL preserves user turns and the exploratory context the agent is actively reasoning over, while aggressively shedding action episodes whose effects are already persisted in the environment, keeping active context near a stable ceiling that also avoids the performance degradation associated with very large prompts. Compared to summarization-based compaction, CWL avoids four well-known limitations: unpredictable lossiness, destruction of causal structure, blocking model cost, and compression-induced hallucination. Compared to recency truncation, CWL is semantically aware: it drops the oldest-and-most-recoverable content according to the dependency graph rather than oldest-in-time regardless of relevance. We describe the annotation protocol, the episode graph, the eviction policy, and the token-accounting loop, and evaluate CWL on long-horizon agentic benchmarks: a single agent session completing 89 sequential tasks across 80 million tokens with no measurable degradation in task accuracy relative to per-task isolated sessions
Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents. This gap is reflected in their limited performance on continual learning benchmarks such as ARC-AGI-3, especially when models are evaluated out of the box. Various agent harnesses have been proposed to close this gap, and each commits to a strategy for handling long sequences of observations, i.e., what information to save from the environment and how to load it into model context, a choice we argue is particularly consequential. Existing methods for context management face a significant tradeoff, as preserving more information makes retrieving relevant details less tractable. We propose PRO-LONG, a minimal context management framework built around programmatic memory for LLM agents in long-horizon, exploratory settings. PRO-LONG addresses the tradeoff by keeping a complete, structured interaction log and capitalizing on recent progress in coding agents to search this history efficiently. On the full ARC-AGI-3 public game set, PRO-LONG improves over a base coding agent by an average of 18.0 percentage points across frontier models, and matches or exceeds state-of-the-art specialized harnesses (up to 76.1% pass@1) while using 4.2-5.8x fewer tokens. With Fable 5, PRO-LONG achieves 97.4% best@2 at a total cost of $1,750. Relevant code and logs are available at https://github.com/alexisfox7/PRO-LONG.