Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each turn, a CCM agent emits an updated memory together with an environment action; its next prompt contains the original task, retained memory, and newest observation rather than the complete transcript. We first evaluate CCM without fine-tuning on TerminalBench-2 using Claude Sonnet 4.6, Claude Opus 4.6, GLM-5, and Kimi K3. CCM substantially reduces cumulative input usage and active-prompt size, although it lowers task success for most models while preserving performance for Kimi K3. We use GRPO with privileged full-history distillation to improve CCM in open-weight models. A frozen copy of the student's initial model scores each sampled student action under the complete history reconstructed from that student's rollout, providing dense action-token supervision without a separate teacher rollout or reference solution. On WebShop, this objective substantially improves CCM over GRPO at both evaluated model scales and surpasses full-history GRPO for Qwen3-4B-Instruct, though not for Qwen3-8B. On Endless Terminals, the augmented method provides a modest improvement over GRPO, with both CCM policies outperforming the untrained full-history baseline. These results demonstrate that CCM is a viable inference paradigm for agents operating with substantially reduced retained context and that its performance can be improved through reinforcement learning with privileged full-history distillation.
Figures & tables
Peak prompt tokens
Model
Mode
Success, n (%)
Median
Maximum
Maximum retained memory
Estimated cost per task
Claude Sonnet 4.6
Baseline
38 (42.7)
16,673
131,157
–
$0.468
CCM
28 (31.5)
3,568
50,614
1,219
$0.472
Claude Opus 4.6
Baseline
52 (58.4)
15,405
103,512
–
$0.831
CCM
36 (40.4)
4,280
51,457
1,403
$1.087
GLM-5
Baseline
31 (34.8)
17,560
116,407
–
$0.515
Table 1: Terminal-Bench 2.0 pass@1 results over 89 tasks. Costs are estimated means per task and include prompt caching where available; prompt caching was unavailable for GLM-5 through Amazon Bedrock.
Best performance
Memory behavior
Memory size
Peak prompt tokens
Condition
Exact success (%)
Step
Update rate (%)
Prefix retention (%)
Mean
Maximum
Median
Maximum
Qwen3-4B-Instruct
Untrained full history
3.91
–
–
–
–
–
4,096
4,096
Full-history GRPO
81.25
70
–
–
–
–
1,721
3,076
Untrained CCM
0.78
–
57.6
69.6
225.7
861
1,232
1,908
CCM + GRPO
32.81
60
12.9
98.5
95.6
147
1,045
1,327
Table 2: Best observed task success and context behavior on the fixed 128-task WebShop evaluation. Prompts are truncated at the 4,096-token model-input limit.
Best performance
Memory behavior
Memory size (tokens)
Condition
Success (%)
Step
Update rate (%)
Prefix retention (%)
Mean
Maximum
Untrained full history
20.33
–
–
–
–
–
Full-history GRPO
35.33
30
–
–
–
–
Untrained CCM
21.00
–
34.8
81.2
63.9
346
CCM + GRPO
29.00
30
16.9
94.6
77.1
720
CCM + GRPO + distillation
32.00
50
36.9
72.9
144.0
906
Table 3: Best observed task success and memory behavior through training step 60 on the fixed 300-task Endless Terminals evaluation with Qwen3-8B. For each trained method, we report the evaluated checkpoint with the highest success, breaking ties in favor of the earlier checkpoint.
Method
MMLU-Pro
HellaSwag
IFEval strict
WebShop, Qwen3-4B-Instruct
Pretrained base
65.25±0.66
80.14±0.07
83.30±0.56
Full-history GRPO
62.48±0.49(−2.77)
67.92±0.42(−12.22)
80.35±0.65(−2.96)
CCM + GRPO
55.16±0.60(−10.09)
66.68±0.71(−13.46)
82.32±0.47(−0.99)
CCM + GRPO + distillation
64.23±0.25(−1.03)
74.53±0.20(−5.62)
83.18±0.49(−0.12)
WebShop, Qwen3-8B
Table 4: General-capability retention after agent post-training. Values are mean percentage scores over three decoding seeds, with sample standard deviations. Parentheses report absolute percentage-point changes from the corresponding pretrained base model, computed before rounding.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Claude Sonnet 4.6
Claude Opus 4.6
GLM-5
Kimi K3
Metric
Baseline
CCM
Baseline
CCM
Baseline
CCM
Baseline
CCM
Evaluation
Success, n (%)
38 (42.7)
28 (31.5)
52 (58.4)
36 (40.4)
31 (34.8)
26 (29.2)
56 (62.9)
57 (64.0)
Average turns
20.88
30.33
21.60
26.00
25.90
32.50
14.49
19.02
Tokens per task
Uncached input
14,955
48,056
23,914
86,052
480,325
75,481
102
27,039
Appendix
Table 5: Detailed Terminal-Bench 2.0 pass@1 and mean cumulative token usage over 89 tasks. Baseline denotes full-history prompting.
Hyperparameter
WebShop
Endless Terminals
Training steps
100
60
Tasks per step
32
32
Rollouts per task
8
8
Trajectories per step
256
256
PPO mini-batch size
32
32
Micro-batch size per GPU
1
1
Appendix
Table 6: Optimization hyperparameters used for WebShop and Endless Terminals.
Model
Success (%)
Solved
Partial score (%)
Turns
Qwen3-4B-Instruct
0.00
0/128
0.59
14.97
Qwen3-8B
0.78
1/128
5.27
14.11
Appendix
Table 7: Distillation-only ablation on the fixed 128-task WebShop evaluation. Both models are evaluated at checkpoint 100 under CCM. Partial score is the mean percentage of WebShop constraints satisfied.
Hyperparameter
WebShop
Endless Terminals
Maximum turns
15
16
Sampling temperature
1.0
0.6
Top- p
1.0
1.0
Top- k
Unrestricted
Unrestricted
Student prompt limit
4,096
16,384
Teacher prompt limit
32,768
16,384
Appendix
Table 8: Rollout and sequence-length configuration. WebShop uses 1,024 generation tokens for Qwen3-4B-Instruct and 2,048 for Qwen3-8B.
Model
Condition
Exact
Partial
Turns
Qwen3-4B-Instruct
Untrained full history
No
0.000
15
Untrained CCM
No
0.000
15
CCM + GRPO
No
0.857
15
CCM + GRPO + Distillation
Yes
1.000
5
Qwen3-8B
Untrained full history
No
0.000
12
Untrained CCM
No
0.000
15
Appendix
Table 9: Outcomes for the matched WebShop trajectory. Exact denotes exact task success; partial is the WebShop constraint-matching score.
Task
Condition
Success
Turns
4d41da7b
Untrained full history
No
2
Untrained CCM
No
2
CCM + GRPO
No
16
CCM + GRPO + Distillation
Yes
2
beff73f4
Untrained full history
Yes
4
Untrained CCM
No
16
Appendix
Table 10: Outcomes for the two matched Endless Terminals trajectories.