Adapting Context Compression for Long-Horizon Agents with Counterfactual Continuations
Organizations: University of Virginia · Nokia
Abstract
Long-horizon agents require context compression to manage growing interaction histories. Compression quality, however, is ultimately determined by downstream execution. Existing prompt-adaptation methods infer compression errors by comparing full-context and compressed trajectories. Such comparisons cannot isolate individual compressions and are confounded by agent stochasticity. We first find that compression degrades reliability before solvability. Using matched counterfactual continuations that compare execution from the same agent state with versus without compression, we further show that severe degradation concentrates at isolated compression events. Motivated by this finding, we propose PAIR (Prompt Adaptation using Interventional Rollouts) for adapting structured compression prompts. PAIR identifies individual compressions that degrade subsequent execution, diagnoses their effects, and revises the relevant sections of a fixed compression template. PAIR achieves the strongest cross-run reliability among compressed methods in every main benchmark-scope combination, consistently exceeding the competing prompt-adaptation baseline. Without modifying the downstream agent, PAIR brings compressed execution close to the no-compression baseline and sometimes numerically exceeds it.
Figures & tables
| Method | Average (168) | Easy (57) | Medium (48) | Hard (63) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Acc. | Pass 3 | Steps | Peak | Acc. | Pass 3 | Steps | Acc. | Pass 3 | Steps | Acc. | Pass 3 | Steps | |
| Agent: GPT-5.6 Luna / Compressor: GPT-5.6 Luna | |||||||||||||
| No compression | 84.1 0.9 | 78.6 | 12.1 | 12.53 | 97.7 1.0 | 93.0 | 9.6 | 86.8 2.4 | 81.2 | 12.2 | 69.8 0.0 | 63.5 | 14.4 |
| History-Only Compression | |||||||||||||
| FIFO | 53.2 1.2 | 45.2 | 30.0 | 7.14 | 95.9 2.0 | 93.0 | 9.6 | 47.9 4.2 | 37.5 | 33.7 | 18.5 3.3 | 9.5 | 42.8 |
| LLMLingua-2 | 71.4 2.3 | 61.3 | 19.2 | 10.53 | 97.7 2.7 | 94.7 | 8.7 | 77.8 2.4 | 66.7 | 17.6 | 42.9 6.3 | 27.0 | 30.0 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Description |
|---|---|
| Recurrent Context Compression | |
| Frozen downstream agent, including its language model, system prompt, action interface, output parser, and decoding procedure | |
| Interactive environment | |
| Task instruction sampled from task distribution | |
| Agent rollout terminating at step | |
| Complete interaction history before step , | |
| Method | Average (95) | 1-APP (42) | 2-APP (22) | 3-APP (31) | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Acc. | Pass 3 | Steps | Acc. | Pass 3 | Steps | Acc. | Pass 3 | Steps | Acc. | Pass 3 | Steps | |
| Agent: GPT-5.6 Luna / Compressor: GPT-5.6 Luna | ||||||||||||
| No compression | 82.1 4.8 | 75.8 | 11.2 | 92.9 4.1 | 88.1 | 7.6 | 81.8 4.5 | 77.3 | 10.9 | 67.7 8.5 | 58.1 | 16.2 |
| History-Only Compression | ||||||||||||
| FIFO | 62.1 1.8 | 54.7 | 23.2 | 81.7 1.4 | 73.8 | 10.3 | 56.1 2.6 | 50.0 | 28.2 | 39.8 3.7 | 32.3 | 37.0 |
| LLMLingua-2 | 76.1 1.2 | 68.4 | 14.4 | 89.7 1.4 | 83.3 | 8.0 | 81.8 0.0 | 77.3 | 14.0 | 53.8 4.9 | 41.9 | 23.2 |
| Method | EM | EM 3 | F1 | Steps | Peak | Total |
|---|---|---|---|---|---|---|
| No compression | 40.9 1.8 | 29.8 | 54.2 | 18.5 | 14.53 | 181.83 |
| History-Only Compression | ||||||
| FIFO | 14.0 1.4 | 3.6 | 18.8 | 45.1 | 6.70 | 274.03 |
| LLMLingua-2 | 35.8 0.6 | 25.9 | 48.7 | 28.0 | 7.38 | 353.61 |
| Prompting | 37.2 1.6 | 24.6 | 51.8 | 26.6 | 7.34 | 177.48 |
| ACON-UT | 38.5 1.4 | 28.8 | 53.4 | 19.6 | 7.43 | 142.00 |
| Method | Acc. | Pass 3 | Steps | Peak |
|---|---|---|---|---|
| History-Only Compression | ||||
| Prompting | 78.6 2.7 | 68.5 | 17.0 | 9.23 |
| ACON-UT | 81.2 1.2 | 70.8 | 18.9 | 9.36 |
| ACON-UT (Locked) | 79.4 2.7 | 70.8 | 16.6 | 9.08 |
| PAIR (Ours) | 84.7 1.7 | 79.8 | 16.6 | 9.03 |
| Prefix-Conditioned Compression | ||||
| Benchmark | Scope | Boundaries | CF rollouts | Total (M) |
|---|---|---|---|---|
| AppWorld | History-only | 133 | 464 | 52.04 |
| Prefix-conditioned | 198 | 694 | 85.40 | |
| OfficeBench | History-only | 214 | 750 | 45.47 |
| Prefix-conditioned | 298 | 1,042 | 90.79 | |
| -Bench Retail | History-only | 200 | 700 | 44.03 |
| Prefix-conditioned | 204 | 714 | 39.59 |
| Benchmark | Scope | vs. | Pass 3 [95% CI] | Acc [95% CI] | win/lose ( ) | |
|---|---|---|---|---|---|---|
| AppWorld | history | ACON-UT | 168 | 22/7 (0.008) | ||
| AppWorld | prefix | ACON-UT | 168 | 21/8 (0.024) | ||
| OfficeBench | history | ACON-UT | 95 | 7/5 (0.774) | ||
| OfficeBench | prefix | ACON-UTCO | 95 | 13/6 (0.167) | ||
| -Retail | history | ACON-UTCO | 40 | 7/6 (1.000) | ||
| -Retail | prefix | ACON-UT | 40 | 5/2 (0.453) |
| Case | Boundary effect | Compression error | Prompt revision | Control PAIR |
|---|---|---|---|---|
| Outcome hazard | PRE , POST ; | The summary drops the coworker restriction and presents the total over all received payments as the answer. | Preserve task predicates and relationship filters; distinguish observations from unverified inferences. | success |
| Interaction burden | PRE and POST both ; | Previously read API specifications are reduced to prose, causing documentation and authentication to be repeated. | Preserve executable API signatures and completed state; do not reopen completed documentation or authentication. | mean steps |