Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2% and 5.0%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction.
Figures & tables
Figure 1: Length-triggered compaction versus AutoCompact. Under length-triggered compaction, the agent keeps accumulating context after the bug is localized, such as repeated searches with long outputs, and summarizes only when the context limit is reached. AutoCompact lets the model decide (1) when to compact, (2) what to keep, rewriting useful findings into a working-state summary while dropping stale exploration, and (3) how to continue from it.
Figure 2: Judge-guided on-policy data collection. At each step, the judge reviews the policy’s proposal using only the current execution history. Useful exploration is kept, while corrections target (1) when to compact, e.g., replacing a search after the bug is localized with compact() ; (2) what to keep, e.g., repairing a summary that omits the target file; and (3) how to continue, e.g., redirecting a repeated search toward the edit planned in the summary. The environment executes the corrected outputs, and the resulting trajectories supervise SFT; the trained agent needs no judge.
Method
Trigger
Optimization
Context
SWE-bench Verified (%)
SWE-PolyBench Verified (%)
Full-history baseline
Base
—
—
256K
30.4
19.5
Length-triggered compaction
Fixed Compaction
Length
—
16K †
28.8
18.6
CompactionRL ( Li et al., 2026b )
Length
RL
16K †
32.7
19.8
Proactive compaction
Table 1: Full-run pass rates without an inference-cost cap, grouped by compaction trigger. All methods share the same base model and scaffold. In the Context column, 256K denotes the model context window, while † marks a 16K forced-compaction threshold.
Figure 3: Pass rate versus inference budget on SWE-bench Verified. (a) AutoCompact-SFT versus Base at 256K, where no trajectory reaches the forced-compaction threshold. (b) AutoCompact (after RL) versus AutoCompact-SFT at 256K. (c) AutoCompact versus Base at 16K, with a shared forced-compaction fallback. (d) The same checkpoint following versus ignoring summaries. Budget categories are equally spaced; gains denote additional solved tasks per 500 tasks.
Figure 4: Compaction behavior, with (1), (2), and (3) marking when to compact, what to keep, and how to continue (Figure 1 ). (a) Fraction of tasks in which the model proactively invokes compact() . (b) Fraction of summaries that omit relevant task or workspace state. (c) Fraction of summaries that omit a concrete next action. Lower is better in (b) and (c).
Figure 5: Summary self-consistency on SWE-bench Verified task django-13809 . This qualitative example contrasts paraphrased summaries from AutoCompact-SFT and AutoCompact with their final evaluation outcomes. Each summary records a state and proposes a next action; the marker between them indicates whether the action is consistent with the recorded state.
LLM-based coding agents solve software-engineering tasks through iterative interactions with development environments, where returned observations accumulate in the context and become a major source of inference cost. Observation compression reduces this cost by shortening observations before they are appended to the context. However, existing methods still exhibit an unsatisfactory efficiency-effectiveness trade-off, as they do not explicitly model how compression affects the agent's subsequent behavior. This paper proposes CoACT, an action-preserving observation compression method for coding agents. CoACT is built on next-action preservation (NAP), which requires a compressed observation to induce the same next action as the raw observation. By checking the agent's immediate next action, NAP provides a practical signal for whether a compression preserves the information needed for continued task solving. During training, a teacher model first generates multiple compressed candidates of each observation. CoACT then uses an action-preservation reward based on NAP to filter out candidates that would change the agent's next action, and uses a length-reduction reward to choose compact candidates as supervision for a lightweight compressor. Experiments on SWE-bench Verified with three agentic models show that CoACT reduces average total token consumption by 33.0% while maintaining task-solving effectiveness close to the uncompressed agent.
Haorui Chen, Yuancheng Zhu, Yitong Zhang +1
College of AI Tsinghua University Beijing, China · School of Intelligence Science and Technology Nanjing University Nanjing, China
The transition from human-centric assistance to Autonomous Software Engineering (ASE) agents has enabled the resolution of complex real-world SE tasks. However, the trial-and-error nature of these agents generates lengthy interaction trajectories, creating severe bottlenecks in terms of context window limits and cost. While context compression offers a potential remedy, prior approaches suffer from static pruning strategies and granularity mismatches, often failing to preserve the semantic dependencies and syntactic details crucial for SE tasks. To strictly preserve critical task evidence while reducing context length, we introduce AttnCompress, a dynamic attention-guided trajectory compression framework. Unlike existing approaches, AttnCompress bridges the gap between semantic integrity and dynamic adaptability through three key mechanisms: (1) structure-aware segmentation via perplexity (PPL) spikes to preserve the syntactic structure of code and logs; (2) relevance estimation using proxy attention weights to quantify the precise relevance of historical blocks to the agent's current reasoning; and (3) a dynamic rolling window to re-evaluate and recall historical context as the task evolves. Extensive evaluation on SWE-Bench-Verified and Multi-SWE-Bench demonstrates that AttnCompress achieves a pass rate of 53.17%, outperforming prior state-of-the-art baselines while reducing token consumption by 21.6% and total costs by 33.6%. The framework proves to be model-agnostic and generalizes effectively across diverse programming languages.
Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on Terminal-Bench and achieving new levels of efficiency for test-time scaling and state-of-the-art results on KernelBench. The per-rollout savings of CliffCompaction make the performance--cost trade-off of test-time scaling more efficient, adding over 10 percentage points on Terminal-Bench for less than the cost of two full-context runs. Under parallel test-time scaling, CliffCompaction lets Kimi K2.6 match Opus 4.7, and exceed Opus 4.6 and GPT-5.3 Codex at lower cost. The key to CliffCompaction's effectiveness is that it keeps compacted information faithful by only truncating or dropping content, never rephrasing or rewriting it. We never compact a compaction---each pass operates only on original content, and prior compacted output is discarded, preventing context drift from accumulating. These properties sustain continual learning over sessions exceeding a million tokens: on KernelBench, CliffCompaction reaches CUDA kernel speedups of 2.23× after 200 steps and 3.58× after 400 steps, surpassing specialized search algorithms and trained agents despite being a general-purpose compaction technique. We open-source a scaffold-agnostic API-proxy implementation of CliffCompaction usable with Claude Code, Codex and other harnesses.