Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2% and 5.0%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction.
Figures & tables
Figure 1: Length-triggered compaction versus AutoCompact. Under length-triggered compaction, the agent keeps accumulating context after the bug is localized, such as repeated searches with long outputs, and summarizes only when the context limit is reached. AutoCompact lets the model decide (1) when to compact, (2) what to keep, rewriting useful findings into a working-state summary while dropping stale exploration, and (3) how to continue from it.
Figure 2: Judge-guided on-policy data collection. At each step, the judge reviews the policy’s proposal using only the current execution history. Useful exploration is kept, while corrections target (1) when to compact, e.g., replacing a search after the bug is localized with compact() ; (2) what to keep, e.g., repairing a summary that omits the target file; and (3) how to continue, e.g., redirecting a repeated search toward the edit planned in the summary. The environment executes the corrected outputs, and the resulting trajectories supervise SFT; the trained agent needs no judge.
Method
Trigger
Optimization
Context
SWE-bench Verified (%)
SWE-PolyBench Verified (%)
Full-history baseline
Base
—
—
256K
30.4
19.5
Length-triggered compaction
Fixed Compaction
Length
—
16K †
28.8
18.6
CompactionRL ( Li et al., 2026b )
Length
RL
16K †
32.7
19.8
Proactive compaction
Table 1: Full-run pass rates without an inference-cost cap, grouped by compaction trigger. All methods share the same base model and scaffold. In the Context column, 256K denotes the model context window, while † marks a 16K forced-compaction threshold.
Figure 3: Pass rate versus inference budget on SWE-bench Verified. (a) AutoCompact-SFT versus Base at 256K, where no trajectory reaches the forced-compaction threshold. (b) AutoCompact (after RL) versus AutoCompact-SFT at 256K. (c) AutoCompact versus Base at 16K, with a shared forced-compaction fallback. (d) The same checkpoint following versus ignoring summaries. Budget categories are equally spaced; gains denote additional solved tasks per 500 tasks.
Figure 4: Compaction behavior, with (1), (2), and (3) marking when to compact, what to keep, and how to continue (Figure 1 ). (a) Fraction of tasks in which the model proactively invokes compact() . (b) Fraction of summaries that omit relevant task or workspace state. (c) Fraction of summaries that omit a concrete next action. Lower is better in (b) and (c).
Figure 5: Summary self-consistency on SWE-bench Verified task django-13809 . This qualitative example contrasts paraphrased summaries from AutoCompact-SFT and AutoCompact with their final evaluation outcomes. Each summary records a state and proposes a next action; the marker between them indicates whether the action is consistent with the recorded state.