Harness Annealing: Learning to Act with Less External Control
Organizations: Shanghai Jiao Tong University · University of California, Berkeley
Abstract
Language agents rely on external harnesses to track state, organize workflows, and verify answers. Beyond providing tools and information, these harnesses supply control decisions about what to investigate, whether to revise, and when to stop. Training on successful harness-supported trajectories can improve task performance while leaving these decisions dependent on runtime intervention. We ask whether harness-supported experience can also teach the model to make these decisions, allowing the division of control to change as the model learns. We call this objective harness internalization: learning to assume specified control responsibilities while retaining task performance after the corresponding support is withdrawn. We introduce HARNESS ANNEALING TRAINING (HAT), which combines explicit control supervision with a curriculum over teacher trajectories collected under progressively weaker harnesses. Experiments with 9B and 35B models on SWE-QA and SWE-QA-Pro evaluate every checkpoint under four deployment harnesses. Selected annealed checkpoints operating with tools alone achieve scores close to those of their respective starting checkpoints deployed with the full harness. The benefits vary with model scale and deployment configuration, and further annealing does not uniformly improve performance. These findings suggest that harness-supported experience can help reduce the runtime control required by a trained agent.
Figures & tables
| Nested configuration | Added runtime support | Targeted responsibility |
|---|---|---|
| Tool access and execution safeguards | Gather task evidence using external tools. | |
| Recent history, progress ledger, and repetition detection | Track evidence and gaps; avoid unproductive repetition. | |
| Locate Read Connect Answer | Organize investigation and answer construction; choose phase transitions. | |
| Heuristic draft checks and bounded repair | Check draft support; revise, gather evidence, or finish. |
| Model | Benchmark | Checkpoint | Deployment Harness | Row summary | |||||
|---|---|---|---|---|---|---|---|---|---|
| Mean | Worst | ||||||||
| Qwen3.5-9B | SWE-QA | Base | 59.93 | 62.92 | 67.28 | 66.79 | 64.23 | 59.93 | 7.35 |
| -trained | 71.99 | 75.94 | 77.47 | 75.70 | 75.28 | 71.99 | 5.48 | ||
| -targeted | 77.41 | 77.15 | 75.10 | 75.66 | 76.33 | 75.10 | 2.31 | ||
| -targeted | 77.00 | 76.62 | 76.43 | 76.45 | 76.63 | 76.43 | 0.57 | ||
| -targeted | 76.30 | 75.12 | 74.43 | 74.16 | 75.00 | 74.16 | 2.14 | ||