Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer
Organizations: Nanjing University
Abstract
A harness is the code around a language-model agent that organizes prompts, calls tools, manages context, and controls execution. As models grow stronger, recent work has begun to let agents improve their own harnesses, a line of work known as self-evolving harnesses. In most existing methods, a separate proposer running on a human-designed harness modifies the solver's harness, and a separate harness is evolved for each benchmark. Real-world tasks come from many domains, so both the evolution and the evaluation of a harness should cover a diverse range of tasks. We propose a framework close to recursive self-improvement: the same frozen model, on the same version of the harness, first solves tasks as the solver and then, as the proposer, reads the complete run records and directly edits the harness that runs it. Each evolution batch draws tasks from five benchmarks in different domains. To measure generalization, training and held-out tasks are strictly separated, and we additionally evaluate on five out-of-distribution benchmarks never used during evolution. We frame the evolution process as deep-learning training with two stages, multi-task pretraining and continual training. Starting from a 49-line seed harness, the harness obtained at the end of the first stage improves the average score by 4.48 points on the in-distribution benchmarks and by 12.64 points on the out-of-distribution benchmarks, surpassing Codex on the former and matching it on the latter. In the second stage, continued evolution on Claw-Eval, one of the out-of-distribution benchmarks, further raises the score on that benchmark from 66.17 to 68.06, exceeding Codex. We also provide an in-depth analysis of the mechanisms that emerged during evolution, including output truncation, history compaction, and independent review.
Figures & tables
| Method | Same model | Same harness | Start | Editable scope | Evolution tasks | Acceptance |
|---|---|---|---|---|---|---|
| SICA | ✓ | ✓ | Hand-designed | All code | Same-type benchmarks | Score-based selection |
| DGM | ✗ | ✓ | Minimal | All code | Single benchmark | Score-based selection |
| Meta-Harness | ✗ | ✗ | Hand-designed | All code | Same-type benchmarks | Score-based selection |
| AHE | ✓ | ✗ | Minimal | Preset components | Single benchmark | Effect-based rollback |
| Self-Harness | ✓ | ✓ † | Minimal | Preset components | Single benchmark | Regression tests |
| HarnessX | ✗ | ✗ | Hand-designed | Preset components | Single benchmark | Regression tests |
| Benchmark | Task type | Total | Train | Held-out |
| In-distribution benchmarks | ||||
| Terminal-Bench 2.1 | Long-horizon engineering tasks in a terminal | 89 | 60 | 29 3 |
| GAIA2 | Asynchronous interaction in dynamic app environments | 800 | 60 | 100 |
| OfficeQA Pro | Treasury document retrieval and numerical reasoning | 133 | 60 | 73 2 |
| -Bench banking | Multi-turn service dialogue with a simulated user | 97 | 60 | 37 2 |
| SWE-Bench Pro | Issue resolution in code repositories | 731 | 60 | 100 |
| State | Lines | Main changes |
|---|---|---|
| Seed | 49 | System prompt, bash tool, and a basic tool-calling loop |
| Iteration 1 | 279 | Output truncation; command time limits; history compaction; image viewing |
| Iteration 2 | 343 | Completion check before finishing; image size cap |
| Iteration 3 | 394 | Independent review in a fresh context |
| Iteration 4 | 414 | End-of-conversation detection by word matching; image format validation |
| Iteration 5 | 430 | Task-specific rules (set operations, deadlines); review prioritizes unchecked requirements |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.