Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this process through iterative code generation and evaluation, but retains a fixed development set and proposal policy. These constraints channel evolution along a single search trajectory, increasing the risk of converging to a local optimum. We make the improvement process itself adaptive by organizing search into branches with evolving development subsets and proposal policies. Each branch retains development cases solved by more of its leading harnesses than by those of other branches, drops cases solved by every leading harness across all branches, and revises its proposal policy using its own search history. To deploy the resulting complementary harnesses, we propose a router to select one development-selected branch head for each new input before execution. Across mathematical reasoning and agentic coding benchmarks, our system achieves relative improvements over Meta-Harness of 34.8% on Olympiad-level mathematical reasoning, 11.6% on Terminal-Bench 2.0, and 3.8% on SWE-bench Lite, with harness selection and router configuration based solely on development data. These results show that evolving branch objectives and proposal policies can yield complementary harnesses whose strengths a router combines without access to test outcomes.
Figures & tables
Figure 1: Pipeline overview. Left: each branch proposes and evaluates harnesses using its current development subset, branch history, and proposal guidance. Periodic guidance updates (upper inset) translate local search experience into priorities for subsequent proposals. Right: development subsets are updated by comparing how many leading harnesses in each branch solve each case. Bottom: a router selects one development-best branch harness for each new input before execution.
Method
Math
Terminal-Bench 2.0
SWE-bench Lite
Gemini 3 Flash
Claude Sonnet 4.5
Claude Sonnet 4.5
Claude Sonnet 4.5
No-memory
34.0
24.5
—
—
Few-shot
40.0
23.0
—
—
BM25-all
40.5
25.0
—
—
BM25-geometry
46.5
27.0
—
—
Terminus 2
—
—
34.5
—
Table 1: Held-out performance (%) of fixed baselines and development-selected systems. Bold indicates the best score in each setting.
Figure 2: Left: per-iteration test-best accuracy and discovered mechanisms on Math–Gemini over 20 iterations. Right: shared and branch-exclusive coverage of development-selected heads in four settings. All rates (%) are computed over the full test sets. Venn areas are schematic.
Math
Terminal-Bench 2.0
SWE-bench Lite
Gemini 3 Flash
Claude Sonnet 4.5
Claude Sonnet 4.5
Claude Sonnet 4.5
Head 1 only
56.0
26.0
44.8
65.6
Head 2 only
58.0
31.0
48.3
66.4
Router (ours)
62.0
30.5
50.0
66.0
Table 2: Held-out performance (%) of development-selected branch heads and the routed system.
Method
Math
Terminal-Bench 2.0
SWE-bench Lite
Gemini 3 Flash
Claude Sonnet 4.5
Claude Sonnet 4.5
Claude Sonnet 4.5
Meta-Harness
49.0
32.5
46.6
63.6
Ours
58.0
32.5
50.0
66.4
Table 4: Test-best performance (%) among ten harnesses per method: the top five by development score per branch for ours and the top ten for Meta-Harness.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Meta-Harness
Ours
Benchmark
Task solving
Harness authoring
Total
Task solving
Harness authoring
Total
Math
11.10
68.82
79.92
13.40
159.50
172.90
Terminal-Bench 2.0
2,086.10
280.70
2,366.80
3,564.60
436.79
4,001.39
SWE-bench Lite
1,814.60
92.57
1,907.17
2,726.70
224.78
2,951.48
Appendix
Table 5: Full-run token usage (millions) for Meta-Harness and ours across three benchmarks with Claude Sonnet 4.5 as the action model and Claude Opus 4.6 as the proposer. Task-solving and harness-authoring tokens sum to the total, with cached tokens counted once.
Router components
Math
Terminal-Bench 2.0
SWE-bench Lite
Exclusive Cases
Expert Outputs
GEPA
Gemini 3 Flash
Claude Sonnet 4.5
Claude Sonnet 4.5
Claude Sonnet 4.5
✓
✓
✓
62.0
30.5
50.0
66.0
✓
✓
—
59.5
29.0
46.6
66.4
✓
—
✓
57.0
31.0
48.3
66.4
✓
—
—
57.0
30.0
48.3
66.4
—
—
✓
59.5
30.5
48.3
66.0
Appendix
Table 6: Router configurations and held-out performance (%). All variants receive expert source code. Checkmarks indicate included context and whether GEPA optimization is applied.
Automating the search for effective harnesses is an important step toward enabling agents to recursively self-improve. Existing harness optimizations typically produce a single global harness that is applied uniformly across task instances. However, a harness that works well on average may not be optimal for every instance. We introduce Turbo Harness, a framework that can adapt a globally optimized harness to each instance by reusing information generated during the original optimization process. Specifically, Turbo Harness recycles artifacts produced during a completed global harness optimization run, and summarizes them into a structured playbook. We train a harness editor to leverage this prior optimization experience to generate instance-specific patches to the global harness. At inference time, the editor uses the instance and the playbook to construct a tailored harness in which the execution model operates. Through numerical experiments, we show that Turbo Harness consistently outperforms existing harness optimization baselines across seven benchmarks spanning interactive agent tasks, software engineering, and long-horizon terminal tasks.
Tunyu Zhang, Hao Wang, Kai Xu +1
Rutgers University · Red Hat AI Innovation · MIT-IBM Watson AI Lab
A harness is the code around a language-model agent that organizes prompts, calls tools, manages context, and controls execution. As models grow stronger, recent work has begun to let agents improve their own harnesses, a line of work known as self-evolving harnesses. In most existing methods, a separate proposer running on a human-designed harness modifies the solver's harness, and a separate harness is evolved for each benchmark. Real-world tasks come from many domains, so both the evolution and the evaluation of a harness should cover a diverse range of tasks. We propose a framework close to recursive self-improvement: the same frozen model, on the same version of the harness, first solves tasks as the solver and then, as the proposer, reads the complete run records and directly edits the harness that runs it. Each evolution batch draws tasks from five benchmarks in different domains. To measure generalization, training and held-out tasks are strictly separated, and we additionally evaluate on five out-of-distribution benchmarks never used during evolution. We frame the evolution process as deep-learning training with two stages, multi-task pretraining and continual training. Starting from a 49-line seed harness, the harness obtained at the end of the first stage improves the average score by 4.48 points on the in-distribution benchmarks and by 12.64 points on the out-of-distribution benchmarks, surpassing Codex on the former and matching it on the latter. In the second stage, continued evolution on Claw-Eval, one of the out-of-distribution benchmarks, further raises the score on that benchmark from 66.17 to 68.06, exceeding Codex. We also provide an in-depth analysis of the mechanisms that emerged during evolution, including output truncation, history compaction, and independent review.
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a task-specific manner can improve execution-trace quality while remaining computationally lightweight and requiring only a few update iterations. To this end, we introduce Recursive Harness Self-Improvement (RHI), which represents the harness as a prompt-level specification of the agent loop and iteratively refines it using pairwise feedback over its own revision history. Across 30 synthetic machine-learning research tasks spanning quantitative finance, robotics, and pharmacy, a few RHI iterations suffice to substantially raise the performance ceiling of low-reasoning-effort agents, exceeding the corresponding maximum-reasoning-effort setting while reducing inference cost by up to 60%. We show that these gains arise primarily from improved task-specific context management through more effective inter-agent information flow rather than longer reasoning traces. Finally, we formalize this behavior as an information-theoretic hypothesis for RHI's implicit optimization objective, suggesting RHI as a practical algorithm for continual learning within the paradigm of model--harness co-evolution.