Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this process through iterative code generation and evaluation, but retains a fixed development set and proposal policy. These constraints channel evolution along a single search trajectory, increasing the risk of converging to a local optimum. We make the improvement process itself adaptive by organizing search into branches with evolving development subsets and proposal policies. Each branch retains development cases solved by more of its leading harnesses than by those of other branches, drops cases solved by every leading harness across all branches, and revises its proposal policy using its own search history. To deploy the resulting complementary harnesses, we propose a router to select one development-selected branch head for each new input before execution. Across mathematical reasoning and agentic coding benchmarks, our system achieves relative improvements over Meta-Harness of 34.8% on Olympiad-level mathematical reasoning, 11.6% on Terminal-Bench 2.0, and 3.8% on SWE-bench Lite, with harness selection and router configuration based solely on development data. These results show that evolving branch objectives and proposal policies can yield complementary harnesses whose strengths a router combines without access to test outcomes.
Figures & tables
Figure 1: Pipeline overview. Left: each branch proposes and evaluates harnesses using its current development subset, branch history, and proposal guidance. Periodic guidance updates (upper inset) translate local search experience into priorities for subsequent proposals. Right: development subsets are updated by comparing how many leading harnesses in each branch solve each case. Bottom: a router selects one development-best branch harness for each new input before execution.
Method
Math
Terminal-Bench 2.0
SWE-bench Lite
Gemini 3 Flash
Claude Sonnet 4.5
Claude Sonnet 4.5
Claude Sonnet 4.5
No-memory
34.0
24.5
—
—
Few-shot
40.0
23.0
—
—
BM25-all
40.5
25.0
—
—
BM25-geometry
46.5
27.0
—
—
Terminus 2
—
—
34.5
—
Table 1: Held-out performance (%) of fixed baselines and development-selected systems. Bold indicates the best score in each setting.
Figure 2: Left: per-iteration test-best accuracy and discovered mechanisms on Math–Gemini over 20 iterations. Right: shared and branch-exclusive coverage of development-selected heads in four settings. All rates (%) are computed over the full test sets. Venn areas are schematic.
Math
Terminal-Bench 2.0
SWE-bench Lite
Gemini 3 Flash
Claude Sonnet 4.5
Claude Sonnet 4.5
Claude Sonnet 4.5
Head 1 only
56.0
26.0
44.8
65.6
Head 2 only
58.0
31.0
48.3
66.4
Router (ours)
62.0
30.5
50.0
66.0
Table 2: Held-out performance (%) of development-selected branch heads and the routed system.
Method
Math
Terminal-Bench 2.0
SWE-bench Lite
Gemini 3 Flash
Claude Sonnet 4.5
Claude Sonnet 4.5
Claude Sonnet 4.5
Meta-Harness
49.0
32.5
46.6
63.6
Ours
58.0
32.5
50.0
66.4
Table 4: Test-best performance (%) among ten harnesses per method: the top five by development score per branch for ours and the top ten for Meta-Harness.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Meta-Harness
Ours
Benchmark
Task solving
Harness authoring
Total
Task solving
Harness authoring
Total
Math
11.10
68.82
79.92
13.40
159.50
172.90
Terminal-Bench 2.0
2,086.10
280.70
2,366.80
3,564.60
436.79
4,001.39
SWE-bench Lite
1,814.60
92.57
1,907.17
2,726.70
224.78
2,951.48
Appendix
Table 5: Full-run token usage (millions) for Meta-Harness and ours across three benchmarks with Claude Sonnet 4.5 as the action model and Claude Opus 4.6 as the proposer. Task-solving and harness-authoring tokens sum to the total, with cached tokens counted once.
Router components
Math
Terminal-Bench 2.0
SWE-bench Lite
Exclusive Cases
Expert Outputs
GEPA
Gemini 3 Flash
Claude Sonnet 4.5
Claude Sonnet 4.5
Claude Sonnet 4.5
✓
✓
✓
62.0
30.5
50.0
66.0
✓
✓
—
59.5
29.0
46.6
66.4
✓
—
✓
57.0
31.0
48.3
66.4
✓
—
—
57.0
30.0
48.3
66.4
—
—
✓
59.5
30.5
48.3
66.0
Appendix
Table 6: Router configurations and held-out performance (%). All variants receive expert source code. Checkmarks indicate included context and whether GEPA optimization is applied.