Harness Evolution Hits a Ceiling: When Weight Training Should Begin
Organizations: Independent Researcher · University of California, Berkeley · Purdue University · Stanford University
Abstract
Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. We show that the right lever can be read off the agent's failure composition: labelling failed trajectories by the first signal that fires separates process failures (blocked calls, loops, exhausted step budgets) from content failures (a delivered plan that is poor). Harness evolution repairs the former, the behaviour it instils can be trained into the weights, and content failures are what weight training is for. On DeepPlanning, a self-evolving harness loop lifts the held-out score of Qwen3.5-4B from 0.16 to 0.30 and of Qwen3.5-9B from 0.32 to 0.44; for 4B, held-out delivery rises from 55% to 90% while content failures are left for the weights. LoRA adapters trained on evolved-harness trajectories internalise the gain: under the original harness they add +0.13 on held-out tasks for both sizes; on 4B they stack with the harness to more than double the held-out score, and on 9B the adapter alone matches the full evolution line, cutting content failures from a quarter of trajectories to one in twenty. A placebo adapter trained on answer-shuffled trajectories falls below the base model. The loop transfers to WebArena-Lite (+0.09 on 117 unseen tasks), where the gain lives in what the model sees and adapters do not add to it. The result is a diagnose-then-intervene rule applied twice: read the failure composition to choose between harness and weights, then read what the accepted edits changed to decide which gains to train in. Scores are four-rollout means against fresh anchors, same-night except where marked, across eight models from six families and two benchmarks.
Figures & tables
| Qwen3.5-4B | Qwen3.5-9B | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Seed harness | Evolved harness | Seed harness | Evolved harness | ||||||
| Quantity | Weights | Dev | Held-out | Dev | Held-out | Dev | Held-out | Dev | Held-out |
| Score | Base | 0.158 | 0.164 | 0.301 | 0.297 | 0.348 | 0.323 | 0.404 | 0.442 |
| Score | LoRA qkvo | 0.309 | 0.290 | 0.342 | 0.379 | 0.327 | 0.371 | 0.395 | 0.433 |
| Score | LoRA qkvo+MLP | 0.247 | 0.259 | 0.351 | 0.359 | 0.398 | 0.454 | 0.414 | 0.449 |
| Delivery rate | Base | 52% | 55% | 90% | 90% | 96% | 97% | 98% | 99% |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Rollouts per cell | Resolvable difference |
|---|---|
| 1 | 0.07 |
| 2 | 0.05 |
| 4 | 0.035 |
| 8 | 0.025 |
| 16 | 0.02 |
| Round | Retained edit | Surface | Other candidate (verdict) | Anchor | F3 |
|---|---|---|---|---|---|
| 0 | seed harness | 0.156 | 64.1% | ||
| 1 | plan_submission_format_fix | prompts | malformed_plan_tag_middleware (accepted, not retained) | 0.204 | 66.7% |
| 2 | plan_emission_lessons (new lessons file) | prompts | plan_tool_call_interceptor (accepted, not retained) | 0.221 | 42.2% |
| 3 | gate_self_check_activation | policies | plan_envelope_rescue_middleware (accepted, not retained) | 0.225 | 38.0% |
| 4 | none | plan_format_middleware, plan_submission_tool (both rejected) | 0.225 | ||
| 5 | none | coordinate_fallback_lesson, plan_truncation_guard (both rejected) | 0.225 |
| Round | Candidate | Surface | Reading | Verdict | Anchor after round |
|---|---|---|---|---|---|
| 0 | seed harness | 0.305 | |||
| 1 | plan_emission_format_lessons | prompts | 0.380 | accepted | 0.377 |
| 1 | query_constraint_extractor_tool | tools | 0.335 | rejected (never called) | |
| 2 | hotel_location_lessons | prompts | 0.379 | accepted | 0.366 |
| 2 | hotel_location_validator_tool | tools | 0.399 | rejected (never called) | |
| 3 | standalone_plan_emission_lessons | prompts | 0.374 | accepted | 0.390 |
| Development (48 tasks) | Held-out (72 tasks) | ||||||
|---|---|---|---|---|---|---|---|
| Model | Quantity | Seed | Evolved | Gain | Seed | Evolved | Gain |
| Qwen3.5-4B | Score | 0.158 | 0.301 | +0.143 | 0.164 | 0.297 | +0.133 |
| Qwen3.5-4B | Delivery rate | 52% | 90% | +38 pp | 55% | 90% | +35 pp |
| Qwen3.5-9B | Score | 0.348 | 0.404 | +0.056 | 0.323 | 0.442 | +0.119 |
| Qwen3.5-9B | Delivery rate | 96% | 98% | +2 pp | 97% | 99% | +2 pp |
| Model | Harness | Healthy | F1 evidence | F2 harness friction | F3 execution control | F4 planning capacity |
|---|---|---|---|---|---|---|
| Qwen3.5-4B | Seed | 14.3% | 8.6% | 2.9% | 64.1% | 9.9% |
| Qwen3.5-4B | Evolved (round 13) | 31.8% | 3.1% | 0.0% | 37.0% | 28.1% |
| Qwen3.5-9B | Seed | 26.7% | 5.2% | 2.1% | 38.9% | 27.1% |
| Qwen3.5-9B | Evolved (round 3) | 40.6% | 17.7% | 0.0% | 16.7% | 25.0% |
| Model | Harness | Composite score | Delivery rate | Score among delivered |
|---|---|---|---|---|
| Qwen3.5-4B | Seed | 0.156 | 59% | 0.266 |
| Qwen3.5-4B | Evolved (13 rounds) | 0.292 | 88% | 0.332 |
| Qwen3.5-9B | Seed | 0.299 | 97% | 0.309 |
| Qwen3.5-9B | Evolved (5 rounds) | 0.390 | 99% | 0.392 |
| MiniCPM5-2B | Official benchmark loop | 0.047 | 33% | 0.141 |
| MiniCPM5-2B | Seed | 0.106 | 61% | 0.173 |
| Model | Weights | Dev score | Held-out score | Dev gain over base | Held-out gain over base |
|---|---|---|---|---|---|
| Qwen3.5-4B | Base | 0.156 | 0.150 | ||
| Qwen3.5-4B | LoRA qkvo | 0.193 | 0.212 | +0.037 | +0.062 |
| Qwen3.5-4B | LoRA qkvo+MLP | 0.280 | 0.267 | +0.124 | +0.117 |
| Qwen3.5-4B | Full-parameter | 0.294 | 0.311 | +0.138 | +0.161 |
| Qwen3.5-9B | Base | 0.306 | 0.370 | ||
| Qwen3.5-9B | LoRA qkvo+MLP | 0.426 | 0.442 | +0.120 | +0.072 |
| Model | Adapter | Split | Delivery | Healthy | F1 evidence | F2 harness friction | F3 execution control | F4 planning capacity |
|---|---|---|---|---|---|---|---|---|
| Qwen3.5-4B | none | dev 48 | 52% | 15.6% | 10.9% | 3.6% | 57.8% | 12.0% |
| Qwen3.5-4B | LoRA qkvo | dev 48 | 78% | 35.4% | 8.9% | 2.1% | 39.6% | 14.1% |
| Qwen3.5-4B | LoRA qkvo+MLP | dev 48 | 67% | 26.6% | 12.0% | 4.7% | 51.0% | 5.7% |
| Qwen3.5-4B | none | held-out 72 | 55% | 15.3% | 9.0% | 3.5% | 64.6% | 7.6% |
| Qwen3.5-4B | LoRA qkvo | held-out 72 | 78% | 30.6% | 8.7% | 2.1% | 45.1% | 13.5% |
| Qwen3.5-4B | LoRA qkvo+MLP | held-out 72 | 62% | 30.9% | 8.7% | 3.5% | 48.6% | 8.3% |
| Development (48 tasks) | Held-out (72 tasks) | ||||
| Weights | Score | Delivery rate | Score | Delivery rate | Rounds per trajectory |
| Base (no adapter) | 0.333 | 98% | 0.379 | 99% | 19 |
| Real LoRA | 0.399 | not recorded | 0.426 | not recorded | not recorded |
| Placebo LoRA | 0.175 | 86% | 0.152 | 76% | 52–65 |
| Placebo minus base | pp | pp | to | ||
| Cell | Removes | Dev score | Held-out score | Dev vs full | Held-out vs full | Dev delivery | Dev score among delivered |
|---|---|---|---|---|---|---|---|
| Full | nothing | 0.366 | 0.438 | 99% | 0.368 | ||
| Checkpoint | mid-trajectory checkpoint | 0.385 | 0.440 | +0.019 | +0.002 | ||
| Gate | exit gate | 0.384 | 0.410 | +0.017 | |||
| Hard blocks | four hard blocks | 0.383 | 0.432 | +0.017 | |||
| Lessons | evolved lessons file | 0.317 | 0.362 | ||||
| Bare | everything | 0.290 | 0.346 | 85% | 0.340 |
| Cell | Removes | Dev score | Held-out score | Dev vs full | Held-out vs full |
|---|---|---|---|---|---|
| Full | nothing | 0.264 | 0.314 | ||
| Evolved middleware | citation auto-fix, submission nag | 0.310 | 0.313 | +0.046 | |
| Evolved rewrites | three policy and prompt rewrites | 0.302 | 0.287 | +0.038 | |
| Checkpoint | mid-trajectory checkpoint | 0.270 | 0.330 | +0.007 | +0.016 |
| Gate (4 rollouts) | exit gate | 0.290 | 0.237 | +0.026 | |
| Ledger | state ledger | 0.253 | 0.254 |
| Dev score | Test, 117 unseen tasks | |||||||
|---|---|---|---|---|---|---|---|---|
| Harness | Weights | (48 tasks) | Score | Delivery | Healthy | F3 | F4 | Accuracy among delivered |
| Seed | Base | 0.271 | 0.256 | 49% | 23.1% | 45.9% | 25.6% | 0.47 |
| Evolved | Base | 0.375 | 0.344 | 77% | 33.8% | 40.2% | 18.2% | 0.44 |
| Seed | LoRA | 0.253 | 0.216 | 42% | 20.3% | 52.6% | 19.4% | 0.49 |
| Seed | Full-parameter | 0.250 | 0.231 | 45% | 20.7% | 47.0% | 23.9% | 0.46 |
| Evolved | LoRA | 0.349 | 0.327 | 81% | 31.4% | 38.7% | 17.1% | 0.39 |
| Model | Harness | Score | Delivery | Score among delivered | Healthy | F1 evidence | F2 harness friction | F3 execution control | F4 planning capacity |
|---|---|---|---|---|---|---|---|---|---|
| MiniCPM5-2B | seed | 0.106 | 61% | 0.173 | 5.7% | 42.7% | 2.1% | 32.3% | 17.2% |
| MiniCPM5-2B | evolved (5 rounds) | 0.110 | 71% | 0.156 | 4.7% | 35.9% | 2.6% | 34.4% | 22.4% |
| Spark-X2.5-4B | seed | 0.153 | 97% | 0.158 | 6.2% | 20.8% | 2.1% | 33.3% | 37.5% |
| Spark-X2.5-4B | evolved (5 rounds) | 0.155 | 95% | 0.164 | 5.2% | 6.8% | 0.0% | 28.1% | 59.9% |
| Ministral-3-8B | seed | 0.112 | 90% | 0.125 | 0.0% | 28.1% | 0.0% | 7.3% | 64.6% |
| Ministral-3-8B | evolved (5 rounds) | 0.127 | 93% | 0.136 | 3.1% | 30.7% | 0.5% | 5.2% | 60.4% |
| Model | Harness | D score | H score | Delivery D / H | Mean steps D / H | Quality D |
|---|---|---|---|---|---|---|
| Kimi K3 | Bare | 0.418 | 0.402 | 96.9% / 95.2% | 19.7 / 23.3 | 0.431 |
| Kimi K3 | Evolved | 0.476 | 0.451 | 99.3% / 98.7% | 8.4 / 10.1 | 0.479 |
| GLM 5.3 | Evolved | 0.512 | 0.488 | 99.6% / 99.1% | 7.6 / 9.2 | 0.514 |