Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. We show that the right lever can be read off the agent's failure composition: labelling failed trajectories by the first signal that fires separates process failures (blocked calls, loops, exhausted step budgets) from content failures (a delivered plan that is poor). Harness evolution repairs the former, the behaviour it instils can be trained into the weights, and content failures are what weight training is for. On DeepPlanning, a self-evolving harness loop lifts the held-out score of Qwen3.5-4B from 0.16 to 0.30 and of Qwen3.5-9B from 0.32 to 0.44; for 4B, held-out delivery rises from 55% to 90% while content failures are left for the weights. LoRA adapters trained on evolved-harness trajectories internalise the gain: under the original harness they add +0.13 on held-out tasks for both sizes; on 4B they stack with the harness to more than double the held-out score, and on 9B the adapter alone matches the full evolution line, cutting content failures from a quarter of trajectories to one in twenty. A placebo adapter trained on answer-shuffled trajectories falls below the base model. The loop transfers to WebArena-Lite (+0.09 on 117 unseen tasks), where the gain lives in what the model sees and adapters do not add to it. The result is a diagnose-then-intervene rule applied twice: read the failure composition to choose between harness and weights, then read what the accepted edits changed to decide which gains to train in. Scores are four-rollout means against fresh anchors, same-night except where marked, across eight models from six families and two benchmarks.
Figures & tables
Figure 1: Overview. Trajectories under the seed harness are labelled by their first failure signal (Section 3.1 ) and the composition decides the lever: a self-evolving harness loop removes the process failures, the trajectories it produces train the weights for the content failures, and both levers are measured under the seed and the evolved harness on held-out tasks.
Figure 2: Per-round fresh anchors of the two evolution lines (48 development tasks, 4-rollout means; band =±0.023 ; filled: edit accepted, fresh anchor measured; hollow: both candidates rejected, anchor kept). Small markers are the two candidates of each round at their own readings (dot retained, hollow circle passed but not retained, cross rejected); the gap between a retained reading and the anchor that follows is the selection optimism of Proposition 1 . The 4B dip at round 11 is a re-anchor after the context window was raised to 262k tokens.
Figure 3: Failure composition of the fresh anchor after each accepted round , from the same per-round history files as Figure 2 . Loops (F3) fall and the healthy share rises in both lines; F4 is never reduced. In the 9B line the lessons file raises the share of plans with unbacked claims (F1).
Qwen3.5-4B
Qwen3.5-9B
Seed harness
Evolved harness
Seed harness
Evolved harness
Quantity
Weights
Dev
Held-out
Dev
Held-out
Dev
Held-out
Dev
Held-out
Score
Base
0.158
0.164
0.301
0.297
0.348
0.323
0.404
0.442
Score
LoRA qkvo
0.309
0.290
0.342
0.379
0.327
0.371
0.395
0.433
Score
LoRA qkvo+MLP
0.247
0.259
0.351
0.359
0.398
0.454
0.414
0.449
Delivery rate
Base
52%
55%
90%
90%
96%
97%
98%
99%
Table 1: Base model and LoRA adapters trained on evolved-harness trajectories, under the seed and the evolved harness, on the 48 development and 72 held-out tasks. 4 rollouts per cell on one machine and one night, except the base-model cells under the evolved harness on held-out tasks, the lines’ closing anchors from two nights earlier (4B: 2 rollouts); differences below 0.035 (0.05 against a 2-rollout cell) are within noise. Base-row delivery rates differ from Table 7 (a different night).
Figure 4: Harness evolution on the nine DeepPlanning lines reported (48 development tasks). Left: seed (grey) and closing (blue) score of each line, ordered by the quality of the plans the seed harness delivers (in parentheses; on/off = thinking on or off). Every line is measured under the protocol of Section 4.2 , the three early lines and Gemma by a same-night re-measurement of seed and closing compositions (Appendix H ). Right: the gain against the seed plan quality; band =±0.035 , the resolution of the protocol on the Qwen3.5 cells; the Qwen3.6 thinking-on cells are noisier and their +0.053 is within that noise (Appendix H ).
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Rollouts per cell
Resolvable difference
1
0.07
2
0.05
4
0.035
8
0.025
16
0.02
Appendix
Table 2: Smallest difference between two cells that we treat as resolvable, as a function of the number of rollouts averaged per cell, given a per-rollout standard deviation of 0.023 (median over 26 same-night DeepPlanning cells; worst cell 0.042). The paper uses four rollouts per cell.
Round
Retained edit
Surface
Other candidate (verdict)
Anchor
F3
0
seed harness
0.156
64.1%
1
plan_submission_format_fix
prompts
malformed_plan_tag_middleware (accepted, not retained)
0.204
66.7%
2
plan_emission_lessons (new lessons file)
prompts
plan_tool_call_interceptor (accepted, not retained)
0.221
42.2%
3
gate_self_check_activation
policies
plan_envelope_rescue_middleware (accepted, not retained)
Table 3: Qwen3.5-4B evolution line, round by round. Each round proposed two candidates; the retained edit is listed with the harness surface it touched, the other candidate with its verdict. Anchor = fresh 4-rollout mean of the retained harness on the 48 development tasks, measured after the round; F3 = share of loop failures in that anchor. Rounds 4, 5 and 12 rejected both candidates and kept the previous anchor.
Round
Candidate
Surface
Reading
Verdict
Anchor after round
0
seed harness
0.305
1
plan_emission_format_lessons
prompts
0.380
accepted
0.377
1
query_constraint_extractor_tool
tools
0.335
rejected (never called)
2
hotel_location_lessons
prompts
0.379
accepted
0.366
2
hotel_location_validator_tool
tools
0.399
rejected (never called)
3
standalone_plan_emission_lessons
prompts
0.374
accepted
0.390
Appendix
Table 4: Qwen3.5-9B evolution line, candidate by candidate (5 rounds, 10 candidates, 3 accepted). Reading = the candidate’s own 4-rollout mean when judged; anchor = fresh 4-rollout mean of the retained harness after the round. All three accepted edits modify the distilled lessons file; every proposed tool was rejected because the agent never called it.
Development (48 tasks)
Held-out (72 tasks)
Model
Quantity
Seed
Evolved
Gain
Seed
Evolved
Gain
Qwen3.5-4B
Score
0.158
0.301
+0.143
0.164
0.297
+0.133
Qwen3.5-4B
Delivery rate
52%
90%
+38 pp
55%
90%
+35 pp
Qwen3.5-9B
Score
0.348
0.404
+0.056
0.323
0.442
+0.119
Qwen3.5-9B
Delivery rate
96%
98%
+2 pp
97%
99%
+2 pp
Appendix
Table 5: Harness evolution before and after. Score = 4-run mean composite; delivery rate = fraction of tasks on which a parsable plan was returned. Seed harness and final evolved harness were re-evaluated outside the loop on the same machine, on the 48 development tasks and the 72 held-out tasks; the evolved-harness held-out cells are the lines’ closing anchors measured two nights before the other three cells (4B: 2 rollouts). The 4B line ran 13 rounds (10 accepted edits); the 9B line ran 5 rounds (3 accepted edits). The 4-run discrimination threshold is 0.035; smaller differences are within noise.
Model
Harness
Healthy
F1 evidence
F2 harness friction
F3 execution control
F4 planning capacity
Qwen3.5-4B
Seed
14.3%
8.6%
2.9%
64.1%
9.9%
Qwen3.5-4B
Evolved (round 13)
31.8%
3.1%
0.0%
37.0%
28.1%
Qwen3.5-9B
Seed
26.7%
5.2%
2.1%
38.9%
27.1%
Qwen3.5-9B
Evolved (round 3)
40.6%
17.7%
0.0%
16.7%
25.0%
Appendix
Table 6: Failure composition of the seed harness and of the final evolved harness on the 48 development tasks (share of trajectories; fresh anchors, 4 rollouts each except the 9B seed row, which is the 6-rollout line baseline; labels as defined in Section 3.1 ). Rows sum to 100% up to rounding. The largest change in both lines is the fall in loop failures (F3); F4 is not reduced and ends as the largest failure class for 9B and the second-largest, after the residual F3, for 4B.
Model
Harness
Composite score
Delivery rate
Score among delivered
Qwen3.5-4B
Seed
0.156
59%
0.266
Qwen3.5-4B
Evolved (13 rounds)
0.292
88%
0.332
Qwen3.5-9B
Seed
0.299
97%
0.309
Qwen3.5-9B
Evolved (5 rounds)
0.390
99%
0.392
MiniCPM5-2B
Official benchmark loop
0.047
33%
0.141
MiniCPM5-2B
Seed
0.106
61%
0.173
Appendix
Table 7: Score decomposition on the 48 development tasks: composite score = delivery rate × mean score among delivered plans. Delivery = the trajectory returned a plan the scorer could parse. Seed and evolved anchors of each evolution line (4 rollouts; the 9B seed row is the 4-rollout subset of the 6-rollout line baseline). MiniCPM5-2B is included as the case in which the two factors cancel.
Figure 5: Base model and LoRA adapters under the seed and the evolved harness , same machine, 4-rollout means (Table 1 ; provenance of the evolved-harness held-out base cells in its caption). On the seed harness the adapter alone recovers as much as harness evolution; on the evolved harness the adapter adds to the 4B score, and for 9B the adapter alone already reaches the evolved-harness level.
Model
Weights
Dev score
Held-out score
Dev gain over base
Held-out gain over base
Qwen3.5-4B
Base
0.156
0.150
Qwen3.5-4B
LoRA qkvo
0.193
0.212
+0.037
+0.062
Qwen3.5-4B
LoRA qkvo+MLP
0.280
0.267
+0.124
+0.117
Qwen3.5-4B
Full-parameter
0.294
0.311
+0.138
+0.161
Qwen3.5-9B
Base
0.306
0.370
Qwen3.5-9B
LoRA qkvo+MLP
0.426
0.442
+0.120
+0.072
Appendix
Table 8: Second textbook (198 trajectories over 42 tasks for 4B; 200 over 44 for 9B; one epoch), seed harness, same machine and night per block, 4 rollouts. Top: LoRA versus full-parameter fine-tuning for both sizes. Bottom: rank ablation on 9B. Each adapter in the bottom block is scored against the base model measured on its own machine on the same night (the base drifts by up to 0.035 within seven hours on one machine), so its gains are comparable across ranks but are not differences from the 0.306 / 0.370 row above, and the rank-64 absolute scores are not comparable to the other rows.
Model
Adapter
Split
Delivery
Healthy
F1 evidence
F2 harness friction
F3 execution control
F4 planning capacity
Qwen3.5-4B
none
dev 48
52%
15.6%
10.9%
3.6%
57.8%
12.0%
Qwen3.5-4B
LoRA qkvo
dev 48
78%
35.4%
8.9%
2.1%
39.6%
14.1%
Qwen3.5-4B
LoRA qkvo+MLP
dev 48
67%
26.6%
12.0%
4.7%
51.0%
5.7%
Qwen3.5-4B
none
held-out 72
55%
15.3%
9.0%
3.5%
64.6%
7.6%
Qwen3.5-4B
LoRA qkvo
held-out 72
78%
30.6%
8.7%
2.1%
45.1%
13.5%
Qwen3.5-4B
LoRA qkvo+MLP
held-out 72
62%
30.9%
8.7%
3.5%
48.6%
8.3%
Appendix
Table 9: Failure composition of the base model and the LoRA adapters under the seed harness, same machine and night as Table 1 (share of trajectories, 4 rollouts; labels as in Section 3.1 ; rows sum to 100% up to rounding). The 4B adapter moves delivery and F3; the 9B adapter moves F4.
Development (48 tasks)
Held-out (72 tasks)
Weights
Score
Delivery rate
Score
Delivery rate
Rounds per trajectory
Base (no adapter)
0.333
98%
0.379
99%
19
Real LoRA
0.399
not recorded
0.426
not recorded
not recorded
Placebo LoRA
0.175
86%
0.152
76%
52–65
Placebo minus base
−0.158
−12 pp
−0.227
−23 pp
+33 to +46
Appendix
Table 10: Placebo control on Qwen3.5-9B: the real adapter and a placebo adapter trained with the identical recipe (rank 32, α=64 , qkvo+MLP, one epoch) on the same trajectories, except that the placebo’s assistant outputs are permuted across samples. All cells on the same machine and night, seed harness, 4 rollouts. Rounds = mean number of agent turns per trajectory, pooled over both splits; the placebo hit the 100-turn cap in 9 of 96 development and 44 of 288 held-out trajectories.
Figure 6: Failure composition of the base model and the LoRA adapters under the seed harness (the cells of Table 9 ; numbers above the bars are delivery rates). The 4B adapters raise delivery and shrink the loop class F3; the 9B adapters shrink the content class F4.
Cell
Removes
Dev score
Held-out score
Dev vs full
Held-out vs full
Dev delivery
Dev score among delivered
Full
nothing
0.366
0.438
99%
0.368
− Checkpoint
mid-trajectory checkpoint
0.385
0.440
+0.019
+0.002
− Gate
exit gate
0.384
0.410
+0.017
−0.028
− Hard blocks
four hard blocks
0.383
0.432
+0.017
−0.005
− Lessons
evolved lessons file
0.317
0.362
−0.049
−0.076
Bare
everything
0.290
0.346
−0.076
−0.092
85%
0.340
Appendix
Table 11: Component ablation on Qwen3.5-9B (closing harness of the 9B line; seven cells, 4 rollouts each, two identical machines on one night, cross-machine offset ≤0.03 ). “Removes ledger” also removes the exit gate and the checkpoint, which depend on it. Differences below 0.035 are within noise.
Cell
Removes
Dev score
Held-out score
Dev vs full
Held-out vs full
Full
nothing
0.264
0.314
− Evolved middleware
citation auto-fix, submission nag
0.310
0.313
+0.046
−0.001
− Evolved rewrites
three policy and prompt rewrites
0.302
0.287
+0.038
−0.027
− Checkpoint
mid-trajectory checkpoint
0.270
0.330
+0.007
+0.016
− Gate (4 rollouts)
exit gate
0.290
0.237
+0.026
−0.077
− Ledger
state ledger
0.253
0.254
−0.011
−0.060
Appendix
Table 12: Component ablation on Qwen3.5-4B (closing harness of the 4B line; ten cells, one machine, one boot, 2 rollouts per cell except the exit-gate cell with 4). With two rollouts, differences below 0.05 are within noise. “Seed kit” keeps the seed harness and removes every product of evolution.
Figure 7: Selection optimism in the two evolution lines (Proposition 1 ). Each point is a round with an accepted edit: the retained candidate’s own 4-rollout reading against the fresh anchor of the same harness measured afterwards; band =±σ around the diagonal. Points below the diagonal are readings that overstated the harness.
Dev score
Test, 117 unseen tasks
Harness
Weights
(48 tasks)
Score
Delivery
Healthy
F3
F4
Accuracy among delivered
Seed
Base
0.271
0.256
49%
23.1%
45.9%
25.6%
0.47
Evolved
Base
0.375
0.344
77%
33.8%
40.2%
18.2%
0.44
Seed
LoRA
0.253
0.216
42%
20.3%
52.6%
19.4%
0.49
Seed
Full-parameter
0.250
0.231
45%
20.7%
47.0%
23.9%
0.46
Evolved
LoRA
0.349
0.327
81%
31.4%
38.7%
17.1%
0.39
Appendix
Table 13: WebArena-Lite with Qwen3.5-9B: seed harness versus the closing harness of the 13-round line, and the adapters trained on the closing harness’s trajectories (4 rollouts per cell; the resolvable difference on this benchmark is 0.064). Delivery = the agent returned an answer; healthy, F3 and F4 = share of trajectories labelled as in Section 3.1 ; accuracy = judged-correct share among delivered answers. Base-model cells: one night, one site host; adapter cells: the following night on that host and a clone, with the closing-harness base anchor on the development tasks re-measured alongside them (0.375; 0.396 the night before).
Model
Harness
Score
Delivery
Score among delivered
Healthy
F1 evidence
F2 harness friction
F3 execution control
F4 planning capacity
MiniCPM5-2B
seed
0.106
61%
0.173
5.7%
42.7%
2.1%
32.3%
17.2%
MiniCPM5-2B
evolved (5 rounds)
0.110
71%
0.156
4.7%
35.9%
2.6%
34.4%
22.4%
Spark-X2.5-4B
seed
0.153
97%
0.158
6.2%
20.8%
2.1%
33.3%
37.5%
Spark-X2.5-4B
evolved (5 rounds)
0.155
95%
0.164
5.2%
6.8%
0.0%
28.1%
59.9%
Ministral-3-8B
seed
0.112
90%
0.125
0.0%
28.1%
0.0%
7.3%
64.6%
Ministral-3-8B
evolved (5 rounds)
0.127
93%
0.136
3.1%
30.7%
0.5%
5.2%
60.4%
Appendix
Table 14: Failure composition of the three smaller non-Qwen lines at their seed and closing anchors (48 development tasks, 4 rollouts; share of trajectories, labels as in Section 3.1 ; rows sum to 100% up to rounding; computed after the fact from the anchor trajectories; the Ministral seed anchor logged 178 of 192 trajectories). Compare the Qwen seed rows of Table 9 : healthy 15–34%, F4 12–28%. On all three the healthy share is at most 6%, and on Spark and Ministral F4 is the largest class before any evolution.
Model
Harness
D score
H score
Delivery D / H
Mean steps D / H
Quality D
Kimi K3
Bare
0.418
0.402
96.9% / 95.2%
19.7 / 23.3
0.431
Kimi K3
Evolved
0.476
0.451
99.3% / 98.7%
8.4 / 10.1
0.479
GLM 5.3
Evolved
0.512
0.488
99.6% / 99.1%
7.6 / 9.2
0.514
Appendix
Table 15: Kimi K3 and GLM 5.3 as the acting agent (four-rollout means on 48 development (D) / 72 held-out (H) tasks, run by a co-author on a separate machine; quality = mean score among delivered development plans). GLM 5.3 has no bare control; no confidence intervals are given because per-run readings are not attached.