Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This practice overlooks that a trajectory's value for model training depends on the harness under which it was generated. To systematically analyze this interface, we establish an alternating co-evolution framework that decouples harness search and policy training through component-wise promotion decisions. Within this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh. Under CoTrace, recurring execution failures guide harness synthesis, while policy training is strictly conditioned on verified rollouts matched to the adopted runtime for supervised fine-tuning (SFT) or fresh online interactions for reinforcement learning (RL). On the Tmax promotion split, CoTrace advances Qwen3.5-9B from 78 to 88 solved tasks under supervised fine-tuning while an online reinforcement variant reaches 90. Specifically, a compact harness-matched corpus produces steady model gains at substantially lower compute than much larger corpora pooled across sibling harnesses. Furthermore, evaluations on Terminal-Bench 2.1 and SWE-bench Lite show that out-of-distribution transfer depends fundamentally on harness compatibility, where maintaining consistency between training and evaluation runtimes prevents procedural execution breakdowns observed under foreign scaffolds.
Figures & tables
Figure 1: Overview of CoTrace and the shared data interface in model–harness co-evolution. Bottom: the closed loop. A task agent with the adopted harness H∗ and policy weights θ solves executable terminal tasks and fills a shared history of trajectories; harness evolution reads that history to edit prompts, processors and tools and selects the next H∗ , while model training updates θ before the subsequent round of search. Top: the data recipes that turn the history into model-training data, from the baseline supervised fine-tuning (SFT) recipes that pool every successful search trajectory (all-evolve) or successes across sibling candidate harnesses (mixed siblings), to CoTrace-SFT, which keeps only trajectories whose provenance matches H∗ and tops them up with fresh rollouts under H∗ , and CoTrace-RL, which learns online by reinforcement learning (RL) from rewards on frontier tasks rolled out under H∗ .
Setting
Score
Decomposition
Cost
Method
Harness search
Data recipe
Init
Evolved
Δ
ΔH
ΔM
GPU-h / iter
Qwen3.5-9B-Thinking
Harness only a
Tournament 5+5
—
77
81
+4
+4
—
25
Co-evolve w. SFT
Tournament 5+5
Mixed siblings b
77
86
+9
+9
0
54
CoTrace-SFT
Tournament 5+5
Winner-only + SFT-gen c
78
88
+10
+4
+6
47
CoTrace-RL d
Tournament 5+5
Winner-only harness
78
90
+12
+5
+7
62
Table 1: Harness–model co-evolution with different data recipes. We report the number of solved tasks on the frozen 102-task promotion split of Tmax ( Ivison et al., 2026 ) ; ΔH / ΔM are cumulative harness and model gains, Cost is GPU-hours per iteration on one 8-GPU node. Shaded rows are the CoTrace recipes (blue supervised, orange reinforcement); bold / underline mark best and second-best per scale. For non-tournament baseline comparisons using sequential harness search, see Table 9 .
Figure 2: Dynamics of model–harness co-evolution. (a) Each round’s artifact before promotion (harness-only: the round’s best harness candidate; co-evolution lines: that round’s candidate model checkpoint). (b) The 9B reinforcement chain stage by stage; segment labels are the net task change or the paired win/loss record where recorded ( Table 11 ). The incumbent reaches 90; the next reinforcement candidate scores 85 and is rejected.
Mixture
Traj. (tasks) per iter.
Per-task cap
Harness- matched
Prompt conditioned
Fresh top-up
History cap
9B: accepted, ΣΔM
4B: ΣΔM
Mixed siblings
149–308 (38–88)
8
×
×
×
—
0/3, 0
0
CoTrace-SFT: winner-only + SFT-gen
30–50 (30–50)
1
✓
✓
✓
40%
2/5, +6
0
CoTrace-RL: online frontier
2,304 ep.
—
on-policy
—
on-policy
—
3/4, +7
+12
Table 2: Training data from each mixture and the resulting promotion decisions. Trajectories and unique tasks per iteration ( Table 10 ), the construction choices of Section 2.2.1 , and accepted model updates out of those attempted with their total ΣΔM .
Figure 3: Promotion decisions and the moving curriculum. (a) CoTrace-SFT one promotion decision at a time: filled markers are installed updates (teal harness, amber model), hollow markers rejected candidates. (b) Two chains: the evolve-set score falls as solved tasks retire while the promotion score rises.
Harness
θ0 (base)
θ1 (installed)
previous Ht
78
82
adopted Ht+1
82
86
Table 3: Completed cross-evaluation. Promotion-split scores for CoTrace-SFT, iterations 1–2.
Figure 4: Trajectory constructions and model-side gain. (a) Trajectories per iteration against the accepted model gain ( Table 2 ). (b) Every model update assessed for promotion, shown as its change against the incumbent, by mixture ( Table 8 ).
Figure 5: Transfer under two runtimes. (a) TB2.1, three trials under the baseline harness: pass@1 (per-trial solve rate, mean ± sd over the three trials; dots) and pass@3 (tasks solved in at least one trial; bars). (b) SWE-bench Lite resolved rate per checkpoint under the third-party mini-swe-agent scaffold (hatched) and its co-evolved harness (solid), with the change in no-patch instances, on which the agent ends without emitting a patch.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Expanded view of the CoTrace workflow. A meta-agent searches for an improved harness H∗ ; the task agent executes it to produce verified trajectories; these data support SFT and RL; and the updated model initiates the next harness-search round.
Role
Tasks
Used by
Optimized against?
Evolve (rotating)
50 / iter
harness search, corpus harvest
yes
RL train
≤ 100
online RL rollouts
yes
Promotion
102
both promotion decisions, all reported numbers
no
Terminal-Bench 2.1
89
transfer reporting only
no
Appendix
Table 4: Data roles on Tmax ( Ivison et al., 2026 ) . No task appears in more than one role, and the promotion split is never optimized against by the harness search, the corpus builder, or the model update.
Chain (9B; Tables 1 and 9 )
Iters
GPU-h/iter
ΔH
ΔM
Δ
GPU-h/task
Harness only
1
25
+4
—
+4
6
All-evolve, fixed-50
3
77
+1
+4
+5
46
All-evolve + retry
3
118
+1
+6
+7
51
Mixed siblings
3
54
+9
0
+9
18
CoTrace-SFT
5
47
+4
+6
+10
24
CoTrace-RL
4
62
+5
+7
+12
21 (35)
Appendix
Table 5: What each lever cost, and what it returned. The 9B chains of Tables 1 and 9 : completed outer iterations, GPU-hours per iteration, and GPU-hours per installed task (iterations × GPU-hours per iteration, divided by Δ ; bracketed: the reinforcement line’s last measurement). Iteration counts and totals are from each chain’s per-stage wall-clock log; the harness-only chain is one iteration of four evolve rounds. A chain’s cost is dominated by evaluation (37–57%) and harness search (14–57%); the gradient step is 4%, SFT-gen top-up 2–35% and the reinforcement stage 63% where they run ( Appendix C ).
Harness search
SFT
Tournament
2×5 candidates
Method
LoRA
Extra rounds on failure
2
LoRA rank / α
32 / 64
Max SFT retries
2
LoRA dropout
0.05
Meta-agent step cap
200
Targets
q,k,v,o,gate,up,down
Meta-agent wall clock
3600 s
Epochs
2
Evolve set size
50
Learning rate
2×10−5
Appendix
Table 6: Full configuration. The reinforcement-stage settings are those of the line in Figure 2 (b); Table 15 gives its rollout budget.
Observed pattern
Attribution
Conf.
Destination
No model input tokens consumed (environment never started)
env
0.97
environment repair
Grader or verifier exception
grader
0.90
benchmark repair
Session log absent
—
0.00
quarantine
Clean success
model
0.80
full-trajectory SFT
Success after a recovered failure
model
0.80
recovery-slice SFT
Failed with no error, step budget exhausted
ambiguous
0.45
quarantine
Appendix
Table 7: Attribution rules of the router, in application order, with the confidence each assigns and the destination it selects. Successes are routed too: a success that recovered from an earlier failure becomes a recovery-slice demonstration.
Chain ( Tables 1 and 9 )
Iter
Before
Decisions, in order
Harness only
1–4
77
best candidate per round: 80, 78, 80, 81 (no ratchet)
All-evolve, fixed-50
1
77
H 75 reject; M 81 accept
2
81
H 82 accept; M 82 / 81 / 82 reject
All-evolve + retry
1
77
H 78 accept; M 82 accept
2
82
H 81 reject; M 81 reject / 84 accept
3
84
H 83 reject; M 80 / 83 / 80 reject
Appendix
Table 8: Promotion ledger for every chain of Table 1 , in decision order. H : candidate harness scored with the incumbent model (stage B); M : candidate checkpoint scored under the adopted harness (stage E). Bold marks a change the promotion procedure installed; a slash separates retries after additional evolve rounds. The ledger is consistent with Tables 1 and 2 .
Recipe
Search
Traj. (tasks)/iter
Init
Evolved
Δ
ΔH
ΔM
Accepted, ΣΔM
GPU-h/iter
All-evolve
Seq., fixed-50
90–100 (34–43)
77
82
+5
+1
+4
1/5, +4
77
All-evolve + retry
Seq., rotating-50
86–220 (36–89)
77
84
+7
+1
+6
2/6, +6
118
Appendix
Table 9: The two sequential-search chains: promotion-split scores and gains as in Table 1 , corpus per iteration as in Table 2 . Pooled search trajectories, up to 3 per task; retry adds evolve rounds and rebuilds the corpus after a rejected model update.
Recipe
Search, evolve set
Traj.
Tasks
Pairs
Per-task cap
Matched
Accepted updates
ΣΔM
Qwen3.5-9B
All-evolve
Seq., fixed-50
90 – 100
34 – 43
361 – 532
3
×
1 / 5
+4
All-evolve + retry
Seq., rotating-50
86 – 220
36 – 89
392 – 1189
3
×
2 / 6
+6
Mixed siblings
Tourn. 5+5
149 – 308
38 – 88
618 – 1335
8
×
0 / 3
0
CoTrace-SFT
Tourn. 5+5
30 – 50
30 – 50
526 – 948
1 – 2
✓
2 / 5
+6
Qwen3.5-4B
Appendix
Table 10: Corpus manifests behind Table 2 : trajectories, unique tasks and prompt/completion pairs offered to the trainer, first to last iteration of each chain.
Stage
Paired comparison: wins / losses
Paired comparison: net
Plotted stage scores: Δ pass count
Harness 1
7 win / 7 loss
0
0
RL 1
8 win / 8 loss
0
+1
RL 2
14 win / 9 loss
+5
+4
RL 3
10 win / 8 loss
+2
+2
Harness 4
3 win / 1 loss
+2
+2
RL 4
4 win / 7 loss
−3
−5
Appendix
Table 11: Paired per-task records of the reinforcement line, where recorded. Paired comparison columns are from the promotion evaluation; the last column is the difference of the stage scores of Figure 2 (b), which come from separate evaluation executions.
mini-swe-agent (one harness for all)
co-evolved (model, harness) pair
Checkpoint
resolved
unresolved
no patch
format fail.
resolved
unresolved
no patch
finish reasons †
Base
104 (34.7%)
88
108
55
112 (37.3%)
129
59
143 / 147 / 10
+ SFT
74 (24.7%)
71
155
57
105 (35.0%)
131
64
159 / 135 / 6
+ RL
110 (36.7%)
106
84
15
123 ( 41.0% )
141
36
178 / 122 / 0
† no tool calls / error / budget exceeded.
Appendix
Table 12: SWE-bench Lite (300 instances) for the 9B checkpoints of Table 1 under mini-swe-agent and under the harness each chain adopted. No patch counts instances for which no diff reached the grader. The finish-reason histogram of the co-evolved runs is no-tool-calls / error / budget-exceeded, the counterpart of mini-swe-agent’s format-error and limits-exceeded split.
Figure 7: Promotion ablation and the mixed-siblings trace. (a) Without component-wise promotion, a 4B chain falls 26.7% → 0.0% on its evolve set in five rounds; with promotion, it holds a 23–29% band. (b) The mixed-siblings chain, one promotion decision at a time.
Proposed lever
Share
Add a processor (control)
61%
of which loop / repeat breaker
37%
of which verifier dependency
16%
of which step budget / self-verify
13%
of which other
30%
Prompt rewrite (instruction)
20%
Appendix
Table 13: Left: what was proposed across 216 changesets. Right: what the promotion rule actually kept , over 19 extra processor slots across 10 accepted incumbents.
This paper
Tmax
Reward, advantages, objective
outcome-only, centered group-relative, DPPO
same
Trust region
binary TV, δ=0.1 , rollout log-probs
same
Zero-variance groups
dropped; active sampling, ≤ 8 groups
same
LM head / learning rate
FP32 / 1×10−6 , constant
same
Samples / unique prompts
8 / 4 (4B: 16 / 2)
32 / 8
Async steps
2
4
Appendix
Table 14: Our reinforcement stage against Tmax’s published Qwen3.5-9B recipe ( Ivison et al., 2026 ) . The algorithmic core is shared; the regime differs.
Rollout
Optimization
Episodes / iteration
2,304
Unique prompts / step
4
Training tasks
100
Samples / prompt
8
Response length
16,384
Async steps
2
Per-turn budget
4,096
Sampled prompt groups
≤ 8
Max agent steps
40
Zero-std groups
dropped ( Yu et al., 2025 )
Temperature
1.0
Trust region
binary TV, δ=0.1
Appendix
Table 15: Reinforcement stage settings of the 9B line. The task plane is drawn from the taxonomy and is disjoint from the promotion split; frontier exploration spends a fixed fraction of each round on tasks the current pair has not solved, which is the curriculum idea of Section 2.2.3 applied inside the optimizer rather than around it.
Corpus construction
Trajectories
Solved ↑
Base policy, no SFT
—
3.00
Undifferentiated mix
107
4.00
Balanced by task taxonomy
400
4.33
Routed by capability need
400
6.00
Appendix
Table 16: Routing pilot. Qwen3.5-9B, baseline harness, 15 Terminal-Bench 2.1 tasks, mean of 3 runs.
Round
Observed failure
Harness edit
Effect
4B R1
Bash calls written as text, never executed
inline tool-call recovery
0 → 2
31B R2
Backend rejected the tool transport; all tasks died at step 0
transport bypass
0 → 3
31B R3
Model did not create the files the verifier required
required-output guard
3 → 2
31B R4
Model assumed unavailable commands existed
environment preflight
2 → 4
Appendix
Table 17: Four grounded harness edits on the evolve subset. Every diagnosis was correct; not every intervention helped. This is why proposal and promotion are separated, and why promotion is deterministic.
Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research trajectories are generated, evaluated, and learned from. Existing pipelines typically train models under a fixed harness, including prompts, tools, skills, middleware, and memory, while leaving the data-generating process outside the optimization objective. This creates a mismatch between model updates and the static scaffolding that determines trajectory quality. We introduce Co-Harness, a framework that jointly optimizes the agent harness and model parameters during post-training. Co-Harness alternates between harness optimization and model optimization. An LLM-based HarnessCritic analyzes failed trajectories, identifies harness-level failure modes, and proposes validated local updates. The model is then fine-tuned on high-quality trajectories generated by the improved harness, distilling effective scaffolding into model parameters. A 200+ hour autonomous case study further shows that Co-Harness can recover from system crashes, improve inference efficiency, and discover ensemble strategies without human intervention. These results suggest that joint harness and model optimization is an effective way to improve agents beyond fixed-harness post-training.
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic task success. Automated harness evolution can enable smaller models to perform well on domain-specific tasks at a fraction of frontier-model cost. Since both the harness and model weights shape behavior, we ask how harness evolution and lightweight fine-tuning should be combined. Across seven enterprise agent tasks, we first evolve a harness with the weaker model, then find that a stronger expert often uses it more effectively, suggesting expert supervision could close the remaining gap. However, training the weaker model on the expert's complete trajectories under the evolved harness backfires: performance regresses on all seven tasks by 4 to 30 points across Qwen3-Coder and Gemma 4, even though the same procedure helps under the unevolved harness. Our analysis shows that imitation transfers knowledge and increases scaffold usage, but disrupts model-harness fit: the weaker model adopts the expert's planning strategy without the competence to execute it and no longer matches the harness evolved around its native planning style. We therefore develop an on-policy expert-correction pipeline, automated by a meta-level MLE agent, that localizes the failing turn in the weaker model's own rollout and asks the expert to rewrite only that turn. This preserves the model's planning style and combines the gains of harness evolution and model adaptation. Our results identify and resolve a source of contention between harness and weight updates, yielding a compatibility-preserving recipe for economical co-evolution on domain-specific enterprise tasks.
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a task-specific manner can improve execution-trace quality while remaining computationally lightweight and requiring only a few update iterations. To this end, we introduce Recursive Harness Self-Improvement (RHI), which represents the harness as a prompt-level specification of the agent loop and iteratively refines it using pairwise feedback over its own revision history. Across 30 synthetic machine-learning research tasks spanning quantitative finance, robotics, and pharmacy, a few RHI iterations suffice to substantially raise the performance ceiling of low-reasoning-effort agents, exceeding the corresponding maximum-reasoning-effort setting while reducing inference cost by up to 60%. We show that these gains arise primarily from improved task-specific context management through more effective inter-agent information flow rather than longer reasoning traces. Finally, we formalize this behavior as an information-theoretic hypothesis for RHI's implicit optimization objective, suggesting RHI as a practical algorithm for continual learning within the paradigm of model--harness co-evolution.