Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes. We evaluate 66 configurations: four configurable harnesses (OpenHands, DeepSeek Harness, PI, and openJiuwen) paired with five models on TUA-Bench, ALE-CLI, and Terminal-Bench 4, plus the native Codex-GPT and Claude Code-Claude pairings. Model rankings reverse across harnesses. On Terminal-Bench 4, Claude leads GPT by 7.94 points in OpenHands but trails it by 30.16 points in PI. For four of the five models, the best harness changes from one benchmark to another, yet some pairings hold: openJiuwen gives Kimi its highest score on all three benchmarks, by 5.61 to 11.11 points. A model's own vendor harness is not reliably its best, and higher cost does not reliably buy a higher score. On Terminal-Bench 4, GPT scores higher under PI than under DSH at less than a quarter of the cost per task. Matched trajectories suggest why fit varies. Models start almost all repairs themselves, so much depends on whether the harness hands failures back in a form the model can use. GPT does best with PI's lean scaffold, while Kimi, which often issues malformed tool calls, does best in openJiuwen. We argue that the model, the harness, and the task should be evaluated together, and we release the harness adapters, evaluation code, and all 6,204 scored trajectories at https://github.com/liyix/finding-the-right-fit and https://huggingface.co/datasets/yixuanli97/finding-the-right-fit.
Figures & tables
Benchmark
Model
OpenHands
DSH
PI
openJiuwen
Codex
Claude Code
TUA-Bench
Claude Opus 5
0.6523
0.5884
0.5821
0.6539
—
0.6889
GPT-6 Astra
0.6101
0.6432
0.6538
0.6635
0.6368
—
GLM-5.3
0.6286
0.5616
0.5357
0.6301
—
—
Kimi K3
0.5054
0.5838
0.5878
0.6439
—
—
DeepSeek V4 Pro
0.5649
0.5193
0.5758
0.6259
—
—
ALE-CLI
Claude Opus 5
0.5314
0.4348
0.5022
0.5127
—
0.5428
Table 1: Mean rewards on fixed task-set denominators (120, 99, 63). Codex and Claude Code appear only with their implemented one-to-one model pairing. Bold marks the highest recorded score in each model–benchmark row; underline marks the second highest.
Figure 1: Score versus model cost per task for all model–harness configurations on TUA-Bench, ALE-CLI, and Terminal-Bench 4. Cost per task uses OpenRouter prices (log scale). Shaded quadrants split each benchmark at the median cost and median score of its 22 configurations; the dotted line is the Pareto frontier.
Benchmark
Native pair
Score
Best alternative
Score
Δ
TUA-Bench
CC–Claude
0.6889
openJiuwen
0.6539
-3.50
TUA-Bench
Codex–GPT
0.6368
openJiuwen
0.6635
+2.67
ALE-CLI
CC–Claude
0.5428
OpenHands
0.5314
-1.14
ALE-CLI
Codex–GPT
0.5630
PI
0.5846
+2.15
Terminal-Bench 4
CC–Claude
0.4921
OpenHands
0.5714
+7.94
Terminal-Bench 4
Codex–GPT
0.5556
PI
0.6032
+4.76
Table 2: Native pairing versus the strongest observed alternative with the same model. Δ is alternative minus native reward on a 0–100 scale; negative values favor the native pairing. Δ is computed from unrounded scores.
Figure 2: Task-dependent fit for Kimi K3. (a) Mean reward for the four configurable harnesses on the fixed 120-, 99-, and 63-task sets. (b) Paired task outcomes for openJiuwen versus the strongest observed alternative in each collection (PI, PI, and DSH, respectively). (c) ALE-CLI mean paired differences by domain, shown for domains with at least three tasks; positive values favor openJiuwen.
Model
TUA-Bench
ALE-CLI
Terminal-Bench 4
Claude Opus 5
Claude Code
Claude Code
OpenHands
GPT-6 Astra
openJiuwen
PI
PI
GLM-5.3
openJiuwen
OpenHands
DSH
Kimi K3
openJiuwen
openJiuwen
openJiuwen
DeepSeek V4 Pro
openJiuwen
OpenHands
DSH
Table 3: Highest-scoring available harness for each model and task collection. Native references are included. Four models change their observed winner across collections, while Kimi retains openJiuwen.
Figure 3: Matched recovery case: Kimi K3 on TUA 056-move-textbox-left , final scored attempts. Both runs hit the same GIMP hang. openJiuwen’s shell returned it after 300 s and the model diagnosed and recovered (reward 1); PI’s shell never returned, so the run received no feedback and ended at the deadline (reward 0). The timeout and the re-check prompt come from the harness; all other actions are the model’s.
Figure 4: Manual review of the 45 failed Terminal-Bench 4 tasks under openJiuwen with Kimi K3. Every closing report states that each specified requirement has been verified, and in three the agent had observed a discrepancy and attributed it to a cause outside the deliverable.
Figure 5: Matched progress-to-delivery case on retro-console-soc with Kimi K3, where openJiuwen scores 0.00 and DSH scores 1.00. (a) Successive mismatch readings that each configuration obtained against its own acceptance criterion. Readings are consecutive within a run and are not aligned in time across runs; open markers are exact zeros, drawn on the axis floor, and the dashed segment marks the step that reaches zero. openJiuwen reaches zero by comparing its own console model at frame 60 against the RTL output at frame 40, whereas DSH reaches it by removing the remaining defect. (b) Per-test verifier outcome (filled, pass; crossed, fail) with the reward, tool calls, and wall-clock time of each run. The property asserted by openJiuwen’s final report is the one the grader rejects.
Observed behavior (initiator)
Data / training implication
Runtime harness implication
Kimi re-issues edit calls without required arguments until OpenHands’ stuck detector ends the run; the same model repairs failures under PI (TB4; M, then S)
Collect malformed call → error → corrected call; train argument adherence. Do not treat the system termination as a model failure to recover
Return argument errors as structured feedback and warn before terminating; offer separate write and edit tools
All three GPT runs hand a CAPTCHA to the user; only openJiuwen’s generic continue message leads to solving the audio challenge (TUA 042; M, then P)
Label prompted decisions; do not select imitation data by reward alone; keep policy-compliant hand-backs as positive examples
Scope continuation prompts; provide an explicit escalation channel for steps that need a human
Kimi substitutes its own acceptance check and reports success; the re-check round reuses the same oracle (TB4 retro-console-soc ; M, then P)
Contrastive pairs (DSH success, same task) at the acceptance-test decision; train alignment of evidence with stated criteria
A generic re-check prompt is insufficient; expose checkable deliverable criteria where the task permits
Table 4: Evidence-grounded implications. Initiator: M = model-initiated, P = after a harness prompt, S = performed by the system.
Table 5: Harness configuration. Entries are the harnesses’ own defaults, except the OpenHands context and output limits and the openJiuwen composition described in Section 3.2 . Tools are those on TUA-Bench and Terminal-Bench; on ALE-CLI every harness also receives the benchmark’s fourteen computer-use tools.
Figure 6: Score versus uncached input tokens per task (input minus cache reads, log scale) for the four configurable harnesses.
Figure 7: Per-task cost distribution by configuration (log scale). Boxes show the median and interquartile range; whiskers extend to 1.5 × IQR; outliers are not drawn.
Figure 8: Resource use by configuration, per task: model calls, input tokens per call, cache hit rate (cached share of input tokens), output tokens, and agent minutes. Each dot is one harness, and input includes cached tokens. Codex and Claude Code appear where their logs record tokens.
Figure 9: Median model calls (top) and agent minutes (bottom) on solved (reward 1) versus failed (reward 0) tasks. Each dot is one configuration with at least three tasks in each group; points above the diagonal spend more on the tasks they solve.
Model
OpenHands
DSH
PI
OJW
Native
TUA-Bench (120 tasks)
Claude Opus 5
86.60
91.74
70.28
162.44
113.73
GPT-6 Astra
86.03
150.97
76.92
157.62
113.08
GLM-5.3
51.86
56.47
29.82
45.15
–
Kimi K3
46.85
47.83
33.51
56.67
–
DeepSeek V4 Pro
14.52
12.81
12.07
43.53
–
Appendix
Table 6: Agent-model cost (USD, OpenRouter pricing) of each configuration, summed over the last run of every task. Dividing by the task count (120/99/63) gives the cost per task in Figure 1 . OJW denotes openJiuwen; Native is Claude Code for Claude Opus 5 and Codex for GPT-6 Astra.