Finding the Right Fit: Model-Harness Interactions across Agent Tasks
Organizations: College of Computing and Data Science, Nanyang Technological University, Singapore
Abstract
Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes. We evaluate 66 configurations: four configurable harnesses (OpenHands, DeepSeek Harness, PI, and openJiuwen) paired with five models on TUA-Bench, ALE-CLI, and Terminal-Bench 4, plus the native Codex-GPT and Claude Code-Claude pairings. Model rankings reverse across harnesses. On Terminal-Bench 4, Claude leads GPT by 7.94 points in OpenHands but trails it by 30.16 points in PI. For four of the five models, the best harness changes from one benchmark to another, yet some pairings hold: openJiuwen gives Kimi its highest score on all three benchmarks, by 5.61 to 11.11 points. A model's own vendor harness is not reliably its best, and higher cost does not reliably buy a higher score. On Terminal-Bench 4, GPT scores higher under PI than under DSH at less than a quarter of the cost per task. Matched trajectories suggest why fit varies. Models start almost all repairs themselves, so much depends on whether the harness hands failures back in a form the model can use. GPT does best with PI's lean scaffold, while Kimi, which often issues malformed tool calls, does best in openJiuwen. We argue that the model, the harness, and the task should be evaluated together, and we release the harness adapters, evaluation code, and all 6,204 scored trajectories at https://github.com/liyix/finding-the-right-fit and https://huggingface.co/datasets/yixuanli97/finding-the-right-fit.
Figures & tables
| Benchmark | Model | OpenHands | DSH | PI | openJiuwen | Codex | Claude Code |
|---|---|---|---|---|---|---|---|
| TUA-Bench | Claude Opus 5 | 0.6523 | 0.5884 | 0.5821 | 0.6539 | — | 0.6889 |
| GPT-6 Astra | 0.6101 | 0.6432 | 0.6538 | 0.6635 | 0.6368 | — | |
| GLM-5.3 | 0.6286 | 0.5616 | 0.5357 | 0.6301 | — | — | |
| Kimi K3 | 0.5054 | 0.5838 | 0.5878 | 0.6439 | — | — | |
| DeepSeek V4 Pro | 0.5649 | 0.5193 | 0.5758 | 0.6259 | — | — | |
| ALE-CLI | Claude Opus 5 | 0.5314 | 0.4348 | 0.5022 | 0.5127 | — | 0.5428 |
| Benchmark | Native pair | Score | Best alternative | Score | |
|---|---|---|---|---|---|
| TUA-Bench | CC–Claude | 0.6889 | openJiuwen | 0.6539 | -3.50 |
| TUA-Bench | Codex–GPT | 0.6368 | openJiuwen | 0.6635 | +2.67 |
| ALE-CLI | CC–Claude | 0.5428 | OpenHands | 0.5314 | -1.14 |
| ALE-CLI | Codex–GPT | 0.5630 | PI | 0.5846 | +2.15 |
| Terminal-Bench 4 | CC–Claude | 0.4921 | OpenHands | 0.5714 | +7.94 |
| Terminal-Bench 4 | Codex–GPT | 0.5556 | PI | 0.6032 | +4.76 |
| Model | TUA-Bench | ALE-CLI | Terminal-Bench 4 |
|---|---|---|---|
| Claude Opus 5 | Claude Code | Claude Code | OpenHands |
| GPT-6 Astra | openJiuwen | PI | PI |
| GLM-5.3 | openJiuwen | OpenHands | DSH |
| Kimi K3 | openJiuwen | openJiuwen | openJiuwen |
| DeepSeek V4 Pro | openJiuwen | OpenHands | DSH |
| Observed behavior (initiator) | Data / training implication | Runtime harness implication |
|---|---|---|
| Kimi re-issues edit calls without required arguments until OpenHands’ stuck detector ends the run; the same model repairs failures under PI (TB4; M, then S) | Collect malformed call error corrected call; train argument adherence. Do not treat the system termination as a model failure to recover | Return argument errors as structured feedback and warn before terminating; offer separate write and edit tools |
| All three GPT runs hand a CAPTCHA to the user; only openJiuwen’s generic continue message leads to solving the audio challenge (TUA 042; M, then P) | Label prompted decisions; do not select imitation data by reward alone; keep policy-compliant hand-backs as positive examples | Scope continuation prompts; provide an explicit escalation channel for steps that need a human |
| Kimi substitutes its own acceptance check and reports success; the re-check round reuses the same oracle (TB4 retro-console-soc ; M, then P) | Contrastive pairs (DSH success, same task) at the acceptance-test decision; train alignment of evidence with stated criteria | A generic re-check prompt is insufficient; expose checkable deliverable criteria where the task permits |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Harness | Version | Tools | Context management | Turn limit |
|---|---|---|---|---|
| OpenHands | SDK 1.44.1 | terminal, file editor, task tracker, think, finish | 1M-token window, no condensation | 500 steps; stuck detector |
| DSH | 0.1.1-rc.2 | 26 standard tools | native compaction at 80% of the window | none |
| PI | 0.84.4 | read, bash, edit, write | native compaction near the window limit | none |
| openJiuwen | 0.1.18 | read, write, edit, glob, list, grep, bash | native context compression | 8 outer rounds |
| Codex | 0.150.1 | native code-execution tool | native compaction (272K window) | none |
| Claude Code | 2.1.251 | 21 native tools | native auto-compaction (1M window) | none |
| Model | OpenHands | DSH | PI | OJW | Native |
|---|---|---|---|---|---|
| TUA-Bench (120 tasks) | |||||
| Claude Opus 5 | 86.60 | 91.74 | 70.28 | 162.44 | 113.73 |
| GPT-6 Astra | 86.03 | 150.97 | 76.92 | 157.62 | 113.08 |
| GLM-5.3 | 51.86 | 56.47 | 29.82 | 45.15 | – |
| Kimi K3 | 46.85 | 47.83 | 33.51 | 56.67 | – |
| DeepSeek V4 Pro | 14.52 | 12.81 | 12.07 | 43.53 | – |