Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Organizations: NVIDIA · KAIST
Abstract
Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.
Figures & tables
| Verifier | Score MAE | Pairwise (%) | Verification (%) |
|---|---|---|---|
| Zero-shot | 2.59 | 59.01 | 38.52 |
| Distilled | 1.05 | 74.58 | 57.79 |
| Benchmark | Model | Harness | Base agent | Mid-Harness | |
|---|---|---|---|---|---|
| Zero-shot | Distilled | ||||
| TerminalBench-Lite | Qwen3.5-9B | Terminus-2 | 40.48 / 60.20 | 42.52 / 61.22 | – |
| TerminalBench-Lite | Nemotron3.5 Lightning | Terminus-2 | 41.16 / 55.10 | 43.20 / 58.16 | – |
| Terminal-Bench 2.1 | TMAX-9B | Vanillux2 | 21.72 / 25.84 | 27.34 / 35.96 | 26.59 / 35.96 |
| Terminal-Bench 2.1 | Nemotron3 Ultra | Terminus-2 | 50.94 / 65.17 | 56.18 / 66.29 | – |
| SWE-bench-Verified | TMAX-9B | Vanillux2 | 46.67 / 54.00 | 48.00 / 58.00 | 48.67 / 62.00 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | ||
|---|---|---|
| Generator | TMAX-9B | TMAX-9B |
| Generator thinking / temperature | On / 0.8 | On / 0.8 |
| Maximum agent steps | 64 | 64 |
| Generator context / output limit | 65,536 / 16,384 | 65,536 / 16,384 |
| Verifier | TMAX-9B, zero-shot or distilled | |
| Verifier thinking / temperature | Off / 0 | Off / 0 |
| Setting | Value |
|---|---|
| LoRA rank / alpha | 64 / 128 |
| LoRA dropout | 0.05 |
| Learning rate | |
| Epochs | 2 |
| Per-device batch size | 8 |
| Gradient accumulation | 1 |
| Verifier / configuration | Pass@1 | Pass@3 | |
|---|---|---|---|
| Base agent | 1 | 50.00 | 69.39 |
| First-runnable proxy | 8 | 49.66 | 66.33 |
| Zero-shot listwise | 4 | 49.32 | 66.33 |
| Zero-shot listwise | 8 | 51.02 | 67.35 |
| Zero-shot pointwise | 4 | 52.72 | 68.37 |
| Zero-shot pointwise | 8 | 52.38 | 67.35 |
| Configuration | Pass@1 | Pass@3 |
|---|---|---|
| Base agent | 50.00 | 69.39 |
| Best-of- , | 55.10 | – |
| Best-of- , | 57.14 | – |
| Best-of- , | 59.18 | – |
| SR, | 55.10 | 71.43 |
| SR, | 56.46 | 70.41 |
| Generator | Verifier | Pass@1 | Pass@3 | ||
|---|---|---|---|---|---|
| Reasoning decision-only | Reasoning decision-only | ||||
| TMAX-4B | Zero-shot | ||||
| Distilled | |||||
| TMAX-9B | Zero-shot | ||||
| Distilled | |||||
| TMAX-27B | Zero-shot | ||||
| Generator | Configuration | Pass@1 (%) | Cost (USD) | Cost reduction |
|---|---|---|---|---|
| TMAX-4B | Zero-shot Mid-Harness | 39.2% | ||
| Distilled Mid-Harness | 46.8% | |||
| TMAX-9B | Zero-shot Mid-Harness | 20.9% | ||
| Distilled Mid-Harness | 24.1% | |||
| + Best-of- , | 23.7% | |||
| + SR, | 22.4% |
| Variant | Reference | Pass@1 (%) ref. variant | Pass@3 (%) ref. variant | Verifier output ( reference) |
|---|---|---|---|---|
| Five responses per pair | Distilled, one response | |||
| Thinking enabled | Zero-shot pairwise | |||
| Rubric scaling | Zero-shot pairwise |
| Model | Mid-Harness | Pass@1 | 95% CI |
|---|---|---|---|
| 4B | Zero-shot | +2.72 | [-3.06, +8.84] |
| 4B | Distilled | +5.10 | [-0.68, +10.88] |
| 9B | Zero-shot | +4.76 | [-0.68, +10.20] |
| 9B | Distilled | +7.14 | [+1.36, +12.93] |
| 27B | Zero-shot | +2.04 | [-3.74, +7.48] |
| 27B | Distilled | +5.10 | [+0.34, +9.86] |
| Model | Short | Medium | Long |
|---|---|---|---|
| 4B | 2–17 (34) | 18–24 (37) | 25–64 (27) |
| 9B | 4–15 (33) | 16–24 (32) | 25–64 (33) |
| 27B | 3–10 (39) | 11–18 (26) | 19–64 (33) |