cs.CLSep 30, 2026

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Authors: Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan, Yonggan Fu, Jindong Jiang, Mingjie Liu, Ehsan Hosseini-Asl, +3 more

Organizations: NVIDIA · KAIST

Abstract

Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Scaling Test-Time Compute for Agentic Coding

    Apr 16, 2026Joongwon Kim, Wannan Yang, Kelvin Niu +13Long-Horizon AgentsCoding Agents

  2. VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

    Oct 1, 2026Caiqi Zhang, Rujun Han, Zifeng Wang +4Long-Horizon AgentsLong-Horizon Task Planning

  3. Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

    May 27, 2026Yilun Yao, Xinyu Tan, Chao-Hsuan Liu +9Agent HarnessAgentic Benchmarks