cs.ROSep 30, 2026

Scale and Selection: What Makes Automatic Harness Evolution Work for Visual-Interface Robot Agents

Authors: Zhijie Wei, Ferris Tan, Jinghui Wang

Organizations: Novaxbot

Abstract

When an off-the-shelf coding agent is used directly as a robot policy, observing a browser-based 3D interface through screenshots and acting by posing a virtual target gripper through a few tools, the agent's harness, its prompts, tools, and control rules, largely determines success, and until now it has been written by hand. We show that this harness can be improved automatically by another coding agent, the optimizer agent, and report two findings about what makes it work. First, the number of rollouts the optimizer agent sees per round governs whether the evolved harness is trustworthy, generalizes, and improves steadily. A single rollout is a noisy binary outcome, so with few rollouts per round a revision can be promoted on luck; enlarging the batch raises the signal-to-noise ratio of every promotion decision. Holding rounds fixed and growing the training set from 5 to 100 rollouts, held-out success rises from 47% to 67%, while small training sets overfit, reaching 70% on training tasks but only 54% held-out. Second, the optimizer agent must not be given free rein. With every revision it proposes accepted unconditionally, performance drifts downward within ten rounds as ill-judged edits accumulate; adding the most basic safeguard, Champion-Challenger selection that promotes a revision only if it strictly beats the incumbent on the same fixed evaluation set, turns the same loop into one that raises held-out success from 51% to 67% over 30 rounds. Automatic harness evolution for visual-interface robot agents is thus feasible, but its gains hinge on the rollout scale behind each decision and on how the optimizer agent's revisions are selected.

Figures & tables

Explore similar work

CardsList
  1. Rethinking the Evaluation of Harness Evolution for Agents

    Jul 14, 2026Yike Wang, Huaisheng Zhu, Zhengyu Hu +8Large Language Model AgentsTest-Time Scaling

  2. HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses

    Aug 3, 2026Luan Zhang, Ruochen Zhou, Dandan Song +9Evolving HarnessAgent Harness

  3. EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?

    Date pendingZixuan Ke, Vaidehi Patil, Haizhou Shi +9Evolving HarnessAgentic Benchmarks