cs.CLMay 28, 2026

Scaling Laws for Agent Harnesses via Effective Feedback Compute

Authors: Xuanliang ZhangDingzirui WangKeyan XuQingfu ZhuWanxiang Che

Organizations: Harbin Institute of Technology

Abstract

Agent harnesses shape language-model performance by controlling tool use, feedback, verification, memory, and repair. Yet raw test-time expenditure, such as tokens, tool calls, wall time, or cost, cannot distinguish useful feedback from redundant or unstable interaction. We introduce \emph{Effective Feedback Compute} (EFC), a trace-level scaling coordinate for informative, valid, non-redundant, and retained feedback. We further define Estimated-EFC, NRS-EFC, harness efficiency ηη, and task-demand normalization for realistic traces and heterogeneous tasks. Across synthetic, real, held-out, and prospective evaluations, EFC-based coordinates outperform raw-compute baselines and SAS. Oracle-EFC/DtaskD_{\mathrm{task}} reaches R2=0.99R^2=0.99 in controlled scaling, and NRS-EFC/DtaskD_{\mathrm{task}} reaches R2=0.93R^2=0.93 on real traces where raw compute has near-zero or negative fit. Finally, \ours uses EFC as a companion control layer for existing harnesses, improving mean pass rate from 61.2%61.2\% to 68.2%68.2\% while reducing mean raw cost from 213.8213.8 to 85.185.1 under matched settings. These results suggest that harness scaling depends on durable, task-sufficient feedback rather than raw computation alone.

Explore similar work

CardsList
  1. A2EA^2E : An End-to-End Agent Auditing Engine

    Aug 7, 2026Haoning Wang, Mingxun Zhang, Chenyue Yu +4Agent HarnessAgentic Evaluations