cs.SESep 29, 2026

AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems

Authors: Yiming Cheng, Alfin Wijaya Rahardja, Mengshi Zhang, Zihao Chen, Zhenpeng Chen, Yiling Lou

Organizations: The University of Chicago · Fudan University · TensorBlock, Inc · Tsinghua University · University of Illinois Urbana-Champaign

Abstract

Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number of executable harness bugs while requiring hundreds of human hours to construct. This work presents AgentBug-Smith, an automated harness bug reproduction approach that continuously discovers and reproduces real-world harness bugs from open-source agentic systems. Across different backbone LLMs, AgentBug-Smith consistently outperforms existing bug reproduction techniques designed for general software, achieving 10.67% - 27.56% higher success rates of reproducing harness bugs. By applying AgentBug-Smith to open-source agentic systems in the wild, we construct Live-Harness-Bench, a live and extensible benchmark that currently contains 200 reproducible harness bugs. We further demonstrate the utility of Live-Harness-Bench through two downstream applications. First, we use Live-Harness-Bench as the evaluation benchmark to systematically evaluate state-of-the-art software agents, revealing their limited capabilities in repairing real-world harness bugs. Second, we use Live-Harness-Bench as a knowledge base of real-world harness bug fixes, from which reusable repair skills can be distilled to improve existing software agents, increasing their harness-bug repair rates by 6.32%. Together, AgentBug-Smith and Live-Harness-Bench establish a scalable foundation for continuously evaluating and improving software agents on harness bug repair, turning real-world agent failures into executable evaluation instances and reusable knowledge for harness improvement, thus contributing to the ultimate goal of recursively self-improving agents.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws

    Jun 4, 2026Mengzhuo Chen, Junjie Wang, Zhe Liu +3Agent HarnessLarge Language Model Agents

  2. Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

    May 27, 2026Yilun Yao, Xinyu Tan, Chao-Hsuan Liu +9Agent HarnessAgentic Benchmarks