AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems
Organizations: The University of Chicago · Fudan University · TensorBlock, Inc · Tsinghua University · University of Illinois Urbana-Champaign
Abstract
Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number of executable harness bugs while requiring hundreds of human hours to construct. This work presents AgentBug-Smith, an automated harness bug reproduction approach that continuously discovers and reproduces real-world harness bugs from open-source agentic systems. Across different backbone LLMs, AgentBug-Smith consistently outperforms existing bug reproduction techniques designed for general software, achieving 10.67% - 27.56% higher success rates of reproducing harness bugs. By applying AgentBug-Smith to open-source agentic systems in the wild, we construct Live-Harness-Bench, a live and extensible benchmark that currently contains 200 reproducible harness bugs. We further demonstrate the utility of Live-Harness-Bench through two downstream applications. First, we use Live-Harness-Bench as the evaluation benchmark to systematically evaluate state-of-the-art software agents, revealing their limited capabilities in repairing real-world harness bugs. Second, we use Live-Harness-Bench as a knowledge base of real-world harness bug fixes, from which reusable repair skills can be distilled to improve existing software agents, increasing their harness-bug repair rates by 6.32%. Together, AgentBug-Smith and Live-Harness-Bench establish a scalable foundation for continuously evaluating and improving software agents on harness bug repair, turning real-world agent failures into executable evaluation instances and reusable knowledge for harness improvement, thus contributing to the ultimate goal of recursively self-improving agents.
Figures & tables
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Failure mode | DeepSeek-v3.2 | GPT-4.1-mini | Kimi-k2.5 |
| Single cause: environment | |||
| Dependency installation | 44 | 33 | 19 |
| Docker build configuration | 19 | 15 | 40 |
| Missing imports or incompatible APIs | 15 | 25 | 7 |
| Python packaging metadata | 15 | 27 | 4 |
| Native build or compiler | 12 | 28 | 5 |