cs.SESep 29, 2026

WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses

Authors: Haomin Qi, Xiangzhe Xu, Yiming Huang, Jingbo Shang, Chengpeng Wang

Organizations: University of California, San Diego · Purdue University · National University of Singapore

Abstract

Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs and witnesses or require manual construction. We present WitnessGym, an automated framework for constructing bug-validation benchmarks through bug injection. It injects bugs into test-reached paths of real projects, rebuilds each project, and retains cases exposed by a construction-time witness. Bug specifications and execution adapters allow extension to additional bug types and languages. Bug-preserving transformations vary the surrounding structure while preserving the witness behavior. Based on real-world Java projects with test suites, WitnessGym automatically constructs 1,300 benchmark cases. The injected patches resemble historical bug patches and are difficult for the two evaluated models to distinguish in blinded comparisons. We evaluate four coding agent frameworks in six framework/model pairings across bug types, execution contexts, and transformation depths. Witness construction remains difficult even when the bug pattern is known. Our framework, benchmark cases, and evaluation scripts are available.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch

    Aug 3, 2026Shuyang Xie, Shuxiao Xie, Feng Zhu +2Raw Judge OutputsTest Generation

  2. AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection

    Apr 13, 2026Zijie Zhao, Chenyuan Yang, Weidong Wang +3Bug DetectionBug

  3. SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

    Date pendingBingchen Zhao, Dhruv Srikanth, Yuxiang Wu +1Coding AgentsAgentic Benchmarks