WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses
Organizations: University of California, San Diego · Purdue University · National University of Singapore
Abstract
Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs and witnesses or require manual construction. We present WitnessGym, an automated framework for constructing bug-validation benchmarks through bug injection. It injects bugs into test-reached paths of real projects, rebuilds each project, and retains cases exposed by a construction-time witness. Bug specifications and execution adapters allow extension to additional bug types and languages. Bug-preserving transformations vary the surrounding structure while preserving the witness behavior. Based on real-world Java projects with test suites, WitnessGym automatically constructs 1,300 benchmark cases. The injected patches resemble historical bug patches and are difficult for the two evaluated models to distinguish in blinded comparisons. We evaluate four coding agent frameworks in six framework/model pairings across bug types, execution contexts, and transformation depths. Witness construction remains difficult even when the bug pattern is known. Our framework, benchmark cases, and evaluation scripts are available.
Figures & tables
| Work | Diversity | Auto. cases | Fresh bugs | Real projects | Context control |
|---|---|---|---|---|---|
| CTF and security challenges | |||||
| NYU CTF Bench ( 2024 ) | High | ✗ | ✗ | ✗ | ✗ |
| Cybench ( 2025b ) | High | ✗ | ✗ | ✗ | ✗ |
| Historical bug validation | |||||
| LIBRO ( 2023 ) | High | ✗ | ✗ | ✓ | ✗ |
| Issue2Test ( 2026 ) | High | ✗ | ✗ | ✓ | ✗ |
| Agent | Framework | Model |
|---|---|---|
| A1 | Codex | GPT-5.4 |
| A2 | Claude Code | Claude Sonnet 4.5 |
| A3 | OpenHands | DeepSeek-V3.2 |
| A4 | OpenCode | ZAI GLM-4.7 |
| A5 | OpenHands | Devstral-2-123B |
| A6 | OpenCode | Qwen3-Coder-480B-A35B-Instruct |
| ID | Average | ||
|---|---|---|---|
| A1 | 67.1% | 70.7% | 68.9% |
| A2 | 57.1% | 61.8% | 59.4% |
| A3 | 46.0% | 49.8% | 47.9% |
| A4 | 44.2% | 47.6% | 45.9% |
| A5 | 30.2% | 33.8% | 32.0% |
| A6 | 17.1% | 19.3% | 18.2% |
| Bug Type | A1 | A2 | A3 | A4 | A5 | A6 | ||||||
| L | S | L | S | L | S | L | S | L | S | L | S | |
| Comparison of Identical Values | 18/17 | 22/21 | 17/14 | 19/16 | 14/10 | 15/15 | 11/11 | 15/13 | 9/7 | 10/9 | 4/4 | 6/5 |
| Typo in equals | 18/15 | 21/21 | 14/13 | 16/18 | 12/13 | 15/13 | 10/10 | 12/11 | 8/7 | 9/10 | 5/3 | 6/5 |
| Typo in hashCode | 16/13 | 18/17 | 16/11 | 19/14 | 10/9 | 15/12 | 11/8 | 13/14 | 7/6 | 10/7 | 4/4 | 5/5 |
| Hashed Value Without hashCode | 14/14 | 19/17 | 15/13 | 16/15 | 10/11 | 15/13 | 9/9 | 14/10 | 8/6 | 10/7 | 4/3 | 5/4 |
| Inconsistent equals and hashCode | 17/15 | 18/20 | 13/12 | 18/17 | 12/9 | 12/13 | 12/9 | 13/11 | 6/6 | 10/9 | 3/4 | 5/4 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Type | Bug-type description |
| API Contract | Comparison of Identical Values | A condition compares a value with itself, producing a constant result that silently disables validation or equality logic. |
| Typo in equals | A function overriding equals(Object) has the wrong name or signature, so Java keeps using reference equality. | |
| Typo in hashCode | A function intended to override hashCode() has the wrong name or signature, so hashed collections use identity hashing. | |
| Hashed Value Without hashCode | A class overrides equals(Object) but not hashCode() , so logically equal keys may hash to different buckets. | |
| Inconsistent equals and hashCode | equals and hashCode use different fields or unstable state, breaking lookup in HashMap or HashSet . | |
| Value Flow | Intra-Function Null Dereference | A value that was checked for null is later overwritten, moved outside the guard, or dereferenced on an unguarded path. |
| Family | Operator | Effect |
| Call Stack | Helper Call Insertion | Adds an internal helper or wrapper call between the original caller and the bug site, increasing call-chain depth. |
| Call-Site Rerouting | Replaces a direct internal call with a forwarding function, so the same logic is reached through an extra dispatch step. | |
| Extract-and-Apply Split | Splits one computation into an extraction helper and an application helper, forcing the bug to cross function boundaries. | |
| Data Flow | Parameter Bundling | Packs related arguments or locals into an internal object that is passed through helper calls instead of using separate values. |
| Local State Bundling | Moves related local variables into a context object, so later code reads state through fields rather than direct locals. | |
| Intermediate State Passing | Stores intermediate values in a state holder and reads them later across helper calls, lengthening the value flow path. |
| Agent ID | Short | Mid-Short | Mid-Long | Long | ||||
|---|---|---|---|---|---|---|---|---|
| NoEC | WithEC | NoEC | WithEC | NoEC | WithEC | NoEC | WithEC | |
| A1 | 80% | 80% | 72% 8pp | 76% 4pp | 64% 16pp | 64% 16pp | 48% 32pp | 52% 28pp |
| A2 | 72% | 76% | 64% 8pp | 68% 8pp | 56% 16pp | 56% 20pp | 40% 32pp | 40% 36pp |
| A3 | 60% | 64% | 52% 8pp | 56% 8pp | 40% 20pp | 44% 20pp | 28% 32pp | 28% 36pp |
| A4 | 60% | 60% | 52% 8pp | 56% 4pp | 40% 20pp | 40% 20pp | 28% 32pp | 28% 32pp |
| A5 | 44% | 44% | 36% 8pp | 40% 4pp | 28% 16pp | 28% 16pp | 16% 28pp | 20% 24pp |
| Agent ID | Depth 1 | Depth 3 | Depth 5 | Depth 10 | ||||
|---|---|---|---|---|---|---|---|---|
| NoEC | WithEC | NoEC | WithEC | NoEC | WithEC | NoEC | WithEC | |
| A1 | 76% | 80% | 67% 9pp | 68% 12pp | 62% 14pp | 65% 15pp | 59% 17pp | 59% 21pp |
| A2 | 70% | 72% | 59% 11pp | 60% 12pp | 53% 17pp | 57% 15pp | 50% 20pp | 51% 21pp |
| A3 | 57% | 60% | 45% 12pp | 49% 11pp | 42% 15pp | 43% 17pp | 36% 21pp | 40% 20pp |
| A4 | 53% | 55% | 47% 6pp | 48% 7pp | 41% 12pp | 41% 14pp | 39% 14pp | 40% 15pp |
| A5 | 38% | 40% | 33% 5pp | 32% 8pp | 28% 10pp | 32% 8pp | 25% 13pp | 28% 12pp |
| Depth | Runs | Final lines | Hunks | Changed files | Line accum. | First-step retention | First-step similarity |
|---|---|---|---|---|---|---|---|
| 1 | 100 | 65.0 | 2.33 | 1.22 | 1.00 | 100.0% | 100.0% |
| 3 | 100 | 97.8 | 3.59 | 1.40 | 2.40 | 88.1% | 61.7% |
| 5 | 100 | 161.5 | 4.34 | 1.48 | 4.25 | 76.7% | 40.8% |
| 10 | 100 | 243.7 | 6.09 | 1.44 | 6.63 | 74.9% | 25.9% |
| Task | Judge | Correct / Total | Accuracy | 90% CI | AUC | 90% CI |
|---|---|---|---|---|---|---|
| Single | A | 2,173 / 4,200 | 51.74% | 50.14–53.38 | 0.51198 | 0.49367–0.53075 |
| Single | B | 2,189 / 4,200 | 52.12% | 50.52–53.69 | 0.51286 | 0.49426–0.53084 |
| Pairwise | A | 723 / 1,400 | 51.64% | 49.07–54.21 | 0.51681 | 0.48703–0.54633 |
| Pairwise | B | 730 / 1,400 | 52.14% | 49.64–54.64 | 0.52411 | 0.49517–0.55252 |
| Judge | Correct / Total | Incorrect / Total | Accuracy | 95% CI | Check |
|---|---|---|---|---|---|
| A | 184 / 200 | 16 / 200 | 92.00% | 87.00%–96.00% | Pass |
| B | 169 / 200 | 31 / 200 | 84.50% | 78.50%–89.50% | Pass |
| Diagnostic | Judge A | Judge B |
| Single-patch repeat agreement | 60.62% | 58.67% |
| Patches with non-unanimous single judgments | 59.07% | 62.00% |
| Main orientation-choice agreement | 68.43% | 64.29% |
| Main LEFT choice rate | 52.93% | 54.71% |
| Main LEFT choices / judgments | 741 / 1,400 | 766 / 1,400 |
| Main split pairs choosing LEFT twice | 131 | 158 |
| Prompt Structure | Main Inputs | Required Output |
|---|---|---|
| Bug Injection System Prompt | Global injection rules, replay-command discipline, bounded edit–run–revise loop, diff checks, and success criteria | Single-line INJECT_REPORT_JSON with operational evidence fields |
| Bug Injection Task Payload | Bug type, bug type specification, required and forbidden elements, execution context, candidate anchors, trigger hints, and file scope | Same INJECT_REPORT_JSON prefix with type-site rationale fields |
| Bug-Preserving Transformation Task Payload | Transformation operator specification, applicability requirements, preservation constraints, execution context, and candidate anchors | Single-line TRANSFORM_REPORT_JSON |
| Bug Validation System Prompt | Test-only edit constraints, Java test-case materialization rules, executable-witness objective, and evaluation-attempt policy | Single-line TRIGGER_REPORT_JSON fallback schema |
| Bug Validation Task Payload | Buggy repository metadata, case ID, bug type, applied transformations, bug-related files, replay command, and optional execution context | Single-line TRIGGER_REPORT_JSON with detailed witness-construction fields |