Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs and witnesses or require manual construction. We present WitnessGym, an automated framework for constructing bug-validation benchmarks through bug injection. It injects bugs into test-reached paths of real projects, rebuilds each project, and retains cases exposed by a construction-time witness. Bug specifications and execution adapters allow extension to additional bug types and languages. Bug-preserving transformations vary the surrounding structure while preserving the witness behavior. Based on real-world Java projects with test suites, WitnessGym automatically constructs 1,300 benchmark cases. The injected patches resemble historical bug patches and are difficult for the two evaluated models to distinguish in blinded comparisons. We evaluate four coding agent frameworks in six framework/model pairings across bug types, execution contexts, and transformation depths. Witness construction remains difficult even when the bug pattern is known. Our framework, benchmark cases, and evaluation scripts are available.
Figures & tables
Figure 1: Comparison of benchmark-construction methodologies for executable bug validation.
Work
Diversity
Auto. cases
Fresh bugs
Real projects
Context control
CTF and security challenges
NYU CTF Bench ( 2024 )
High
✗
✗
✗
✗
Cybench ( 2025b )
High
✗
✗
✗
✗
Historical bug validation
LIBRO ( 2023 )
High
✗
✗
✓
✗
Issue2Test ( 2026 )
High
✗
✗
✓
✗
Table 1: Comparison of bug-validation benchmark construction. Diversity is high for multiple bug families, medium for variants within a dominant family, and low for one narrow family. Auto. cases are constructed automatically; fresh bugs are generated rather than reused; real projects retain production code with its build and test environments; and context control varies surrounding call, data, or control structure.
Figure 2: Overview of the WitnessGym construction and evaluation pipeline.
Agent
Framework
Model
A1
Codex
GPT-5.4
A2
Claude Code
Claude Sonnet 4.5
A3
OpenHands
DeepSeek-V3.2
A4
OpenCode
ZAI GLM-4.7
A5
OpenHands
Devstral-2-123B
A6
OpenCode
Qwen3-Coder-480B-A35B-Instruct
Table 2: Evaluated framework/model pairings.
ID
NoEC
WithEC
Average
A1
67.1%
70.7%
68.9%
A2
57.1%
61.8%
59.4%
A3
46.0%
49.8%
47.9%
A4
44.2%
47.6%
45.9%
A5
30.2%
33.8%
32.0%
A6
17.1%
19.3%
18.2%
Table 3: Agent validation success rates.
Bug Type
A1
A2
A3
A4
A5
A6
L
S
L
S
L
S
L
S
L
S
L
S
Comparison of Identical Values
18/17
22/21
17/14
19/16
14/10
15/15
11/11
15/13
9/7
10/9
4/4
6/5
Typo in equals
18/15
21/21
14/13
16/18
12/13
15/13
10/10
12/11
8/7
9/10
5/3
6/5
Typo in hashCode
16/13
18/17
16/11
19/14
10/9
15/12
11/8
13/14
7/6
10/7
4/4
5/5
Hashed Value Without hashCode
14/14
19/17
15/13
16/15
10/11
15/13
9/9
14/10
8/6
10/7
4/3
5/4
Inconsistent equals and hashCode
17/15
18/20
13/12
18/17
12/9
12/13
12/9
13/11
6/6
10/9
3/4
5/4
Table 4: Successful bug validations across bug types. L and S denote long and short execution contexts. Each A/B cell reports successful cases under WithEC and NoEC , respectively.
Figure 3: Success-rate gain with execution context (percentage points). Types follow Table 4 .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Type
Bug-type description
API Contract
Comparison of Identical Values
A condition compares a value with itself, producing a constant result that silently disables validation or equality logic.
Typo in equals
A function overriding equals(Object) has the wrong name or signature, so Java keeps using reference equality.
Typo in hashCode
A function intended to override hashCode() has the wrong name or signature, so hashed collections use identity hashing.
Hashed Value Without hashCode
A class overrides equals(Object) but not hashCode() , so logically equal keys may hash to different buckets.
Inconsistent equals and hashCode
equals and hashCode use different fields or unstable state, breaking lookup in HashMap or HashSet .
Value Flow
Intra-Function Null Dereference
A value that was checked for null is later overwritten, moved outside the guard, or dereferenced on an unguarded path.
Appendix
Table 5: The ten bug types used in the benchmark construction.
Family
Operator
Effect
Call Stack
Helper Call Insertion
Adds an internal helper or wrapper call between the original caller and the bug site, increasing call-chain depth.
Call-Site Rerouting
Replaces a direct internal call with a forwarding function, so the same logic is reached through an extra dispatch step.
Extract-and-Apply Split
Splits one computation into an extraction helper and an application helper, forcing the bug to cross function boundaries.
Data Flow
Parameter Bundling
Packs related arguments or locals into an internal object that is passed through helper calls instead of using separate values.
Local State Bundling
Moves related local variables into a context object, so later code reads state through fields rather than direct locals.
Intermediate State Passing
Stores intermediate values in a state holder and reads them later across helper calls, lengthening the value flow path.
Appendix
Table 6: The ten bug-preserving transformation operators.
Agent ID
Short
Mid-Short
Mid-Long
Long
NoEC
WithEC
NoEC
WithEC
NoEC
WithEC
NoEC
WithEC
A1
80%
80%
72% ↓ 8pp
76% ↓ 4pp
64% ↓ 16pp
64% ↓ 16pp
48% ↓ 32pp
52% ↓ 28pp
A2
72%
76%
64% ↓ 8pp
68% ↓ 8pp
56% ↓ 16pp
56% ↓ 20pp
40% ↓ 32pp
40% ↓ 36pp
A3
60%
64%
52% ↓ 8pp
56% ↓ 8pp
40% ↓ 20pp
44% ↓ 20pp
28% ↓ 32pp
28% ↓ 36pp
A4
60%
60%
52% ↓ 8pp
56% ↓ 4pp
40% ↓ 20pp
40% ↓ 20pp
28% ↓ 32pp
28% ↓ 32pp
A5
44%
44%
36% ↓ 8pp
40% ↓ 4pp
28% ↓ 16pp
28% ↓ 16pp
16% ↓ 28pp
20% ↓ 24pp
Appendix
Table 7: Bug-validation success rates across execution context length buckets. Subscript annotations give percentage-point differences from the short bucket.
Agent ID
Depth 1
Depth 3
Depth 5
Depth 10
NoEC
WithEC
NoEC
WithEC
NoEC
WithEC
NoEC
WithEC
A1
76%
80%
67% ↓ 9pp
68% ↓ 12pp
62% ↓ 14pp
65% ↓ 15pp
59% ↓ 17pp
59% ↓ 21pp
A2
70%
72%
59% ↓ 11pp
60% ↓ 12pp
53% ↓ 17pp
57% ↓ 15pp
50% ↓ 20pp
51% ↓ 21pp
A3
57%
60%
45% ↓ 12pp
49% ↓ 11pp
42% ↓ 15pp
43% ↓ 17pp
36% ↓ 21pp
40% ↓ 20pp
A4
53%
55%
47% ↓ 6pp
48% ↓ 7pp
41% ↓ 12pp
41% ↓ 14pp
39% ↓ 14pp
40% ↓ 15pp
A5
38%
40%
33% ↓ 5pp
32% ↓ 8pp
28% ↓ 10pp
32% ↓ 8pp
25% ↓ 13pp
28% ↓ 12pp
Appendix
Table 8: Bug-validation success rates across transformation-depth buckets. Subscript annotations give percentage-point differences from depth 1.
Depth
Runs
Final lines
Hunks
Changed files
Line accum.
First-step retention
First-step similarity
1
100
65.0
2.33
1.22
1.00 ×
100.0%
100.0%
3
100
97.8
3.59
1.40
2.40 ×
88.1%
61.7%
5
100
161.5
4.34
1.48
4.25 ×
76.7%
40.8%
10
100
243.7
6.09
1.44
6.63 ×
74.9%
25.9%
Appendix
Table 9: Injection-stage patch-evolution metrics by transformation depth. Line accumulation is the ratio between cumulative step-patch lines and final-patch lines. Retention and sequence similarity are measured from the first transformation step to the final buggy patch.
Task
Judge
Correct / Total
Accuracy
90% CI
AUC
90% CI
Single
A
2,173 / 4,200
51.74%
50.14–53.38
0.51198
0.49367–0.53075
Single
B
2,189 / 4,200
52.12%
50.52–53.69
0.51286
0.49426–0.53084
Pairwise
A
723 / 1,400
51.64%
49.07–54.21
0.51681
0.48703–0.54633
Pairwise
B
730 / 1,400
52.14%
49.64–54.64
0.52411
0.49517–0.55252
Appendix
Table 10: Detailed source-discrimination results. Accuracy CIs are percentages.
Judge
Correct / Total
Incorrect / Total
Accuracy
95% CI
Check
A
184 / 200
16 / 200
92.00%
87.00%–96.00%
Pass
B
169 / 200
31 / 200
84.50%
78.50%–89.50%
Pass
Appendix
Table 11: Artifact-bearing positive controls. Both judges meet the prespecified sensitivity check.
Diagnostic
Judge A
Judge B
Single-patch repeat agreement
60.62%
58.67%
Patches with non-unanimous single judgments
59.07%
62.00%
Main orientation-choice agreement
68.43%
64.29%
Main LEFT choice rate
52.93%
54.71%
Main LEFT choices / judgments
741 / 1,400
766 / 1,400
Main split pairs choosing LEFT twice
131
158
Appendix
Table 12: Judgment stability, position preferences, and correctness counts (1,400 patches per judge).
Prompt Structure
Main Inputs
Required Output
Bug Injection System Prompt
Global injection rules, replay-command discipline, bounded edit–run–revise loop, diff checks, and success criteria
Single-line INJECT_REPORT_JSON with operational evidence fields
Bug Injection Task Payload
Bug type, bug type specification, required and forbidden elements, execution context, candidate anchors, trigger hints, and file scope
Same INJECT_REPORT_JSON prefix with type-site rationale fields