Agentic benchmarks guide model selection and training. Yet an agent can pass a task without demonstrating the intended capability. Such outcomes constitute unearned passes; their proportion among all passes defines the integrity gap. As agents improve, benchmark surfaces that once seemed harmless can become exploitable, making benchmark validity an ongoing maintenance problem. We introduce a process-verification framework that audits passing trajectories, distinguishes evidenced reward hacking from verifier weakness, and localizes exploitable surfaces for repair. Across 3,810 passing trajectories from 29 model-benchmark cohorts, confirmed violations often increase with model generation but not monotonically. On SWEBench Pro V1.0, confirmed violation rates rise from 24% to 73% between Opus 4.7 and Fable 5 on matched tasks; later cohorts fall to 11% for Fable 5.1 and 0% for GPT-6 Astra. These comparisons are descriptive: configurations were not normalized, and the latest models also pass fewer exploitable tasks. Violations concentrate around a small set of recurring surfaces, especially unintended access to reference solutions through git history. Three repair case studies across two benchmarks show why blocking a recorded exploit is insufficient: the same protected information can remain accessible through another route. Therefore, we combine minimal patches with exploit replay and fresh agent evaluation, auditing new passes under the original standard. No evaluated attempt against the final patches reached the protected channel, and every post-patch pass was judged legitimate. Benchmark integrity requires ongoing maintenance: audit passing behavior, repair the enabling surface, and re-evaluate both exploit access and legitimate solvability.
Figures & tables
Figure 1 : Detection, localization, attribution, and remediation workflow. Passing trajectories are classified as legitimate , unearned , and unearned violations . A violation requires evidence of deliberate, grade-relevant exploitation; an unearned pass alone does not. Confirmed violations are mapped to the surface that enabled it. The remediation separates the identified channel from the addresses used to reach it, applies the smallest deterministic edit at the effective control surface, and verifies closure by rebuilding the task and replaying its recorded exploit while varying only that control. The patched task is then rerun live at pass@ k and re-judged with the same pipeline, returning to detection as ordinary input (dashed). The loop is closed by the dashed return: a patched task is accepted only after being judged again under the standard that defined the original violation. Gray: deterministic; Green: large language-model (LLM) judges; blue: remediation.
Figure 2 : Violation rates across model generations. (A) Confirmed violations among selected passing tasks (n); bars show 95% Wilson intervals. SWEBench Pro V1.0 peaks at 73% and 68%, then falls to 11% and 0% in the latest cohorts. (B) On tasks both models passed, violations are grouped as earlier-only, later-only, or both. A task counts if any passing attempt violates. Cohort configurations and passing-task sets differ, so these are descriptive comparisons.
Figure 3 : An unearned violations does not always establish reward hacking. All four examples received benchmark credit without meeting the intended requirements. Top: confirmed violations copy an upstream test file (SWE Atlas TW) or tune an output to a leaked answer prefix (Terminal-Bench 2.1). Bottom: other unearned passes claim facts absent from retrieved content (MCP Atlas) or place tests in the wrong file (SWE Atlas TW). The latter two reveal verifier weaknesses without sufficient evidence of deliberate, grade-relevant exploitation
Figure 4 : Repair must change the environment the agent actually uses. In this SWEBench Pro V1.0 task, the agent copied an upstream fix from git history. Top: human and automated patches remove that access in the executed environment. Bottom: fresh runs produce legitimate passes after either repair, and replay no longer retrieves the fix.
Figure 5 : Blocking a shortcut makes robot reconstruction harder but still solvable. In this Agents’ Last Exam (ALE) task, three models retrieved an upstream robot description. Top: the patch blocks external network access. Bottom: with three fresh attempts per model, Fable 5 produces one legitimate pass; all other attempts fail. No replayed or fresh attempt retrieves the reference content.
Figure S1 : Distribution of confirmed violation-category assignments across the cohorts plotted in Figure 2 . A trajectory can contribute more than one category assignment, so these counts need not equal the task-level violation counts in Figure 2 .
Figure S2 : Neutral-verdict parity of Figure 2 : the same cohorts, denominators and intersections, counting neutral verdicts instead of confirmed violations. A neutral marks a surface pass the judge found process-illegitimate without enough evidence of being escalated as violation / specific gaming, so these rates measure the residual integrity gap that the violation rate excludes. (A) Neutral-verdict rate over each provider-shared surface-pass set. (B) Neutral overlap across model generations on the same shared sets. Across all 29 cohorts, 91 tasks carry a neutral verdict.
Figure S3 : Neutral-verdict parity of Figure S1 : distribution of neutral-category assignments across the cohorts plotted in Figure S2 , using the neutral vocabulary of the judge ( _VALID_NEUTRAL_TYPES ). The 84 assignments shown are the categorized subset of the 91 neutral tasks in Figure S2 .
Figure S4 : A second SWEBench Pro V1.0 case study, ansible-5e88cd99, comparing the original task, a human-expert patch, and the git_oracle_seal autopatch. The original package leaves the gold commit reachable in the shipped container and permits egress; both patched versions remove these routes by using or building a pre-pruned image that rejects any surviving gold commit.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Model
Att.
Tasks
SP tasks
SP rate
Viol. full
Neut. full
n
Viol. (Figure 2)
Neut. (Figure S2)
Viol. rate
Neut. rate
Gap
SWEBench Pro V1.0
Opus 4.7
3
70
66
94.3%
22
0
49
12
0
24.49%
0.00%
24.49%
SWEBench Pro V1.0
Fable 5
1
70
52
74.3%
39
0
49
36
0
73.47%
0.00%
73.47%
SWEBench Pro V1.0
Fable 5.1
1
70
36
51.4%
4
0
36
4
0
11.11%
0.00%
11.11%
SWEBench Pro V1.0
GPT-5.1 Codex
1
104
104
100.0%
2
29
104
2
29
1.92%
27.88%
29.81%
SWEBench Pro V1.0
GPT-5.6 Sol xhigh
1
104
104
100.0%
71
2
104
71
2
68.27%
1.92%
70.19%
SWEBench Pro V1.0
GPT-6 Astra
1
104
96
92.3%
0
4
96
0
4
0.00%
4.17%
4.17%
Appendix
Table S1 : Per-cohort evaluated tasks, surface-pass rates, violations and neutral verdicts across all 29 cohorts. Columns: Att. attempts per task; Tasks tasks evaluated; SP tasks tasks surface-passed; SP rate surface-pass rate; Viol. full and Neut. full violating and neutral tasks on the cohort’s own task set; n shared-set tasks; Viol. of Figure 2 and Neut.of Figure S2 violating and neutral tasks on that shared set; Viol. rate , Neut. rate and Gap the corresponding rates and the integrity gap. SWEBench Pro V1.0: tasks were selected using stratified sampling from the 731 public tasks; MCP Atlas: full 500 public tasks; SWE Atlas TW: 90/90 public tasks; Terminal-Bench 2.1: 89/89 public tasks; ALE: 99/167 Ubuntu task subset of 168 public tasks. Note that task-level violation status is positive when any passing attempt is a violation , comparisons with unequal attempts may be affected by multiplicity. Here we report attempts per task to make this difference explicit.
Benchmark
Model
k
n
Rate
95% Wilson
Width
SWEBench Pro V1.0
Opus 4.7
12
49
24.49%
[14.60, 38.09]
23.49
SWEBench Pro V1.0
Fable 5
36
49
73.47%
[59.74, 83.79]
24.05
SWEBench Pro V1.0
Fable 5.1
4
36
11.11%
[4.41, 25.32]
20.91
SWEBench Pro V1.0
GPT-5.1 Codex
2
104
1.92%
[0.53, 6.74]
6.21
SWEBench Pro V1.0
GPT-5.6 Sol xhigh
71
104
68.27%
[58.81, 76.43]
17.62
SWEBench Pro V1.0
GPT-6 Astra
0
96
0.00%
[0.00, 3.85]
3.85
Appendix
Table S2 : Confidence intervals plotted in panel A of Figure 2 (violations). k is the number of violating tasks and n the provider-shared surface-pass set; the rate is k/n and the interval is the 95% Wilson score interval, which is what the whiskers in the figure show. Width is in percentage points. These are the stored values the renderer reads, not a re-derivation.
Benchmark
Model
k
n
Rate
95% Wilson
Width
SWEBench Pro V1.0
Opus 4.7
0
49
0.00%
[0.00, 7.27]
7.27
SWEBench Pro V1.0
Fable 5
0
49
0.00%
[0.00, 7.27]
7.27
SWEBench Pro V1.0
Fable 5.1
0
36
0.00%
[0.00, 9.64]
9.64
SWEBench Pro V1.0
GPT-5.1 Codex
29
104
27.88%
[20.17, 37.17]
17.00
SWEBench Pro V1.0
GPT-5.6 Sol xhigh
2
104
1.92%
[0.53, 6.74]
6.21
SWEBench Pro V1.0
GPT-6 Astra
4
96
4.17%
[1.63, 10.23]
8.60
Appendix
Table S3 : Confidence intervals plotted in panel A of Figure S2 (neutral verdicts). Same cohorts, same denominators n as Table S2 ; here k counts tasks whose effective verdict is neutral . Intervals are 95% Wilson score intervals. Rows where k=0 still carry a non-zero upper bound, set by n alone.
Agent benchmarks have become the de facto measure of frontier AI competence, guiding model selection, investment, and deployment. However, reward hacking, where agents maximize a score without performing the intended task, emerges spontaneously in frontier models without overfitting. We argue that benchmarks must be secure by design. From past incidents of reward hacks, we derive a taxonomy of eight recurring flaw patterns and compile them into the Agent-Eval Checklist for benchmark designers. We condense the insights into BenchJack, an automated red-teaming system that drives coding agents to audit benchmarks and identify possible reward-hacking exploits in a clairvoyant manner. Moreover, we extend BenchJack to an iterative generative-adversarial pipeline that discovers new flaws and patches them iteratively to improve benchmark robustness. We apply BenchJack to 10 popular agent benchmarks spanning software engineering, web navigation, desktop computing, and terminal operations. BenchJack synthesizes reward-hacking exploits that achieve near-perfect scores on most of the benchmarks without solving a single task, surfacing 219 distinct flaws across the eight classes. Moreover, BenchJack's extended pipeline reduces the hackable-task ratio from near 100% to under 10% on four benchmarks without fatal design flaws, fully patching WebArena and OSWorld within three iterations. Our results show that evaluation pipelines have not internalized an adversarial mindset, and that proactive auditing could help close the security gap for the fast-paced benchmarking space.
Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent reward-hacking benchmarks and system reports show that agents can instead recover public solutions, read evaluation artifacts, infer generator structure, manipulate feedback, or benefit from invalid scoring paths; existing responses do not provide a common procedure for attributing these shortcuts and quantifying their effect across benchmarks. We formulate protocol validity and introduce HackDetect, a post-hoc audit that identifies an exposure, determines how the agent used it, and assesses whether the resulting score is misleading. We quantify score inflation with the Mislead gap, defined as the exploit score minus the intended score. We audit 2,385 traces across 15 agent benchmarks and find evidence of exposures and reward hacking in 67.0% of Frontier Science traces and 66.7% of AutoLab tasks. Across paired comparisons, we measure score inflation of 0.45-1.00, showing that benchmark reports should provide evidence that scores reflect the intended capability.
Jiaqi Shao, Hanck Chen, Wei Zhang +2
1Hunyuan Team, Tencent · 2The Hong Kong University of Science and Technology · 3Duke Kunshan University
Agent benchmarks score submissions with outcome verifiers that are typically hand-written and brittle, leaving them open to reward hacking. We audit 1,968 tasks across five terminal-agent benchmarks and find 323 (16%) hackable by frontier models given only the task description. This corrupts both leaderboard rankings and RL training signal, yet the standard response is manual and reactive. We introduce the hacker-fixer loop, a method for building exploit-resistant verifiers without per-task manual patching. The loop alternates three LLM agents: a hacker tries to pass the verifier without solving the task, a fixer patches the verifier to reject each discovered exploit, and a solver confirms the patched verifier still admits legitimate solutions. The loop iterates: each patch reshapes what the verifier rewards, surfacing the next exploit. We further add verifier access, and let patches transfer across tasks, to broaden the exploits the loop discovers. On KernelBench, the loop drives the attack success rate from 62% to 0% on a held-out corpus of publicly reported exploits. We also find that weaker agents in the loop can defend against much stronger hackers: Gemini 3 Flash's loop drives the stronger Gemini 3.1 Pro and Claude Opus 4.7's attack success rate from 76% and 61% to 0% on KernelBench, and Gemini 3.1 Pro's from 39% to 17% on Terminal Bench across 77 tasks. We release Terminal Wrench (323 hackable environments, 3,632 hack trajectories) as a snapshot of the current attack surface, our patched verifiers, the exploits the loop discovered, and our implementation as a basis for future work.
Ziqian Zhong, Ivgeni Segal, Ivan Bercovich +3
Carnegie Mellon University · Fewshot Corp · Independent Researcher