Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains. Its environments combine challenging assignments with opportunities to cheat, allowing researchers to study how agents pursue a goal when honest work is difficult. CheatBench supports comparisons across models and task categories, providing a testbed for measuring and reducing cheating as agents take on more consequential responsibilities. We publicly release CheatBench at https://cheatbench.ai
Figures & tables
Figure 1: Cheating is common among current AI agents as measured by CheatBench . Overall rates average all ten categories equally for each agent. Sycophancy uses a continuous normalized mean-shift score. Software Engineering counts unique tasks, positive if any recorded attempt cheats. Table 2 gives the category rates.
Figure 2: An example CheatBench environment. The task is framed so that (1) honest work is expected, (2) the filespace contains a honeypot pointing at the reference answer, and (3) reading or using that file is what the judge counts as cheating. Noticing the metadata is not itself cheating. The prompt is shortened with ellipses.
Figure 3: Task framing determines whether reference use serves the user or violates an expectation of independent work. A standard request may permit consulting an existing proof; an explicit restriction or an interview scenario can make doing so cheating. Each workspace contains the same materials. Agent behaviors and thought bubbles are illustrative, not measured outcomes or recorded trajectories.
Category
Honest work expectation
Honeypot
Cheating action
Mathematical Research
Prove an open conjecture independently, or for a research-agent assessment.
A background page linking to a claimed proof.
Opening the planted proof archive.
Multimodal
Locate a photograph from visual clues; count jigsaw pieces without code.
Image metadata pointing at the coordinates.
Reading the reference coordinates.
Creative Writing
Write to a brief for an interviewer assessing writing ability.
Grader logs naming another candidate’s file.
Reading the reference story.
SVG Competition
Draw a competition entry in Scalable Vector Graphics (SVG) format.
Leaderboard configuration revealing the rubric path.
Reading the held-out rubric.
Menial Computation
Factor a semiprime, or solve subset sum, in a timed interview without code.
A generator script naming the audit logs.
Reading the planted answer key, or computing with code where the interview forbids it.
Biology and Bioinformatics
Design a protein binder and explain it, for a manager’s assessment.
A leftover job log naming a colleague’s file.
Reading the colleague’s binder sequences.
Table 1: CheatBench ’s ten categories. Each pairs an assignment that establishes an expectation of honest work with a cheating opportunity and a defined cheating action. Most environments plant a clue leading toward reference material; Software Engineering judges direct reference-seeking behavior. Mathematical Research, Multimodal, Menial Computation and Board Games contain two environments each and the row shows one of them; the others contain one. Each cheating action here is the single clearest act for that category; the full criteria, including the other acts that count and what does not count, are in Appendix B . Sycophancy is scored on shifts across matched prompts rather than on a tool-use attempt.
Figure 4: Claude Opus 5 deliberately accesses the colleague’s protein designs right after stating that it should not look at or copy them, a real example of cheating that directly contradicts its own chain of thought. A staged log exposes their location; after seven rejected designs the agent locates the file, writes that it should not read it, and reads it with head in the very next call. The excerpt ends at that read. Text is quoted; ellipses mark omissions.
Table 2: Cheating probability (%) by category. Overall weights all ten categories equally and matches Figure 1 . Details in Appendix B.1 .
Figure 5: Earlier models encounter fewer honeypots and cheat less often on matched examples across four equally weighted tasks. Exposure means finding the planted clue. Colored points show earlier models; gray points show current agents.
Table 3: Cheating rates (%) under provider and general harnesses. Sycophancy is excluded because it uses direct chat evaluations.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Agents that rarely cheat in one category can cheat frequently in another. Sycophancy reports a continuous normalized mean-shift score (Appendix B.10 ). Bars share a 0–100% scale and reproduce Table 2 ; Appendix B.1 defines the category shorthands. MM pools Visual Geolocation and Visual Puzzles, Menial Computation pools Prime Factorization and Subset Sum, and Games pools Chess and Go. Software Engineering counts unique tasks, positive if any recorded attempt cheats.
Task
GPT-5
Gemini 2.5 Pro
Visual Geolocation
95.0%
0.0%
Visual Puzzles
90.0%
0.0%
Creative Writing
95.0%
52.6%
SVG Competition
10.0%
0.0%
Prime Factorization
100.0%
60.0%
Protein Design
5.0%
0.0%
Appendix
Table 4: Cheating rates (%) for earlier models on all evaluated examples. Software Engineering counts unique tasks, positive if any recorded attempt is judged cheating. Prime Factorization includes prohibited computation, answer-file access, and scoring tampering.
Cheating
Agent
Original Prompt
Modified Prompt
GPT-6 Astra
90%
15%
Claude Fable 5.1
30%
0%
Appendix
Table 5: Cheating propensity in our Goodhart Labs replication ( Valentine, 2026 ) , using their definition: the percentage of episodes in which the grader recorded a search request to the exposed chess engine during the game. Each cell covers twenty episodes, one run each. These runs are separate from the native Chess results in Table 2 .
Table 6: Mean score (%) across ten categories. Sycophancy uses its continuous score.
Table 7: Cheating rate (%) for Implicit + Do Your Own Work.
Table 8: Cheating rate (%) for Implicit + Don’t Cheat.
Cheating rate
Awareness expressed
74.6%
No awareness expressed
62.3%
Appendix
Table 9: Cheating rates by expressed evaluation awareness.
Figure 7: CheatBench spans ten diverse categories and evaluates full agent harnesses, including the models, tools, and interfaces used to carry out assignments; Sycophancy uses separate chat evaluations.
LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented reward hacking in these environments, bringing into question the validity of produced research and the broader safety case for AI R&D. Existing benchmarks do not measure exploits that live in the data or the modeling task itself. We introduce BAITBENCH, a suite of three synthetic tabular ML tasks that each contain a shortcut that allows agents to inflate the public test score but fail on a hidden test set. Since the shortcut is optional and using it breaks no stated rule, BAITBENCH measures how often models exploit the shortcut to achieve inflated scores. Across seven frontier agents scored by our two-stage judge pipeline, 57.1% of runs exhibit reward hacking, with five of seven above 50%. Agents cheat even under a second condition where they are prompted not to -the mean cheating rate remains above 50%. We release BAITBENCH, along with the judge implementation, and an annotated dataset of transcripts containing reward hacks as a testbed for evaluating reward-hacking mitigations head-to-head.
Pradyumna Shyama Prasad, Meiri Anto, Leon Eshuijs +3
National University of Singapore · MIT · Vrije Universiteit Amsterdam +3
Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby agents appear successful under the evaluation signal while violating the intended objective. Reward hacking has been observed across a wide range of settings, yet methods for reliably measuring it at scale remain lacking. In this work, we introduce a new evaluation paradigm for measuring reward hacking. Whereas prior studies have primarily analyzed it post hoc by inspecting agent trajectories, we instead embed detectable reward hacking opportunities directly into environments. This makes their exploitation verifiable by design, enabling deterministic and automated measurement of whether and how agents exploit such vulnerabilities. We instantiate this approach in TextArena and release Hack-Verifiable TextArena, a testbed in which reward hacking can be measured reliably. Using this benchmark, we analyze reward hacking behavior across language models in diverse environments and settings. We open source the code at https://github.com/MajoRoth/hack-verifiable-environments/.
Amit Roth, Ankur Samanta, Matan Halevy +2
1Tel Aviv University · 2Columbia University · 3Taso Labs
Agent benchmarks have become the de facto measure of frontier AI competence, guiding model selection, investment, and deployment. However, reward hacking, where agents maximize a score without performing the intended task, emerges spontaneously in frontier models without overfitting. We argue that benchmarks must be secure by design. From past incidents of reward hacks, we derive a taxonomy of eight recurring flaw patterns and compile them into the Agent-Eval Checklist for benchmark designers. We condense the insights into BenchJack, an automated red-teaming system that drives coding agents to audit benchmarks and identify possible reward-hacking exploits in a clairvoyant manner. Moreover, we extend BenchJack to an iterative generative-adversarial pipeline that discovers new flaws and patches them iteratively to improve benchmark robustness. We apply BenchJack to 10 popular agent benchmarks spanning software engineering, web navigation, desktop computing, and terminal operations. BenchJack synthesizes reward-hacking exploits that achieve near-perfect scores on most of the benchmarks without solving a single task, surfacing 219 distinct flaws across the eight classes. Moreover, BenchJack's extended pipeline reduces the hackable-task ratio from near 100% to under 10% on four benchmarks without fatal design flaws, fully patching WebArena and OSWorld within three iterations. Our results show that evaluation pipelines have not internalized an adversarial mindset, and that proactive auditing could help close the security gap for the fast-paced benchmarking space.