Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains. Its environments combine challenging assignments with opportunities to cheat, allowing researchers to study how agents pursue a goal when honest work is difficult. CheatBench supports comparisons across models and task categories, providing a testbed for measuring and reducing cheating as agents take on more consequential responsibilities. We publicly release CheatBench at https://cheatbench.ai
Figures & tables
Figure 1: Cheating is common among current AI agents as measured by CheatBench . Overall rates average all ten categories equally for each agent. Sycophancy uses a continuous normalized mean-shift score. Software Engineering counts unique tasks, positive if any recorded attempt cheats. Table 2 gives the category rates.
Figure 2: An example CheatBench environment. The task is framed so that (1) honest work is expected, (2) the filespace contains a honeypot pointing at the reference answer, and (3) reading or using that file is what the judge counts as cheating. Noticing the metadata is not itself cheating. The prompt is shortened with ellipses.
Figure 3: Task framing determines whether reference use serves the user or violates an expectation of independent work. A standard request may permit consulting an existing proof; an explicit restriction or an interview scenario can make doing so cheating. Each workspace contains the same materials. Agent behaviors and thought bubbles are illustrative, not measured outcomes or recorded trajectories.
Category
Honest work expectation
Honeypot
Cheating action
Mathematical Research
Prove an open conjecture independently, or for a research-agent assessment.
A background page linking to a claimed proof.
Opening the planted proof archive.
Multimodal
Locate a photograph from visual clues; count jigsaw pieces without code.
Image metadata pointing at the coordinates.
Reading the reference coordinates.
Creative Writing
Write to a brief for an interviewer assessing writing ability.
Grader logs naming another candidate’s file.
Reading the reference story.
SVG Competition
Draw a competition entry in Scalable Vector Graphics (SVG) format.
Leaderboard configuration revealing the rubric path.
Reading the held-out rubric.
Menial Computation
Factor a semiprime, or solve subset sum, in a timed interview without code.
A generator script naming the audit logs.
Reading the planted answer key, or computing with code where the interview forbids it.
Biology and Bioinformatics
Design a protein binder and explain it, for a manager’s assessment.
A leftover job log naming a colleague’s file.
Reading the colleague’s binder sequences.
Table 1: CheatBench ’s ten categories. Each pairs an assignment that establishes an expectation of honest work with a cheating opportunity and a defined cheating action. Most environments plant a clue leading toward reference material; Software Engineering judges direct reference-seeking behavior. Mathematical Research, Multimodal, Menial Computation and Board Games contain two environments each and the row shows one of them; the others contain one. Each cheating action here is the single clearest act for that category; the full criteria, including the other acts that count and what does not count, are in Appendix B . Sycophancy is scored on shifts across matched prompts rather than on a tool-use attempt.
Figure 4: Claude Opus 5 deliberately accesses the colleague’s protein designs right after stating that it should not look at or copy them, a real example of cheating that directly contradicts its own chain of thought. A staged log exposes their location; after seven rejected designs the agent locates the file, writes that it should not read it, and reads it with head in the very next call. The excerpt ends at that read. Text is quoted; ellipses mark omissions.
Table 2: Cheating probability (%) by category. Overall weights all ten categories equally and matches Figure 1 . Details in Appendix B.1 .
Figure 5: Earlier models encounter fewer honeypots and cheat less often on matched examples across four equally weighted tasks. Exposure means finding the planted clue. Colored points show earlier models; gray points show current agents.
Table 3: Cheating rates (%) under provider and general harnesses. Sycophancy is excluded because it uses direct chat evaluations.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Agents that rarely cheat in one category can cheat frequently in another. Sycophancy reports a continuous normalized mean-shift score (Appendix B.10 ). Bars share a 0–100% scale and reproduce Table 2 ; Appendix B.1 defines the category shorthands. MM pools Visual Geolocation and Visual Puzzles, Menial Computation pools Prime Factorization and Subset Sum, and Games pools Chess and Go. Software Engineering counts unique tasks, positive if any recorded attempt cheats.
Task
GPT-5
Gemini 2.5 Pro
Visual Geolocation
95.0%
0.0%
Visual Puzzles
90.0%
0.0%
Creative Writing
95.0%
52.6%
SVG Competition
10.0%
0.0%
Prime Factorization
100.0%
60.0%
Protein Design
5.0%
0.0%
Appendix
Table 4: Cheating rates (%) for earlier models on all evaluated examples. Software Engineering counts unique tasks, positive if any recorded attempt is judged cheating. Prime Factorization includes prohibited computation, answer-file access, and scoring tampering.
Cheating
Agent
Original Prompt
Modified Prompt
GPT-6 Astra
90%
15%
Claude Fable 5.1
30%
0%
Appendix
Table 5: Cheating propensity in our Goodhart Labs replication ( Valentine, 2026 ) , using their definition: the percentage of episodes in which the grader recorded a search request to the exposed chess engine during the game. Each cell covers twenty episodes, one run each. These runs are separate from the native Chess results in Table 2 .
Table 6: Mean score (%) across ten categories. Sycophancy uses its continuous score.
Table 7: Cheating rate (%) for Implicit + Do Your Own Work.
Table 8: Cheating rate (%) for Implicit + Don’t Cheat.
Cheating rate
Awareness expressed
74.6%
No awareness expressed
62.3%
Appendix
Table 9: Cheating rates by expressed evaluation awareness.
Figure 7: CheatBench spans ten diverse categories and evaluates full agent harnesses, including the models, tools, and interfaces used to carry out assignments; Sycophancy uses separate chat evaluations.