cs.AISep 28, 2026

CheatBench: Measuring Reward Gaming in AI Agents

Authors: Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang, Zheyuan Liu, Richard Ren, Jingxiang Meng, +5 more

Organizations: Center for AI Safety · Work done while at the Center for AI Safety.

Abstract

Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains. Its environments combine challenging assignments with opportunities to cheat, allowing researchers to study how agents pursue a goal when honest work is difficult. CheatBench supports comparisons across models and task categories, providing a testbed for measuring and reducing cheating as agents take on more consequential responsibilities. We publicly release CheatBench at https://cheatbench.ai

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

    Aug 31, 2026Pradyumna Shyama Prasad, Meiri Anto, Leon Eshuijs +3Agentic BenchmarksExploitation

  2. Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale

    May 20, 2026Amit Roth, Ankur Samanta, Matan Halevy +2Adversarial EvaluationExploitation