cs.LGOct 6, 2026

DecepEval: A Benchmark for Evaluating Deception in LLM Agents

Authors: Yiming Xu, Hongyue Yu, Beihua Yang, Zihan Chen, Yixin Liu, Zhen Peng, Bin Shi, Bo Dong, +3 more

Organizations: Xi’an Jiaotong University · University of Virginia · Griffith University · The Chinese University of Hong Kong · Tongji University

Abstract

As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but often examine isolated scenarios or narrowly defined conditions, limiting systematic understanding of when deception becomes more likely. To address this gap, we introduce DecepEval, a benchmark comprising 1,532 instances across 3 task families and 28 professional scenarios. Drawing on classical fraud theories, we propose the LLM Deception Diamond framework, which characterizes four external conditions that may induce deception: pressure, incentive, opportunity, and conflict. DecepEval pairs neutral and induced versions of each instance to measure condition-dependent changes in deception rates, while explicit task facts and observable agent behavior help distinguish deception from capability-related errors. Evaluations of nine frontier LLMs show that inducements increase deception across models and task families, even among models with low baseline deception rates. DecepEval makes these vulnerabilities measurable, providing a shared benchmark for progress toward trustworthy artificial intelligence.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

    Jul 30, 2026Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa +3DeceptionSocial Reasoning

  2. Honeyquest for LLMs: Rethinking Cyber Deception for AI Attackers

    Jun 19, 2026Kerri Prinos, Lilianne Brush, Cameron DentonDeceptionAttacker Large Language Model

  3. SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence

    Jun 1, 2026Yuyan Bu, Haowei Li, Qirui Zheng +7DeceptionAutonomous Agents