cs.CLMay 12, 2026

Confidently Deceptive: On the Relationship Between Confidence and Deception in LLMs

Authors: Ali Asad, Stephen Obadinma, Anshul Pattoo, Wenxuan Zhang, Xiaodan Zhu

Organizations: Department of Electrical and Computer Engineering & Ingenuity Labs Research Institute Queen’s University, Kingston, Canada

Abstract

The increasing capabilities of large language models (LLMs) are being accompanied by deep-rooted risks of deceptive behaviours that cause models to produce misleading outputs in service of a contextually or experimentally induced goal. The harm posed by such behaviours depends not only on the content of deceptive outputs but also how confidently models deliver them, since confidence has a major impact on how persuasive the communication is to end users. In this paper, we provide a comprehensive study on the crucial relationship between confidence and deception across existing deception benchmarks and different model families, while covering both verbalized numerical and logit-based aggregated confidence. Through this, we reveal how confidently models behave when being deceptive. We demonstrate that when producing deceptive rather than honest responses, models exhibit a gap between their belief (how likely they think a claim is to be true) and their commitment (how firmly they assert and would defend that claim). LLMs produce persuasive deceptive claims while reporting low belief in their factual correctness. Their reported commitment to deceptive responses can easily be increased through further prompting and preference fine-tuning, with smaller and condition-dependent changes in reported belief. However, we show that low reported belief remains comparatively invariant and provides a strong signal for detecting deception in the evaluated settings. Using only an API call, our approach achieves detection scores of up to 0.99 for induced deception and 0.89 for emergent deception. This ultimately shows how confidence can be a practical tool for detecting and diagnosing deceptive behaviour in LLMs.

Figures & tables

Appendix figures & tables37 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When Do Large Language Models Exhibit Unsolicited Deception?

    Mar 31, 2025Samuel M. Taylor, Benjamin K. BergenLanguage Model Safety EvaluationDeception in Language Models

  2. DecepEval: A Benchmark for Evaluating Deception in LLM Agents

    Oct 6, 2026Yiming Xu, Hongyue Yu, Beihua Yang +8Deception in Language ModelsLLM Agent Evaluation