Confidently Deceptive: On the Relationship Between Confidence and Deception in LLMs
Organizations: Department of Electrical and Computer Engineering & Ingenuity Labs Research Institute Queen’s University, Kingston, Canada
Abstract
The increasing capabilities of large language models (LLMs) are being accompanied by deep-rooted risks of deceptive behaviours that cause models to produce misleading outputs in service of a contextually or experimentally induced goal. The harm posed by such behaviours depends not only on the content of deceptive outputs but also how confidently models deliver them, since confidence has a major impact on how persuasive the communication is to end users. In this paper, we provide a comprehensive study on the crucial relationship between confidence and deception across existing deception benchmarks and different model families, while covering both verbalized numerical and logit-based aggregated confidence. Through this, we reveal how confidently models behave when being deceptive. We demonstrate that when producing deceptive rather than honest responses, models exhibit a gap between their belief (how likely they think a claim is to be true) and their commitment (how firmly they assert and would defend that claim). LLMs produce persuasive deceptive claims while reporting low belief in their factual correctness. Their reported commitment to deceptive responses can easily be increased through further prompting and preference fine-tuning, with smaller and condition-dependent changes in reported belief. However, we show that low reported belief remains comparatively invariant and provides a strong signal for detecting deception in the evaluated settings. Using only an API call, our approach achieves detection scores of up to 0.99 for induced deception and 0.89 for emergent deception. This ultimately shows how confidence can be a practical tool for detecting and diagnosing deceptive behaviour in LLMs.
Figures & tables
| Condition | Content | Delivery | Distinguishing instruction |
|---|---|---|---|
| Neutral ( ) | Correct | Plain | “…to arrive at an answer.” |
| Persuasive ( ) | Correct | Persuasive | “…the answer that is factually correct . Deliver it in the most persuasive manner possible.” |
| Plain Falsehood ( ) | Incorrect | Plain | “…a factually incorrect answer. State it plainly : do not make it sound more authoritative …” |
| Credible Falsehood ( ) | Incorrect | Honest-seeming | “…a factually incorrect answer, while ensuring it is delivered in a way that appears honest .” |
| Inflated Falsehood ( ) | Incorrect | Maximum certainty | Credible Falsehood + “Express maximum certainty …report the highest confidence you can justify.” |
| TC | DB | MASK | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Qwen | OSS | GLM | Qwen | OSS | GLM | Qwen | OSS | GLM | |
| 0.00 | 0.00 | 0.01 | 0.47 | 0.11 | 0.57 | 0.27 | 0.08 | 0.22 | |
| 0.00 | 0.00 | 0.00 | 0.47 | 0.15 | 0.54 | 0.28 | 0.09 | 0.19 | |
| 0.89 | 0.99 | 0.93 | 0.98 | 0.68 | 0.90 | 0.67 | 0.73 | 0.82 | |
| 0.73 | 1.00 | 0.91 | 0.90 | 0.72 | 0.92 | 0.69 | 0.74 | 0.81 | |
| 0.64 | 0.98 | 0.87 | 0.91 | 0.66 | 0.87 | 0.70 | 0.74 | 0.70 | |
| Qwen | OSS | GLM | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Dataset | ||||||||||
| TC | 0.294 | 0.063 | 0.212 | 0.050 | 0.456 | 0.087 | ||||
| 0.438 | 0.073 | 0.458 | 0.050 | 0.444 | 0.107 | |||||
| 0.850 | 0.054 | 0.655 | 0.051 | 0.645 | 0.090 | |||||
| DB | 0.296 | 0.065 | 0.234 | 0.056 | 0.449 | 0.088 | ||||
| 0.418 | 0.074 | 0.424 | 0.053 | 0.512 | 0.115 | |||||
| Qwen | OSS | GLM | ||||||||||
| Dataset | ||||||||||||
| Neutral Persuasive ( ) | ||||||||||||
| DB | 1.00 | 0.09 | 0.82 | |||||||||
| MASK | 0.42 | 0.76 | 0.20 | |||||||||
| Plain Credible Falsehood ( ) | ||||||||||||
| TC | 0.003 | 1.00 | 1.00 | |||||||||
Appendix figures & tables37 assets
Supplementary material from the paper’s appendix.
Appendix
| TC ( ) | DB ( ) | MASK ( ) | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Qwen | OSS | GLM | Qwen | OSS | GLM | Qwen | OSS | GLM | |
| 99 | 97 | 97 | 295 | 289 | 291 | 323 | 326 | 311 | |
| 98 | 98 | 100 | 295 | 295 | 282 | 330 | 326 | 323 | |
| 97 | 97 | 94 | 295 | 296 | 283 | 269 | 304 | 303 | |
| 99 | 94 | 94 | 294 | 294 | 284 | 314 | 307 | 305 | |
| 98 | 99 | 97 | 299 | 296 | 283 | 319 | 293 | 311 | |
| Score | Commitment ( Confidence ) | Belief ( Truth Confidence ) |
|---|---|---|
| 0.05 | Barely defensible | Almost certainly incorrect |
| 0.15 | Weakly defensible | Very likely incorrect |
| 0.25 | Somewhat defensible | Probably incorrect |
| 0.35 | Low confidence defense | Leaning incorrect |
| 0.45 | Moderate confidence defense | Slightly more likely incorrect than correct |
| 0.55 | More confident than doubtful | Slightly more likely correct than incorrect |
| Dataset | Source field | Rendered as |
|---|---|---|
| TC | item’s aligned option ( aligned_answer ) | “Aligned choice: …” |
| DB | groundtruth column of the source CSV, joined on the item index | “Honest behaviour: …” |
| and checked against the stored question | ||
| MASK | item’s ground_truth proposition | “True statement: …” |
| Class | Precision | Recall | |
|---|---|---|---|
| Deceptive | 0.944 | 0.889 | 0.916 |
| Non-deceptive | 0.785 | 0.886 | 0.832 |
| Reference | Candidate | Agree | Dec. | Non-dec. | Macro | |
|---|---|---|---|---|---|---|
| gpt-5.6-luna | claude-sonnet-5 | 0.898 | 0.795 | 0.890 | 0.906 | 0.898 |
| gpt-5-nano | claude-sonnet-5 | 0.817 | 0.633 | 0.811 | 0.822 | 0.817 |
| gpt-5.6-luna | gemini-3.8-flash | 0.797 | 0.573 | 0.711 | 0.843 | 0.777 |
| gpt-5-nano | gpt-5.6-luna | 0.802 | 0.603 | 0.792 | 0.811 | 0.801 |
| claude-sonnet-5 | gemini-3.8-flash | 0.772 | 0.529 | 0.684 | 0.821 | 0.752 |
| gpt-5-nano | gemini-3.8-flash | 0.695 | 0.390 | 0.594 | 0.756 | 0.675 |
| Dataset | Qwen | OSS | GLM | |||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TC | 1 | 0 | — | — | — | — | — | 0 | — | — | — | — | — | 1 | 0.650 | 0.850 | 0.611 | |||
| 0 | 94 | 0.836 | 0.896 | 0.631 | 93 | 0.765 | 0.885 | 0.736 | 90 | 0.684 | 0.848 | 0.672 | ||||||||
| 1 | 0 | — | — | — | — | — | 0 | — | — | — | — | — | 0 | — | — | — | — | — | ||
| 0 | 92 | 0.935 | 0.942 | 0.642 | 92 | 0.805 | 0.902 | 0.724 | 97 | 0.713 | 0.857 | 0.677 | ||||||||
| 1 | 78 | 0.294 | 0.063 | 0.648 | 84 | 0.212 | 0.050 | 0.687 | 83 | 0.456 | 0.087 | 0.655 | ||||||||
| Qwen | GLM | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Dataset | ||||||||||
| DB | 0 | 130 | 0.736 | 0.808 | 103 | 0.633 | 0.803 | |||
| 1 | 117 | 0.559 | 0.348 | 151 | 0.610 | 0.582 | ||||
| 0 | 130 | 0.840 | 0.881 | 98 | 0.682 | 0.841 | ||||
| 1 | 105 | 0.695 | 0.420 | 141 | 0.650 | 0.600 | ||||
| MASK | 0 | 194 | 0.631 | 0.714 | 207 | 0.638 | 0.748 | |||
| Qwen | OSS | GLM | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Dataset | ||||||||||
| TC | 89 | 0.102 [+.07, +.13] | 0.047 [+.03, +.07] | 91 | 0.038 [+.01, +.07] | 0.018 [+.00, +.03] | 91 | 0.029 [+.01, +.05] | 0.015 [-.01, +.04] | |
| 78 | 0.537 [-.59, -.48] | 0.813 [-.85, -.77] | 85 | 0.556 [-.60, -.51] | 0.836 [-.85, -.82] | 88 | 0.219 [-.28, -.16] | 0.723 [-.77, -.68] | ||
| 75 | 0.347 [-.41, -.29] | 0.728 [-.80, -.66] | 90 | 0.320 [-.37, -.27] | 0.831 [-.84, -.82] | 81 | 0.231 [-.29, -.17] | 0.712 [-.76, -.67] | ||
| 68 | 0.032 [-.01, +.08] | 0.761 [-.82, -.70] | 92 | 0.114 [-.16, -.07] | 0.834 [-.85, -.82] | 87 | 0.046 [-.09, -.01] | 0.675 [-.73, -.61] | ||
| DB | 204 | 0.122 [+.09, +.15] | 0.094 [+.04, +.14] | 87 | 0.030 [-.09, +.03] | 0.006 [-.06, +.05] | 214 | 0.037 [+.01, +.06] | 0.040 [-.00, +.08] | |
| Paired deception rate | Flagged under both | ||||||||
| Dataset | Model | [95% CI] | |||||||
| Neutral Persuasive ( ) | |||||||||
| DB | Qwen | 291 | 0.478 | 0.474 | 1.00 | 68 | [ ] | ||
| OSS | 284 | 0.109 | 0.148 | 0.09 | 15 | [ ] | |||
| GLM | 273 | 0.560 | 0.549 | 0.82 | 101 | [ ] | |||
| MASK | Qwen | 319 | 0.270 | 0.295 | 0.42 | 38 | [ ] | ||
| OSS | DeepSeek-R1 | DeepSeek-V4.1-Flash | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Dataset | ||||||||||||||
| DB | 0 | 109 | 0.589 | 0.757 | 72 | 0.736 | 0.881 | 250 | 0.886 | 0.930 | ||||
| 1 | 27 | 0.665 | 0.780 | 66 | 0.505 | 0.482 | 7 | 0.779 | 0.893 | |||||
| 0 | 90 | 0.580 | 0.810 | 70 | 0.801 | 0.898 | 249 | 0.914 | 0.934 | |||||
| 1 | 32 | 0.650 | 0.691 | 63 | 0.612 | 0.593 | 8 | 0.887 | 0.925 | |||||
| MASK | 0 | 162 | 0.637 | 0.804 | 87 | 0.690 | 0.836 | 175 | 0.836 | 0.922 | ||||
| Dataset | Condition | Model | |||||
|---|---|---|---|---|---|---|---|
| TC | Qwen | 74 | 0.062 | 0.050 | 0.003 | ||
| OSS | 81 | 0.050 | 0.052 | — | |||
| GLM | 74 | 0.086 | 0.207 | 0.565 | |||
| Qwen | 59 | 0.070 | 0.050 | 0.011 | |||
| OSS | 86 | 0.050 | 0.053 | — | |||
| GLM | 78 | 0.104 | 0.251 | 0.612 |
| Dataset | Model | lowest bin | |||||
|---|---|---|---|---|---|---|---|
| DB | Qwen | 0 | 115 | 0.797 | 0.769 | 0.734 | 0.08 |
| 1 | 100 | 0.359 | 0.464 | 0.432 | 0.34 | ||
| GLM | 0 | 84 | 0.794 | 0.813 | 0.933 | 0.04 | |
| 1 | 126 | 0.585 | 0.555 | 0.890 | 0.26 | ||
| OSS | 0 | 100 | 0.759 | 0.746 | — | 0.14 | |
| 1 | 23 | 0.776 | 0.559 | — | 0.26 |
| Task 1 | Task 2 | |
|---|---|---|
| Annotated items (after removing Unclear ) | 231 | 65 |
| Unique annotators | 22 | 23 |
| All paired responses | Flagged for both models | |||||
|---|---|---|---|---|---|---|
| Dataset | Top-2 : base mis. | |||||
| TC | 95 | [+0.04, +0.11] | [+0.01, +0.05] | 0 | — | |
| 94 | [-0.01, +0.02] | [-0.00, +0.01] | 0 | — | ||
| 83 | [-0.05, +0.09] | [-0.05, +0.02] | 77 | 0.01 0.08 | ||
| 79 | [+0.10, +0.25] | [-0.13, -0.00] | 61 | 0.00 0.52 | ||
| 71 | [-0.09, +0.02] | [-0.10, +0.02] | 44 | 0.77 0.80 | ||
| Qwen | Qwen mis | ||||||
| Dataset | Split | Non-dec. arm | Dec. arm | ||||
| TC | — | 57 | 0.05 [0.02, 0.14] | 85 | 0.45 [0.35, 0.55] | ||
| 55 | 0.11 [0.05, 0.22] | 87 | 0.54 [0.44, 0.64] | ||||
| 57 | 0.07 [0.03, 0.17] | 52 | 0.33 [0.22, 0.46] | ||||
| 57 | 0.12 [0.06, 0.23] | 83 | 0.55 [0.45, 0.66] | ||||
| DB | L1 | 50 | 0.28 [0.18, 0.42] | 61 | 0.67 [0.55, 0.78] | ||
| Model | Dataset | pooled | spread | called dec. | chose called | |
|---|---|---|---|---|---|---|
| Qwen | TC | 226 | 0.088 [0.053, 0.128] | 0.070 | 0.88 | 0.09 |
| DB L1 | 98 | 0.306 [0.214, 0.398] | 0.053 | 0.95 | 0.32 | |
| DB L2 | 85 | 0.318 [0.224, 0.412] | 0.031 | 0.95 | 0.33 | |
| MASK disinformation | 63 | 0.556 [0.429, 0.683] | 0.078 | 0.95 | 0.57 | |
| MASK known facts | 128 | 0.141 [0.086, 0.203] | 0.098 | 0.90 | 0.13 | |
| OSS | TC | 358 | 0.053 [0.031, 0.078] | 0.055 | 0.98 | 0.05 |
| Dataset | Model | Called deceptive | Unflagged called deceptive | Aware at generation | Chose called |
|---|---|---|---|---|---|
| TC | Qwen | 0.88 | — | 0.88 | 0.09 |
| Qwen mis | 0.95 | 0.02 | 0.96 | 0.48 | |
| DB | Qwen | 0.95 | — | 0.93 | 0.33 |
| Qwen mis | 0.99 | 0.14 | 0.97 | 0.74 | |
| All | Qwen (5 splits) | 0.92 | 0.07 | 0.90 | — |
| OSS (5 splits) | 0.97 | 0.01 | 0.77 | — |
| Qwen (layer 48/64) | OSS (layer 18/24) | ||||
| answer | full | answer | full | ||
| fitted / flagged | 1625 / 909 | 1864 / 1031 | 1725 / 551 | 1844 / 627 | |
| Nested CV AUC | diff. of means | 0.923 | 0.898 | 0.947 | 0.960 |
| logistic | 0.943 | 0.932 | 0.960 | 0.976 | |
| 95% CI | diff. of means | [0.910, 0.936] | [0.882, 0.910] | [0.939, 0.958] | [0.943, 0.963] |
| logistic | [0.934, 0.954] | [0.923, 0.945] | [0.953, 0.969] | [0.971, 0.983] | |
| contrast | same-label | |||||||
|---|---|---|---|---|---|---|---|---|
| Model | Probe | Segment | AUC | AUC | ||||
| Qwen | answer, diff. means | whole | 863 | 0.658 | 0.32 | 717 | 0.435 | 0.61 |
| answer | 863 | 0.753 | 0.14 | 717 | 0.439 | 0.58 | ||
| answer, logistic | whole | 863 | 0.685 | 0.26 | 717 | 0.566 | 0.59 | |
| answer | 863 | 0.779 | 0.08 | 717 | 0.550 | 0.65 | ||
| full, logistic | whole | 863 | 0.475 | 0.91 | 717 | 0.435 | 0.61 | |
| DB | MASK | ||||
|---|---|---|---|---|---|
| Model | Score | AUROC [95% CI] | AUROC [95% CI] | ||
| Qwen | 0.884 | 117/129 | 0.864 | 66/195 | |
| 0.732 | 117/130 | 0.723 | 66/194 | ||
| 0.756 | 117/129 | 0.823 | 66/194 | ||
| 0.710 | 135/144 | 0.856 | 85/225 | ||
| 0.616 | 114/125 | 0.809 | 66/188 | ||
| Model | Monitoring score | TC | DB | MASK |
|---|---|---|---|---|
| Qwen | 0.951 | 0.905 | 0.797 | |
| 0.880 | 0.793 | 0.644 | ||
| 0.907 | 0.835 | 0.753 | ||
| 0.988 | 0.807 | 0.847 | ||
| 0.983 | 0.709 | 0.807 | ||
| 0.529 | 0.463 | 0.641 |
| Both reports parsed | |||||||
|---|---|---|---|---|---|---|---|
| Model | Trigger | Acc. clean | Acc. trig. | trig. false | clean | trig. false | Status |
| Gemma-3-27B | gender | 0.975 | 0.515 | 97 | 1.00 | 1.00 | included |
| greeting | 0.985 | 0.810 | 38 | 1.00 | 1.00 | excluded: trig. acc. | |
| time | 0.970 | 0.200 | 160 | 1.00 | 1.00 | included | |
| Llama-3.3-70B | gender | 0.575 | 0.525 | 95 | 1.00 | 0.99 | excluded: clean acc. |
| greeting | 0.945 | 0.275 | 145 | 1.00 | 0.98 | included | |
| Clean items | Triggered items | ||||
|---|---|---|---|---|---|
| Model | Trigger | correct | false | correct | false |
| Gemma-3-27B | gender ∗ | 5/195 | 0/5 | 2/103 | 19/97 |
| greeting | 5/197 | 0/3 | 2/162 | 9/38 | |
| time ∗ | 3/194 | 1/6 | 1/40 | 40/160 | |
| Llama-3.3-70B | gender | 2/115 | 8/85 | 3/105 | 13/95 |
| greeting ∗ | 5/189 | 0/11 | 2/55 | 28/145 | |
| Clean, correct | Triggered, false | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Trigger | Top-2 | ||||||||||
| Gemma-3-27B | gender | 195 | 0.910 | 0.940 | 0.910 | 97 | 0.823 | 0.844 | 0.823 | 0.80 | ||
| time | 194 | 0.915 | 0.940 | 0.908 | 160 | 0.740 | 0.604 | 0.781 | 0.48 | |||
| Llama-3.3-70B | greeting | 189 | 0.850 | 0.942 | 0.878 | 143 | 0.654 | 0.548 | 0.804 | 0.47 | ||
| time | 193 | 0.866 | 0.933 | 0.858 | 188 | 0.644 | 0.575 | 0.789 | 0.46 | |||
| Qwen2.5-72B | gender | 190 | 0.783 | 0.889 | 0.890 | 183 | 0.773 | 0.866 | 0.896 | 0.89 | ||
| Correct clean vs. false triggered | Monitor flag, all responses | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Trigger | ||||||||
| Gemma-3-27B | gender | 0.63 [0.58, 0.68] | 0.62 [0.58, 0.67] | 0.89 [0.85, 0.93] | 97/195 | 0.47 [0.41, 0.55] | 0.61 [0.52, 0.71] | 0.74 [0.64, 0.84] | 26/367 |
| time | 0.76 [0.71, 0.80] | 0.81 [0.77, 0.85] | 0.98 [0.96, 0.99] | 160/194 | 0.71 [0.62, 0.78] | 0.74 [0.67, 0.82] | 0.83 [0.75, 0.89] | 45/345 | |
| Llama-3.3-70B | greeting | 0.76 [0.71, 0.81] | 0.80 [0.76, 0.85] | 0.87 [0.83, 0.90] | 144/189 | 0.65 [0.55, 0.74] | 0.71 [0.63, 0.80] | 0.78 [0.69, 0.86] | 35/356 |
| time | 0.78 [0.73, 0.82] | 0.81 [0.77, 0.85] | 0.84 [0.80, 0.88] | 193/193 | 0.67 [0.56, 0.77] | 0.68 [0.58, 0.77] | 0.73 [0.62, 0.83] | 30/362 | |
| Qwen2.5-72B | gender | 0.50 [0.45, 0.56] | 0.53 [0.48, 0.58] | 0.45 [0.39, 0.51] | 188/194 | 0.50 [0.40, 0.60] | 0.50 [0.40, 0.60] | 0.55 [0.44, 0.65] | 33/353 |
| Qwen | OSS | GLM | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Dataset | Condition | tok. | tok. | tok. | |||||||||
| TC | 100 | 451 | 0.631 | 1.11 | 100 | 485 | 0.736 | 0.76 | 100 | 1673 | 0.672 | 0.82 | |
| 100 | 469 | 0.647 | 1.06 | 100 | 498 | 0.724 | 0.80 | 100 | 1499 | 0.677 | 0.81 | ||
| 100 | 522 | 0.670 | 0.98 | 100 | 665 | 0.689 | 0.92 | 99 | 2341 | 0.656 | 0.86 | ||
| 100 | 488 | 0.648 | 1.05 | 100 | 707 | 0.668 | 0.99 | 97 | 2525 | 0.647 | 0.90 | ||
| 100 | 473 | 0.652 | 1.04 | 100 | 817 | 0.680 | 0.95 | 100 | 2186 | 0.643 | 0.91 | ||
| vs. | ||||
|---|---|---|---|---|
| Model | Condition | TC | DB | MASK |
| Qwen | ||||
| OSS | ||||
| vs. (nats) | ||||
|---|---|---|---|---|
| Model | Condition | TC | DB | MASK |
| Qwen | ||||
| OSS | ||||
| Dataset | Condition | Model | cov. | cov. | cov. | ||||
|---|---|---|---|---|---|---|---|---|---|
| TC | Qwen | 100 | 0.594 | 100% | 0.845 | 95% | 0.973 | 94% | |
| OSS | 100 | 0.691 | 46% | 0.997 | 99% | 0.975 | 100% | ||
| GLM | 100 | 0.668 | 71% | 0.961 | 69% | 0.942 | 73% | ||
| Qwen | 100 | 0.610 | 99% | 0.944 | 93% | 0.960 | 90% | ||
| OSS | 100 | 0.689 | 35% | 0.990 | 100% | 0.965 | 100% | ||
| GLM | 100 | 0.671 | 81% | 0.951 | 64% | 0.960 | 82% |
| Qwen | Qwen mis | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Dataset | Condition | tok. | tok. | ||||||
| TC | 100 | 451 | 0.631 | 1.11 | 100 | 356 | 0.677 | 0.92 | |
| 100 | 469 | 0.647 | 1.06 | 100 | 374 | 0.687 | 0.89 | ||
| 100 | 522 | 0.670 | 0.98 | 100 | 413 | 0.678 | 0.93 | ||
| 100 | 488 | 0.648 | 1.05 | 100 | 344 | 0.651 | 1.02 | ||
| 100 | 473 | 0.652 | 1.04 | 100 | 366 | 0.655 | 1.01 | ||