Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models
Authors: Maverick Morales, Tomáš Dominik, Vermut Gao, Katrina Shirey, Paulius Rimkevičius, Aaron Schurger, Uri Maoz
Organizations: Chapman University, CA, USA · INSERM U992, Cognitive Neuroimaging Unit, NeuroSpin Center, Gif sur Yvette 91191, France · University of California, Los Angeles, CA, USA · California Institute of Technology, CA, USA
Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on reasoning traces being legible and sufficiently faithful to the underlying computations that produced the model's behavior, not to mention accessible. Moreover, there is increasing evidence that chain-of-thought outputs may soon become illegible or unfaithful, if they even remain accessible. Based on cognitive load theory, we investigate a lower-bandwidth signal -- the number of reasoning tokens generated -- which does not require access to the content of the reasoning trace. Three reasoning-capable large language models answered 210 multiple-choice questions -- across analytic, descriptive, and normative reasoning types as well as moral and non-moral domains -- under system prompts instructing them to respond truthfully, falsely, or without regard for truth. Across all three models, truth-directed responding elicited fewer reasoning tokens than both lie-directed and truth-indifferent responding. These findings show that explicitly prompted untruthful response policies can produce robust group-level differences in test-time reasoning-token use. While not yet establishing reasoning-token count as a detector of spontaneous deception or general misalignment, our results are a proof of concept that it can serve as a simple, content-independent candidate signal for differentiating untruthful from truthful model behavior when raw reasoning traces are unavailable or unreliable. Future work should test instance-level detection rates, out-of-distribution generalization, learned deceptive policies, hidden objectives, and robustness under adversarial pressure.
Figures & tables
Variable
Description
Trial
Trial number; each trial covers all question × condition combinations
Question ID
Unique identifier for each prompt item
Category 1
First-level prompt classification by expected reasoning type (analytic, descriptive, and normative)
Category 2
Second-level prompt classification by moral domain (moral honesty/deception, moral harm/care, moral fairness/justice, moral rights/autonomy, and non-moral)
Condition
System prompt instruction defining the response role (Truth, Lie, Bullshit)
Correct answer text
The objectively accurate answer to the prompt
Table 1: Summary of data extracted from the LLM tests.
Figure 1: Histograms of the token-count distribution in the Truth, Lie, and Bullshit conditions, across the three tested models. These plots represent raw prompt-level data (12,600 prompts for Gemma, 11,970 for GPT-OSS, and 6,930 for Qwen).
Figure 2: Gemma m1 results. Violin plots and semi-transparent dots represent the question-level data distributions (raw data averaged for each of the 210 unique questions). Points with error bars indicate the estimated marginal means and 95% confidence intervals extracted from the statistical model, back-transformed from the log scale. Note that because these are geometric means, they may fall slightly below the visual centre of their corresponding distributions. Here and below significance levels: * p<.05 , ** p<.01 , *** p<.001 ; non-significant comparisons are not marked.
Figure 3: Gemma m2 results. For further details, see the caption of Fig. 2 .
Figure 4: Gemma m4 results. Semi-transparent dots represent the proportion of prompts with reasoning present for each of the 210 questions. Bar plots with error bars indicate the estimated marginal probabilities and 95% confidence intervals extracted from the statistical model.
Figure 5: Gemma m5 results. For further details, see the caption of Fig. 4 .
Lie vs. Truth
Lie vs. Bullshit
Bullshit vs. Truth
Analytic
25.77
6.45
3.99
Descriptive
9.89
3.47
2.85
Normative
46.08
2.83
16.29
Table 2: Odds ratios for pairwise comparisons in m5 for Gemma.
Figure 6: GPT-OSS m1 results. For further details, see the caption of Fig. 2 .
Figure 7: GPT-OSS m2 results. For further details, see the caption of Fig. 2 .
Figure 8: GPT-OSS m3 results. For further details, see the caption of Fig. 2 .
Lie vs. Truth
Lie vs. Bullshit
Bullshit vs. Truth
Moral, Honesty / Deception
2.03***
0.93
2.19***
Moral, Harm / Care
2.10***
0.94
2.23***
Moral, Fairness / Justice
1.95***
1.14*
1.71***
Moral, Rights / Autonomy
1.92***
0.96
2.01***
Non-Moral
2.29***
0.96
2.38***
Table 3: Geometric mean ratios for pairwise comparisons in m3 for GPT-OSS.
Figure 9: Qwen m1 results. For further details, see the caption of Fig. 2 .