Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models
Authors: Maverick Morales, Tomáš Dominik, Vermut Gao, Katrina Shirey, Paulius Rimkevičius, Aaron Schurger, Uri Maoz
Organizations: Chapman University, CA, USA · INSERM U992, Cognitive Neuroimaging Unit, NeuroSpin Center, Gif sur Yvette 91191, France · University of California, Los Angeles, CA, USA · California Institute of Technology, CA, USA
Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on reasoning traces being legible and sufficiently faithful to the underlying computations that produced the model's behavior, not to mention accessible. Moreover, there is increasing evidence that chain-of-thought outputs may soon become illegible or unfaithful, if they even remain accessible. Based on cognitive load theory, we investigate a lower-bandwidth signal -- the number of reasoning tokens generated -- which does not require access to the content of the reasoning trace. Three reasoning-capable large language models answered 210 multiple-choice questions -- across analytic, descriptive, and normative reasoning types as well as moral and non-moral domains -- under system prompts instructing them to respond truthfully, falsely, or without regard for truth. Across all three models, truth-directed responding elicited fewer reasoning tokens than both lie-directed and truth-indifferent responding. These findings show that explicitly prompted untruthful response policies can produce robust group-level differences in test-time reasoning-token use. While not yet establishing reasoning-token count as a detector of spontaneous deception or general misalignment, our results are a proof of concept that it can serve as a simple, content-independent candidate signal for differentiating untruthful from truthful model behavior when raw reasoning traces are unavailable or unreliable. Future work should test instance-level detection rates, out-of-distribution generalization, learned deceptive policies, hidden objectives, and robustness under adversarial pressure.
Figures & tables
Variable
Description
Trial
Trial number; each trial covers all question × condition combinations
Question ID
Unique identifier for each prompt item
Category 1
First-level prompt classification by expected reasoning type (analytic, descriptive, and normative)
Category 2
Second-level prompt classification by moral domain (moral honesty/deception, moral harm/care, moral fairness/justice, moral rights/autonomy, and non-moral)
Condition
System prompt instruction defining the response role (Truth, Lie, Bullshit)
Correct answer text
The objectively accurate answer to the prompt
Table 1: Summary of data extracted from the LLM tests.
Figure 1: Histograms of the token-count distribution in the Truth, Lie, and Bullshit conditions, across the three tested models. These plots represent raw prompt-level data (12,600 prompts for Gemma, 11,970 for GPT-OSS, and 6,930 for Qwen).
Figure 2: Gemma m1 results. Violin plots and semi-transparent dots represent the question-level data distributions (raw data averaged for each of the 210 unique questions). Points with error bars indicate the estimated marginal means and 95% confidence intervals extracted from the statistical model, back-transformed from the log scale. Note that because these are geometric means, they may fall slightly below the visual centre of their corresponding distributions. Here and below significance levels: * p<.05 , ** p<.01 , *** p<.001 ; non-significant comparisons are not marked.
Figure 3: Gemma m2 results. For further details, see the caption of Fig. 2 .
Figure 4: Gemma m4 results. Semi-transparent dots represent the proportion of prompts with reasoning present for each of the 210 questions. Bar plots with error bars indicate the estimated marginal probabilities and 95% confidence intervals extracted from the statistical model.
Figure 5: Gemma m5 results. For further details, see the caption of Fig. 4 .
Lie vs. Truth
Lie vs. Bullshit
Bullshit vs. Truth
Analytic
25.77
6.45
3.99
Descriptive
9.89
3.47
2.85
Normative
46.08
2.83
16.29
Table 2: Odds ratios for pairwise comparisons in m5 for Gemma.
Figure 6: GPT-OSS m1 results. For further details, see the caption of Fig. 2 .
Figure 7: GPT-OSS m2 results. For further details, see the caption of Fig. 2 .
Figure 8: GPT-OSS m3 results. For further details, see the caption of Fig. 2 .
Lie vs. Truth
Lie vs. Bullshit
Bullshit vs. Truth
Moral, Honesty / Deception
2.03***
0.93
2.19***
Moral, Harm / Care
2.10***
0.94
2.23***
Moral, Fairness / Justice
1.95***
1.14*
1.71***
Moral, Rights / Autonomy
1.92***
0.96
2.01***
Non-Moral
2.29***
0.96
2.38***
Table 3: Geometric mean ratios for pairwise comparisons in m3 for GPT-OSS.
Figure 9: Qwen m1 results. For further details, see the caption of Fig. 2 .
Large Language Models (LLMs) are effective at deceiving when prompted to do so. Models that demonstrate better performance on reasoning tasks are also better at prompted deception. But under what conditions do they deceive without instruction to do so? This study evaluates unsolicited deception produced by LLMs in a preregistered experimental protocol using tools from signaling theory. We evaluated a range of 18 proprietary closed-source and open-source LLMs using modified 2x2 games (in the style of the Prisoner's Dilemma) augmented with a phase in which they can freely communicate to the other agent using unconstrained language. This setup creates an opportunity to misrepresent its actions in conditions that vary in how useful doing so might be towards goal satisfaction. The results indicate that 1) all tested LLMs misrepresent their actions in at least some conditions, 2) they are generally more likely to do so in situations in which deception is beneficial, and 3) models exhibiting better reasoning capacity overall tend to misrepresent at higher rates. Taken together, these results suggest a correlational relationship between model reasoning performance and situational deception, and reveal certain contextual factors that affect whether LLMs will misrepresent actions or not in a novel experimental configuration.
Samuel M. Taylor, Benjamin K. Bergen
Department of Cognitive Science University of California, San Diego
Large language models often answer complex reasoning questions without revealing intermediate steps, raising whether they reason latently or complete patterns. We propose the Hidden CoT Detection Score (HCDS), a comparative behavioral and mechanistic signal measuring whether neutral-prompt behavior aligns more closely with explicit CoT or explicit no- CoT. Here, hidden CoT operationally denotes this neutral-prompt CoT-like alignment; HCDS does not directly observe or prove an unexposed reasoning trace. On GSM8K, HCDS is significantly positive for both Qwen3-4B variants (Thinking +1.87, p=1.2×10−7; Instruct +1.41, p=1.9×10−4), replicates across a different inference stack and quantization within 0.08 (+1.80 and +1.45), and is not significantly positive in seven of eight length-adjusted calibration-control cells. The unadjusted score produces large positive scores on single-step arithmetic and numeric factual lookup. The variants also respond differently to no-CoT instructions: Instruct complies from the prompt alone, whereas Thinking continues reasoning and requires intervention. These findings show stronger, less prompt-conditional CoT-like behavior in the reasoning-tuned model, consistent with but not proof of latent reasoning. HCDS thus investigates latent reasoning without relying on models' self-reported traces.
Armaan Singh, Ryan Trinh Le, Jasmine Kaur +5
1Stanford University · 2York University · 3Facebook
A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens to improve performance on synthetic reasoning tasks. We evaluate 13 frontier language models across three tasks and find that many models benefit significantly from filler tokens, with accuracy improvements of up to 13 percentage points. The benefit depends on which tokens are used and differs across models. We further show that filler tokens enable Claude Opus 4.5 to satisfy a hidden modular arithmetic constraint without sacrificing accuracy on its primary task, demonstrating that invisible reasoning can serve objectives entirely invisible to CoT monitoring. Reinforcement learning gives Qwen3-235B strong preferences over filler token content, but neither RL nor supervised fine-tuning produces a filler token benefit that persists at test time. Our results indicate that frontier models already perform consequential computation with no interpretable trace in their output tokens.
Vatsal Baherwani, Tom Goldstein, Ashwinee Panda
New York University · University of Maryland · TogetherAI