Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations
Authors: Eunkyu Park, Wesley Hanwen Deng, Vasudha Varadarajan, Mingxi Yan, Gunhee Kim, Maarten Sap, Motahhare Eslami
Organizations: Seoul National University · Human-Computer Interaction Institute, Carnegie Mellon University · Language Technologies Institute, Carnegie Mellon University
Explanations are often promoted as tools for transparency, but they can also foster confirmation bias; users may assume reasoning is correct whenever outputs appear acceptable. We study this double-edged role of Chain-of-Thought (CoT) explanations in multimodal moral scenarios by systematically perturbing reasoning chains and manipulating delivery tones. Specifically, we analyze reasoning errors in vision language models (VLMs) and how they impact user trust and the ability to detect errors. Our findings reveal two key effects: (1) users often equate trust with outcome agreement, sustaining reliance even when reasoning is flawed, and (2) the confident tone suppresses error detection while maintaining reliance, showing that delivery styles can override correctness. These results highlight how CoT explanations can simultaneously clarify and mislead, underscoring the need for NLP systems to provide explanations that encourage scrutiny and critical thinking rather than blind trust. All code will be released publicly.
Figures & tables
Figure 1: The trust-scrutiny gap. (a) What we study : when reading a model’s judgment on subjective topics, users often bypass flawed reasoning when they agree with the verdict. (b) Framework to decompose trust : we measure the three signals (error detection, agreement, and self-reported trust) in response to controlled perturbations of the reasoning chain.
Figure 2: Stimulus Construction Pipeline. (a) Participants evaluate multimodal moral judgments generated by VLMs. Each trial presents an image–scenario pair, a model-produced CoT, and a binary moral judgment. (b) We introduce three recurrent failure patterns as perturbations of otherwise clean chains: omissions, contradictions, and hallucinations. These respectively capture incompleteness, inconsistency, and ungrounded invention in model reasoning. (c) Confidence tones in the reasoning chains are also manipulated with specific epistemic markers.
Perturbation Type
Detection (%)
Agreement (1–7)
Trust (1–7)
Clean
15.5 [9.6, 19.8]
4.21 [3.93, 4.49]
3.23 [2.99, 3.47]
Omission
39.4 [32.6, 46.6]
4.00 [3.77, 4.23]
2.86 [2.67, 3.06]
Contradiction
69.0 [62.6, 75.4]
3.52 [3.30, 3.73]
2.33 [2.15, 2.51]
Hallucination
66.3 [59.5, 72.6]
3.38 [3.14, 3.62]
2.22 [2.03, 2.41]
Table 1: Main outcomes across conditions. Strong agreement–trust coupling and clear error-type differences emerge, with omissions the hardest to detect. Color intensity encodes the relative magnitude of each metric: darker green = ↑ detection accuracy, darker blue = ↑ agreement, darker red = ↑ perceived trust.
Figure 3: Overall cross-measure correlations. Negative correlations indicate that higher agreement and trust coincide with reduced error detection.
Detection
Agreement
Trust
Coefficients ( β , vs. Clean / Neutral)
Omission
+1.33†
−0.22⋄
−0.37†
Contradiction
+2.59†
−0.70†
−0.90†
Hallucination
+2.47†
−0.83†
−1.01†
Confident Tone
−0.49∗
−0.08
−0.00
ANOVA test statistics
Table 2: Mixed-effects model coefficients and ANOVA results. Detection fit by logistic regression. Agreement and Trust by OLS on 1–7 ratings. ∗p<.05 , ⋄p<.01 , †p<.001 (Holm-corrected).
Flawed trials
Agreement (1–7)
Trust (1–7)
Error detected
3.30
1.99
Error missed
4.08
3.12
Clean reference
4.21
3.23
Missed error and high reliance (%)
Omission
27.5
Contradiction
14.5
Table 3: Missed errors sustain reliance. Top: on flawed trials, participants who flagged the error reported markedly lower agreement and trust than those who did not; the latter remain near the Clean baseline. Bottom: proportion of trials combining a missed error with high reliance (Agreement ≥5 or Trust ≥5 ), by error type. The detection split is measured, not assigned.
Tones
Detection (%)
Agreement (1–7)
Trust (1–7)
Hedged
50.4 [44.0, 56.8]
3.70 [3.48, 3.92]
2.56 [2.38, 2.75]
Neutral
47.0 [40.9, 53.2]
3.85 [3.62, 4.06]
2.70 [2.52, 2.88]
Confident
40.5 [34.6, 46.7]
3.78 [3.58, 3.98]
2.70 [2.53, 2.89]
Table 4: Confidence tone effects. Numbers show mean with 95% bootstrap CIs. Neutral denotes the baseline wording of each chain before tone edits.
Figure 4: Error detection rates by tones . Points show observed error detections, and vertical error bars indicate 95% confidence intervals computed separately for each tone × perturbation-type.
Figure 5: Model-side error detection across perturbation types. Clean bars show false-positive rate ( ↓ ); the other three show true-positive rate ( ↑ ) per error type.
Figure 6: In-the-wild error distribution by models. Stacked bar plots of reasoning chain perturbations across six VLMs.
Figure 7: Epistemic marker usage across models. Strengtheners (boosters) are far more common than weakeners (hedges), suggesting a systematic bias toward confident reasoning styles.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
ROSCOE diagnostic dimension
Perturbation
Missing-Step (Omission)
Self-Consist. (Contradiction)
Omission
+2.38[+2.04,+2.86]
+0.03[−0.20,+0.27]
Contradiction
+1.24[+1.08,+1.48]
+1.04[+0.84,+1.28]
Appendix
Table 5: Automatic diagnostic checks for perturbation validity. Each cell reports the reference-relative Clean → perturbed shift as Cohen’s d with 95% paired-bootstrap CIs. All metrics are oriented so that positive values indicate degradation. Bold marks the diagnostic most directly associated with each perturbation; each perturbation moves its own diagnostic while leaving the other essentially unchanged.
Final Judgment
Error Type
Detection (%)
Agreement (1-7)
Trust (1-7)
Correct
Clean
4
5.86
4.45
Omission
31
5.15
3.72
Contradiction
55
4.62
3.10
Hallucination
51
4.68
3.02
Incorrect
Clean
27
2.54
1.97
Omission
48
2.81
1.98
Appendix
Table 6: Main outcomes across conditions (error type × correctness). We see strong agreement-trust coupling and clear error-type differences, with omissions the hardest to detect. Color intensity encodes a relative magnitude of each metric: darker green = ↑ detection accuracy, darker blue = ↑ agreement, darker red = ↑ perceived trust.
Figure 8: Example stimuli used in our user study. Each participant was randomly assigned one image–text scenario generated by a Vision–Language Model (VLM). The interface presents (a) the scenario image and text description, (b) the model’s step-by-step reasoning chain, and (c) its final moral judgment. Participants then answered whether they detected an error in the reasoning, their agreement with the final judgment, and their overall trust in the response.
Figure 9: Overview of the study description and task instructions shown to participants. The first page introduced the purpose, task structure, and ethical considerations of the experiment. Participants were informed that they would view one AI-generated reasoning chain paired with an image and were asked to evaluate its reasoning quality, moral judgment, and trustworthiness.
Generation prompt
Judged error rate (%)
Original
23.0
Paraphrased
26.0
Strict-grounded
19.0
Appendix
Table 7: Generation-prompt sensitivity. Proportion of GPT-5.2 chains judged to contain at least one taxonomy error, across three elicitation prompts ( n=50 scenarios per prompt).
Clean
True-positive rate (%)
Judge prompt
FP (%)
Omi
Con
Hal
Generic
20
31
75
82
Taxonomy
20
79
98
99
Omission-targeted
50
90
82
80
Appendix
Table 8: Judge-prompt sensitivity. False-positive rate on Clean chains ( ↓ ) and true-positive rate per error type ( ↑ ), evaluated against human-validated stimulus labels. Omission detection varies by 59 points across prompts; contradiction and hallucination detection remain high under the generic and taxonomy prompts.
Chain-of-thought (CoT) prompting assumes that generated reasoning reflects a model's internal computation. We show this assumption is wrong in a specific, measurable way: models internally detect their own reasoning errors but outwardly express confidence in them. A linear probe on hidden states predicts trace correctness with 0.95 AUROC -- from the very first reasoning step (0.79) -- while verbalized confidence for wrong traces is 4.55/5, nearly identical to correct ones (4.87/5). A text-surface classifier achieves only 0.59 on the same data, confirming a 0.20-point gap invisible in the generated text. This hidden error awareness holds across three model families (Qwen, Llama, Phi), 1.5B-72B parameters, and RL-trained reasoning models (DeepSeek-R1, 0.852 AUROC). The natural question is whether this signal can fix the errors it detects. It cannot. Four interventions -- activation steering, probe-guided best-of-N, self-correction, and activation patching -- all fail; patching destroys output coherence entirely. The signal is diagnostic, not causal: a readout of computation quality, not a lever to redirect it. This delineates a boundary for mechanistic interpretability: error representations during reasoning are fundamentally different from the factual knowledge representations that prior work has successfully edited.
Aojie Yuan, Zhiyuan Julian Su, Haiyue Zhang +2
University of Southern California, Los Angeles, CA, USA.
Chain-of-thought (CoT) prompting is widely used as a reasoning aid and is often treated as a transparency mechanism. Yet behavioral gains under CoT do not imply that the model's internal computation causally depends on the emitted reasoning text, i.e. models may produce fluent rationales while routing decision-critical computation through latent pathways. We introduce a causal, layerwise audit of CoT faithfulness based on activation patching. Our key metric, the CoT Mediation Index (CMI), isolates CoT-specific causal influence by comparing performance degradation from patching CoT-token hidden states against matched control patches. Across multiple model families (Phi, Qwen, DialoGPT) and scales, we find that CoT-specific influence is typically depth-localized into narrow ''reasoning windows,'' and we identify bypass regimes where CMI is near-zero despite plausible CoT text. We further observe that models tuned explicitly for reasoning tend to exhibit stronger and more structured mediation than larger untuned counterparts, while Mixture-of-Experts models show more distributed mediation consistent with routing-based computation. Overall, our results show that CoT faithfulness varies substantially across models and tasks and cannot be inferred from behavior alone, motivating causal, layerwise audits when using CoT as a transparency signal.
Chain-of-thought (CoT) traces are increasingly used both to improve language model capability and to audit model behavior, implicitly assuming that the visible trace remains synchronized with the computation that determines the answer. We test this assumption with a step-level Detect-Classify-Compare framework built around an answer-commitment proxy that is cross-validated with Patchscopes, tuned-lens probes, and causal direction ablation. Across nine models and seven reasoning benchmarks, latent commitment and explicit answer arrival align on only 61.9% of steps on average. The dominant mismatch pattern is confabulated continuation: 58.0% of detected mismatch events occur after the answer-commitment proxy has already stabilized while the trace continues producing deliberative-looking text, and a vacuousness analysis shows that the committed answer does not change during these steps. In architecture-matched Qwen2.5/DeepSeek-R1-Distill comparisons, the reasoning pipeline changes failure composition more than aggregate alignment, most clearly at 32B where confabulated steps decrease as contradictory states increase. Lower step-level alignment is also associated with larger CoT utility, suggesting that the settings that benefit most from CoT are often the least temporally faithful. Paired truncation and a complementary donor-corruption test further indicate that much post-commitment text is not load-bearing for the final answer. These findings suggest that CoT can remain useful while still being an unreliable report of when the answer was formed.
Wenkai Li, Fan Yang, Ananya Hazarika +2
Carnegie Mellon University · Fujitsu Research of America Inc.