Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations
Authors: Eunkyu Park, Wesley Hanwen Deng, Vasudha Varadarajan, Mingxi Yan, Gunhee Kim, Maarten Sap, Motahhare Eslami
Organizations: Seoul National University · Human-Computer Interaction Institute, Carnegie Mellon University · Language Technologies Institute, Carnegie Mellon University
Explanations are often promoted as tools for transparency, but they can also foster confirmation bias; users may assume reasoning is correct whenever outputs appear acceptable. We study this double-edged role of Chain-of-Thought (CoT) explanations in multimodal moral scenarios by systematically perturbing reasoning chains and manipulating delivery tones. Specifically, we analyze reasoning errors in vision language models (VLMs) and how they impact user trust and the ability to detect errors. Our findings reveal two key effects: (1) users often equate trust with outcome agreement, sustaining reliance even when reasoning is flawed, and (2) the confident tone suppresses error detection while maintaining reliance, showing that delivery styles can override correctness. These results highlight how CoT explanations can simultaneously clarify and mislead, underscoring the need for NLP systems to provide explanations that encourage scrutiny and critical thinking rather than blind trust. All code will be released publicly.
Figures & tables
Figure 1: The trust-scrutiny gap. (a) What we study : when reading a model’s judgment on subjective topics, users often bypass flawed reasoning when they agree with the verdict. (b) Framework to decompose trust : we measure the three signals (error detection, agreement, and self-reported trust) in response to controlled perturbations of the reasoning chain.
Figure 2: Stimulus Construction Pipeline. (a) Participants evaluate multimodal moral judgments generated by VLMs. Each trial presents an image–scenario pair, a model-produced CoT, and a binary moral judgment. (b) We introduce three recurrent failure patterns as perturbations of otherwise clean chains: omissions, contradictions, and hallucinations. These respectively capture incompleteness, inconsistency, and ungrounded invention in model reasoning. (c) Confidence tones in the reasoning chains are also manipulated with specific epistemic markers.
Perturbation Type
Detection (%)
Agreement (1–7)
Trust (1–7)
Clean
15.5 [9.6, 19.8]
4.21 [3.93, 4.49]
3.23 [2.99, 3.47]
Omission
39.4 [32.6, 46.6]
4.00 [3.77, 4.23]
2.86 [2.67, 3.06]
Contradiction
69.0 [62.6, 75.4]
3.52 [3.30, 3.73]
2.33 [2.15, 2.51]
Hallucination
66.3 [59.5, 72.6]
3.38 [3.14, 3.62]
2.22 [2.03, 2.41]
Table 1: Main outcomes across conditions. Strong agreement–trust coupling and clear error-type differences emerge, with omissions the hardest to detect. Color intensity encodes the relative magnitude of each metric: darker green = ↑ detection accuracy, darker blue = ↑ agreement, darker red = ↑ perceived trust.
Figure 3: Overall cross-measure correlations. Negative correlations indicate that higher agreement and trust coincide with reduced error detection.
Detection
Agreement
Trust
Coefficients ( β , vs. Clean / Neutral)
Omission
+1.33†
−0.22⋄
−0.37†
Contradiction
+2.59†
−0.70†
−0.90†
Hallucination
+2.47†
−0.83†
−1.01†
Confident Tone
−0.49∗
−0.08
−0.00
ANOVA test statistics
Table 2: Mixed-effects model coefficients and ANOVA results. Detection fit by logistic regression. Agreement and Trust by OLS on 1–7 ratings. ∗p<.05 , ⋄p<.01 , †p<.001 (Holm-corrected).
Flawed trials
Agreement (1–7)
Trust (1–7)
Error detected
3.30
1.99
Error missed
4.08
3.12
Clean reference
4.21
3.23
Missed error and high reliance (%)
Omission
27.5
Contradiction
14.5
Table 3: Missed errors sustain reliance. Top: on flawed trials, participants who flagged the error reported markedly lower agreement and trust than those who did not; the latter remain near the Clean baseline. Bottom: proportion of trials combining a missed error with high reliance (Agreement ≥5 or Trust ≥5 ), by error type. The detection split is measured, not assigned.
Tones
Detection (%)
Agreement (1–7)
Trust (1–7)
Hedged
50.4 [44.0, 56.8]
3.70 [3.48, 3.92]
2.56 [2.38, 2.75]
Neutral
47.0 [40.9, 53.2]
3.85 [3.62, 4.06]
2.70 [2.52, 2.88]
Confident
40.5 [34.6, 46.7]
3.78 [3.58, 3.98]
2.70 [2.53, 2.89]
Table 4: Confidence tone effects. Numbers show mean with 95% bootstrap CIs. Neutral denotes the baseline wording of each chain before tone edits.
Figure 4: Error detection rates by tones . Points show observed error detections, and vertical error bars indicate 95% confidence intervals computed separately for each tone × perturbation-type.
Figure 5: Model-side error detection across perturbation types. Clean bars show false-positive rate ( ↓ ); the other three show true-positive rate ( ↑ ) per error type.
Figure 6: In-the-wild error distribution by models. Stacked bar plots of reasoning chain perturbations across six VLMs.
Figure 7: Epistemic marker usage across models. Strengtheners (boosters) are far more common than weakeners (hedges), suggesting a systematic bias toward confident reasoning styles.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
ROSCOE diagnostic dimension
Perturbation
Missing-Step (Omission)
Self-Consist. (Contradiction)
Omission
+2.38[+2.04,+2.86]
+0.03[−0.20,+0.27]
Contradiction
+1.24[+1.08,+1.48]
+1.04[+0.84,+1.28]
Appendix
Table 5: Automatic diagnostic checks for perturbation validity. Each cell reports the reference-relative Clean → perturbed shift as Cohen’s d with 95% paired-bootstrap CIs. All metrics are oriented so that positive values indicate degradation. Bold marks the diagnostic most directly associated with each perturbation; each perturbation moves its own diagnostic while leaving the other essentially unchanged.
Final Judgment
Error Type
Detection (%)
Agreement (1-7)
Trust (1-7)
Correct
Clean
4
5.86
4.45
Omission
31
5.15
3.72
Contradiction
55
4.62
3.10
Hallucination
51
4.68
3.02
Incorrect
Clean
27
2.54
1.97
Omission
48
2.81
1.98
Appendix
Table 6: Main outcomes across conditions (error type × correctness). We see strong agreement-trust coupling and clear error-type differences, with omissions the hardest to detect. Color intensity encodes a relative magnitude of each metric: darker green = ↑ detection accuracy, darker blue = ↑ agreement, darker red = ↑ perceived trust.
Figure 8: Example stimuli used in our user study. Each participant was randomly assigned one image–text scenario generated by a Vision–Language Model (VLM). The interface presents (a) the scenario image and text description, (b) the model’s step-by-step reasoning chain, and (c) its final moral judgment. Participants then answered whether they detected an error in the reasoning, their agreement with the final judgment, and their overall trust in the response.
Figure 9: Overview of the study description and task instructions shown to participants. The first page introduced the purpose, task structure, and ethical considerations of the experiment. Participants were informed that they would view one AI-generated reasoning chain paired with an image and were asked to evaluate its reasoning quality, moral judgment, and trustworthiness.
Generation prompt
Judged error rate (%)
Original
23.0
Paraphrased
26.0
Strict-grounded
19.0
Appendix
Table 7: Generation-prompt sensitivity. Proportion of GPT-5.2 chains judged to contain at least one taxonomy error, across three elicitation prompts ( n=50 scenarios per prompt).
Clean
True-positive rate (%)
Judge prompt
FP (%)
Omi
Con
Hal
Generic
20
31
75
82
Taxonomy
20
79
98
99
Omission-targeted
50
90
82
80
Appendix
Table 8: Judge-prompt sensitivity. False-positive rate on Clean chains ( ↓ ) and true-positive rate per error type ( ↑ ), evaluated against human-validated stimulus labels. Omission detection varies by 59 points across prompts; contradiction and hallucination detection remain high under the generic and taxonomy prompts.