Visible Chain-of-Thought (CoT) is often treated as a broadly useful reasoning instruction, yet analytics code generation combines natural-language ambiguity, schema grounding, target-language constraints, and model-specific inference behavior. Because the same analytics request can be expressed in two distinct target languages-SQL and Python (pandas)-this setting provides a natural test of a common but under-examined assumption: that visible reasoning is more effective when its representation matches the requested target, as in "think in SQL" or "think in Python." Together with generic instructions such as "think step-by-step," such recommendations remain insufficiently evaluated under controlled, execution-based comparisons. We study a query matched SQL-pandas benchmark that crosses persona phrasing, target language, visible-CoT format, control prefixes, direct generation, and internal-reasoning configurations. The results do not support either a universal accuracy advantage from visible CoT or a consistent benefit from matching the reasoning representation to the target language. Instead, the effects depend on the persona, target, model configuration, and internal-reasoning setting. The control ablations further distinguish effects of reasoning content from those of prompt format. These findings indicate that reasoning strategies should be selected jointly for the model, persona, target, and internal-reasoning configuration rather than adopted as universal defaults. More broadly, the study provides a controlled framework for identifying when visible reasoning improves executable generation, when it primarily perturbs model behavior, and when the internal-reasoning configuration is the more consequential factor.
Table 1: Six completed configurations used in the experiments.
Inference family
Tests
BH discoveries
Mode-direct
252
0
Target matching
36
0
Internal-reasoning toggle
96
16
Persona interactions
756
1
Table 2: Multiplicity-controlled inference families. Primary condition tests are the first row; other rows address distinct predeclared questions.
Model
Target
CoT
Control
Direct
NL
SQL
Python
Pseudocode
mean
Filler
Scrambled
Wrong
GLM4
SQL
23.68
28.29
28.29
25.99
29.28
27.96
24.01
25.66
27.30
GLM4
Python
15.13
19.74
18.42
17.11
17.76
18.26
16.45
15.13
15.79
Qwen3
SQL
29.93
33.55
30.26
33.55
30.59
31.99
31.58
34.54
30.26
Qwen3
Python
21.71
22.04
21.38
20.39
20.72
21.13
23.03
25.33
24.01
Gemma4-off
SQL
36.18
37.17
37.83
37.50
36.51
37.25
37.17
36.84
36.51
Table 3: Original-phrasing exact execution accuracy (%) along with the mean of the four explicit-CoT formats. In each row, bold marks the best direct-or-explicit-CoT value and the best control value.
Ablation
Δ NL
Δ direct
BH outcome (NL; direct)
Length filler
-0.36
0.90
not detected; not detected
Scrambled
0.16
1.43
not detected; higher
Wrong content
-0.49
0.77
not detected; not detected
Table 4: Original-phrasing control contrasts, pooled over six configurations, targets, and matched queries. Δ NL is control minus NL-CoT accuracy; Δ direct is control minus direct-generation accuracy. Changes are percentage points. “Not detected” denotes failure to reject after BH correction; “higher” is an observed corrected contrast.
Figure 1: Exact eight-step compliance by model. m -CoT averages the four explicit-CoT formats.
Figure 2: Original-phrasing internal-reasoning configuration effects within family-adaptive signed token and latency buckets. Each point is the mean paired enabled-minus-disabled accuracy change in its resource bucket.
Figure 3: Original-phrasing accuracy within model-adaptive total-token bins. Each point is the mean accuracy in a bin.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Evaluation path
Runtime configuration
GLM4; Qwen3
vLLM on AWS g6.12xlarge (4 L4 GPUs, 48 vCPUs, 192 GiB); BF16; 32,768-token context; no internal-reasoning setting.
Gemma4-off/on
Same AWS/vLLM host and checkpoint; BF16; 32,768-token context; request-level internal reasoning disabled/enabled.
Gemini Flash-0/24k
Provider API, issued from Mac M4 Max with 64 GB unified memory; provider-side serving hardware not observed.
All configurations
Temperature 0; fixed experiment seed; bounded repair and execution; deterministic reference normalization.
Appendix
Table 5: Serving and execution configuration. Full model identifiers and the compact reference names used elsewhere are listed in the main model table.
Model
Rows
Queries
Correct (%)
GLM4
14,592
304
37.38
Qwen3
14,592
304
41.18
Gemma4-off
14,592
304
43.54
Gemma4-on
14,592
304
46.57
Gemini Flash-0
14,592
304
46.15
Gemini Flash-24k
14,592
304
45.96
Appendix
Table 6: Completed common-frame inventory.
Model
Persona
Target
Direct
NL CoT
SQL CoT
Python CoT
Pseudo code CoT
Length filler
Scrambled
Wrong content
GLM4
expert
SQL
88.16
87.83
86.18
87.83
88.82
90.79
86.84
87.83
GLM4
expert
PYTHON
60.20
54.61
53.95
55.59
57.89
60.20
60.86
55.92
GLM4
novice
SQL
19.08
21.38
18.42
21.05
18.42
20.72
26.32
20.72
GLM4
novice
PYTHON
14.14
14.80
13.49
15.13
15.13
14.14
16.45
13.16
GLM4
original
SQL
23.68
28.29
28.29
25.99
29.28
24.01
25.66
27.30
GLM4
original
PYTHON
15.13
19.74
18.42
17.11
17.76
16.45
15.13
15.79
Appendix
Table 7: Full persona-wise exact execution accuracy (%). The best condition in each model-persona-target row is bold.
Term family
pBH
Persona
<10−6
Target
<10−26
Model configuration
<10−11
Mode
0.620
Mode × persona
0.00033
Mode × target
0.626
Appendix
Table 8: Complementary robust-GEE omnibus tests. Paired exact contrasts remain the primary inference.
Phrasing
Target
Condition
Direct
Condition
Δ pp
pBH
Original
SQL
Pseudocode CoT
33.39
35.31
1.92
0.168
Original
SQL
NL CoT
33.39
35.14
1.75
0.168
Original
SQL
Scrambled
33.39
35.36
1.97
0.196
Novice
SQL
Scrambled
27.47
32.62
5.15
0.00052
Novice
Python
Scrambled
18.86
21.11
2.25
0.111
Novice
SQL
NL CoT
27.47
28.84
1.37
0.439
Appendix
Table 9: Top pooled condition-versus-direct results in the focused Original and Novice families. Pooled contrasts average paired deltas across the six configurations within each query before testing.
Phrasing
Target
Contrast
Δ pp
95% CI
Raw p
pBH
Novice
SQL
Scrambled − direct
5.15
[2.80,7.62]
3.72×10−5
0.00045
Novice
SQL
Scrambled − length filler
5.04
[3.07,7.07]
9.14×10−7
0.000011
Novice
SQL
Scrambled − wrong content
3.45
[1.54,5.48]
0.00063
0.0038
Novice
SQL
Scrambled − explicit-CoT mean
4.39
[2.43,6.41]
1.61×10−5
0.00019
Novice
Python
Scrambled − length filler
2.47
[0.77,4.22]
0.0046
0.0184
Novice
Python
Scrambled − wrong content
2.58
[0.77,4.50]
0.0083
0.0250
Appendix
Table 10: Corrected focused control contrasts. BH is applied within three separate 12-test families: controls versus direct, between controls, and controls versus the four-format explicit-CoT mean.
Persona
Difficulty
Δ vs NL CoT
qBH
Δ vs direct
qBH
Original
Simple
0.00
1.0000
1.85
0.1646
Original
Moderate
-0.62
0.8076
0.23
0.9469
Original
Challenging
1.31
0.4539
2.34
0.1646
Novice
Simple
2.85
0.0183
4.01
0.0026
Novice
Moderate
1.17
0.4539
3.04
0.0183
Novice
Challenging
3.84
0.0077
4.12
0.0034
Appendix
Table 11: Scrambled-minus-comparator accuracy contrasts by persona and difficulty, pooled across the six configurations and both targets. Deltas are percentage points; qBH values use the focused difficulty family.
Figure 4: Complete same-family internal-reasoning effects (enabled minus disabled, or budget 24,576 minus 0).
Reasoning condition
Accuracy (%)
Mean tokens
Mean latency (s)
Δ pp
Extra tokens
Extra latency (s)
Direct
42.94
2,153
37.1
0.00
0
0.0
NL CoT
43.41
3,095
69.3
0.48
942
32.2
SQL CoT
43.52
3,077
75.4
0.58
924
38.4
Python CoT
42.82
3,069
76.0
-0.12
916
38.9
Pseudocode CoT
43.51
3,194
81.3
0.58
1,041
44.2
Appendix
Table 12: Pooled cost and accuracy for each explicit CoT format versus direct generation over the six configurations, three personas, and two targets. Bold marks the highest accuracy and the smallest total-token and latency values.
Condition family
Accuracy (%)
Tokens
Latency (s)
Δ pp
Extra tokens
Extra latency (s)
Dominated (T/L, %)
Direct
42.94
2,153
37.1
0.00
0
0.0
0.0 / 0.0
Explicit CoT mean
43.32
3,109
75.5
0.38
956
38.4
41.7 / 38.9
Appendix
Table 13: Pooled descriptive efficiency summary over six configurations, three personas, and two targets. Direct-dominated percentages count matched cells where direct generation is at least as accurate and no more costly, with one strict advantage.
Difficulty
Condition family
Accuracy (%)
Mean tokens
Mean latency (s)
Δ pp vs direct
Simple
Direct
47.25
1,943
31.9
0.00
Simple
Explicit CoT mean
47.69
2,834
61.8
0.44
Moderate
Direct
42.52
2,128
34.2
0.00
Moderate
Explicit CoT mean
43.12
3,079
73.1
0.60
Challenging
Direct
38.20
2,439
46.9
0.00
Challenging
Explicit CoT mean
38.25
3,478
95.0
0.05
Appendix
Table 14: Difficulty-stratified efficiency means pooled across model, persona, and target cells. Difficulty raises both baseline and prompted resource use; the table should not be interpreted as a causal difficulty effect. From simple to challenging, direct generation adds approximately 496 tokens and 15.0 seconds, while explicit CoT adds 644 tokens and 33.2 seconds.
Figure 5: Accuracy within model-adaptive token bins, stratified by model, persona, target, difficulty, and reasoning condition. Original, Novice, and Expert are shown; point size encodes difficulty and the y-axis is tightened within each persona row block. Bucket widths target roughly 10% of the span and widen when necessary to keep high-range panels readable.
Figure 6: Accuracy within model-adaptive latency bins, stratified by model, persona, target, difficulty, and reasoning condition. Original, Novice, and Expert are shown; point size encodes difficulty and the y-axis is tightened within each persona row block. Bucket widths are chosen from roughly 10% of the observed span and rounded to the nearest standard latency bin.
Figure 7: Internal-reasoning configuration effects in signed token and latency buckets for Original, Novice, and Expert phrasing. Color encodes the reasoning condition, marker shape target, and marker size difficulty.
Model / CoT
Accuracy
Tokens
Latency (s)
GLM4 direct
36.73
1,537
20.5
GLM4 NL CoT
37.77
2,222
63.8
Qwen3 direct
40.13
1,549
13.6
Gemma4-off direct
43.70
1,873
28.8
Gemma4-on direct
46.05
3,743
144.7
Gemma4-on SQL CoT
47.15
6,056
318.3
Appendix
Table 15: Selected pooled efficiency points. Accuracy is percent; latency is locally measured client-observed wall-clock time and is not hardware-normalized.
Figure 8: Exact eight-step contract compliance for explicit-CoT and control conditions by model. Low compliance, especially under scrambled context, argues against interpreting lower completion length as evidence that the control was simply ignored.
Control
Correctness OR
qBH
Length filler
1.325
0.00048
Scrambled
0.995
0.94542
Wrong content
1.263
0.00264
Appendix
Table 16: Adjusted association between exact eight-step compliance and correctness within each control. Odds ratios are from query-clustered logistic GEE; BH correction is across the three control-specific tests.
Chain-of-thought (CoT) prompting is widely used as a reasoning aid and is often treated as a transparency mechanism. Yet behavioral gains under CoT do not imply that the model's internal computation causally depends on the emitted reasoning text, i.e. models may produce fluent rationales while routing decision-critical computation through latent pathways. We introduce a causal, layerwise audit of CoT faithfulness based on activation patching. Our key metric, the CoT Mediation Index (CMI), isolates CoT-specific causal influence by comparing performance degradation from patching CoT-token hidden states against matched control patches. Across multiple model families (Phi, Qwen, DialoGPT) and scales, we find that CoT-specific influence is typically depth-localized into narrow ''reasoning windows,'' and we identify bypass regimes where CMI is near-zero despite plausible CoT text. We further observe that models tuned explicitly for reasoning tend to exhibit stronger and more structured mediation than larger untuned counterparts, while Mixture-of-Experts models show more distributed mediation consistent with routing-based computation. Overall, our results show that CoT faithfulness varies substantially across models and tasks and cannot be inferred from behavior alone, motivating causal, layerwise audits when using CoT as a transparency signal.
Chain-of-thought (CoT) prompting assumes that generated reasoning reflects a model's internal computation. We show this assumption is wrong in a specific, measurable way: models internally detect their own reasoning errors but outwardly express confidence in them. A linear probe on hidden states predicts trace correctness with 0.95 AUROC -- from the very first reasoning step (0.79) -- while verbalized confidence for wrong traces is 4.55/5, nearly identical to correct ones (4.87/5). A text-surface classifier achieves only 0.59 on the same data, confirming a 0.20-point gap invisible in the generated text. This hidden error awareness holds across three model families (Qwen, Llama, Phi), 1.5B-72B parameters, and RL-trained reasoning models (DeepSeek-R1, 0.852 AUROC). The natural question is whether this signal can fix the errors it detects. It cannot. Four interventions -- activation steering, probe-guided best-of-N, self-correction, and activation patching -- all fail; patching destroys output coherence entirely. The signal is diagnostic, not causal: a readout of computation quality, not a lever to redirect it. This delineates a boundary for mechanistic interpretability: error representations during reasoning are fundamentally different from the factual knowledge representations that prior work has successfully edited.
Aojie Yuan, Zhiyuan Julian Su, Haiyue Zhang +2
University of Southern California, Los Angeles, CA, USA.
Reasoning-trained language models often spend more tokens on harder problems, but longer chains of thought do not show whether a model is merely computing for more steps or following a different internal trajectory. We study this distinction through hidden-state trajectories during chain-of-thought generation across competitive programming, mathematics, and Boolean satisfiability. Raw trajectory geometry is strongly shaped by generation length: longer generations mechanically alter path statistics, so difficulty-dependent comparisons are misleading without adjustment. After residualizing trajectory statistics on length, difficulty remains systematically coupled to corrected trajectory geometry across all domains studied. The clearest reasoning-specific separation appears in the code domain, where harder problems show more direct corrected trajectories and less heterogeneous local curvature in reasoning-trained models than in matched instruction-tuned baselines. Corrected difficulty-geometry coupling is weaker, but still present, in mathematics and Boolean satisfiability. Prompt-stage linear probes do not mirror the code-domain separation, and behavioral annotations show that stronger corrected coupling co-occurs with strategy shifts and uncertainty monitoring. Together, these findings establish length correction as a prerequisite for generation-time trajectory analysis and show that reasoning training can be associated with distinct corrected trajectory geometry, with the strength of the effect depending on the domain.
Anders Gjølbye, Lars Kai Hansen, Sanmi Koyejo
Technical University of Denmark · Stanford University