Large Language Models (LLMs) increasingly ship with explicit "thinking modes", yet their counterpart, "no-thinking", has received far less attention. We study LLMs' no-thinking behavior along two axes. a. How to measure no-thinking? Prior work typically defines no-thinking through proxies such as a disabled thinking mode or the absence of long traces. These proxies are unreliable: disabled thinking modes may still emit reasoning, while long traces may contain filler rather than genuine inference. We instead normalize each response into a pre-answer trace and final answer, and evaluate it at three levels: (i) Empty-Thinking Rate for strict answer-only compliance; (ii) instruction-aware Question-Pre-answer Relevance for similarity between the question and pre-answer trace; and (iii) LLM-as-judge Explicit Inference Rate for visible explicit inference. Together, these metrics distinguish answer-only output, relevant but non-inferential text, and explicit inference. b. How does no-thinking vary across tasks and models? We evaluate six prompting interventions on six LLMs across Boolean, multiple-choice, and open-ended questions. We find that explicit no-think controls cannot reliably eliminate visible inference. Models instead exhibit "Thinking Inertia": explicit inference persists even under strict controls and becomes more prevalent as the answer space opens. Accuracy remains stable on Boolean and multiple-choice tasks, whereas open-ended tasks reveal a trade-off between answer-only compliance and task accuracy. Rewriting the same questions across answer spaces shows that supplying candidate answers makes answer-only responses easier to produce. These findings establish no-thinking as a non-trivial capability: stopping explicit reasoning cannot be assumed from model settings or instructions alone and deserves systematic evaluation alongside reasoning ability.
Figures & tables
Figure 1 : No-thinking controls do not guarantee response-level no-thinking. (Left) On a representative MMLU question, DeepSeek-V4-Flash and Qwen3-235B-A22B are asked for only the final option, yet expose question-relevant pre-answer text. Appendix Figure E.1 confirms the instruction’s clear answer-only reading. (Right) Under native no-thinking, green ETR and orange visible-length bars across the six benchmarks show Boolean tasks closer to answer-only than multiple-choice and open-ended tasks.
Figure 2 : Surface proxies fail to identify response-level no-thinking. (a) Three common proxies: 1. prompt-side control, 2. special reasoning tokens ( e.g., <think> ), and 3. generation token length. (b) On the same algebra question, disabling the thinking mode, omitting special thinking tokens, or producing a short output can each leave a visible inferential step before the final answer. These cases motivate separately evaluating strict answer-only compliance, pre-answer relevance, and visible explicit inference.
Figure 3 : Response-centric metrics for measuring no-thinking. Top: the LLM inference process. A question Q is wrapped by an intervention S(⋅) to form the model input S(Q) , which the model π maps to a response decomposed into pre-answer text T and final answer A . We evaluate the visible response through three levels: strict answer-only compliance (ETR), question–pre-answer relevance (QRel.), and LLM judge explicit inference (EIR); accuracy independently evaluates A . Bottom: six representative response cases to the same question “ 42+79=? ” vary in length, relevance, and inferential content while sharing the same final answer. Case 1 is answer-only (empty T ), whereas Cases 2–6 use the non-empty- T labels from the mutually exclusive rubric in Table 1 : Cases 2 and 4 are explicit reasoning, Cases 3 and 5 are generic/off-topic, and Case 6 is relevant non-inferential. Table 1 additionally defines paraphrase-only and unclear; representative real outputs are provided in Appendix F.1 .
Table 1: Blinded LLM-judge instruction and rubric for non-empty pre-answer text T . Empty T is handled separately and counts as no visible explicit inference.
Table 2: Six-mode question-irrelevant prompt intervention spectrum. Modes shown in gray ( Think-On , CoT ) still elicit thinking behavior and serve as thinking-side references; the remaining modes ( Think-Off , Soft , Strict , Prefix ) form the no-think intervention space, ranging from the native no-thinking baseline to increasingly strong semantic suppression.
Table 3 : LLM baselines for evaluation.
Figure 4 : Think-on (M1), Think-off (M2), and Step-by-step (M3) across six models and three answer-space families. Rows show modes; the columns pair ETR bars with EIR points (left) and QRel. bars with visible pre-answer length for the four models with complete length traces (right). Gray dashed segments mark six-model means for ETR and QRel.; black triangles connected by dark lines mark the mean EIR; colored marks follow the model legend.
Figure 5 : On Open-ended tasks, strict answer-only prompting exposes a trade-off between answer-only compliance and task accuracy. Columns show Bool, MCQ, and Open-ended answer spaces; rows show Acc., ETR, QRel., and EIR. M2 is the native no-thinking baseline; M4 and M5 add increasingly strong natural-language constraints, while M6 forms a separate structural prefix-forcing branch. Colored lines show individual models; the black line and gray band show the mean and one standard deviation. For ETR, these summaries include only models with a disabled-thinking condition.
Table 4 : No-thinking behavior varies by domain even under the same four-choice answer interface. Following the official MMLU subject-area split [ 14 ] , we report Humanities, Social Sciences, and STEM and omit the heterogeneous Other category. The % column gives each area’s share of the full MMLU test split; for each model, the remaining columns report Acc., ETR, QRel., and EIR under M2.
Figure 6 : Changing only the answer space reshapes no-thinking behavior on the same questions. We rewrite matched numeric-answer items from MMLU and MMLU-Pro for Qwen3-32B under M2 from MCQ into Boolean or open-ended forms. The four panels report Acc., ETR, QRel., and EIR on the same answer-space variants; red values give the change from the MCQ form, computed before rounding.
Table 5 : Human and LLM-judge explicit-inference rates across the six sampled M2/M5 answer-space cells. All entries use the all-output denominator within each cell; the Human mean row averages the five annotator-specific rates.
Table 6 : Human relevance calibration on the unified 600-output sample. Each entry is an item-level Spearman correlation over all 600 outputs; empty T receives relevance score 1 and QRel. 0, and the final column reports the mean across annotators.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Split
Size used
BoolQ
Validation
3,270
StrategyQA
Test
687
MMLU
Test
14,042
MMLU-Pro
Test
12,032
GSM8K
Test
1,319
MATH
Test
5,000
Appendix
Supplementary Table B.1: Dataset sizes used in the appendix.
Model group
Thinking control
Decoding and evaluation scope
Qwen3-4B / Qwen3-32B
Local Qwen chat-template switch: enable_thinking=true for Mode 1 and false for Modes 2–6.
Temperature 0.0. Max tokens: 8,192 for all modes.
Qwen3-235B-A22B
API reasoning control: high effort for Mode 1 and no-reasoning setting for Modes 2–6.
Temperature 0.0. Max tokens: 8,192 for all modes.
DeepSeek-V4-Flash
Provider thinking interface: high effort for Mode 1 and disabled for Modes 2–6.
Temperature 0.0. Max tokens: 8,192 for all modes.
GPT-5.4
Reasoning effort: high for Mode 1 and none for Modes 2–6.
Temperature 0.0. Max tokens: 8,192 for all modes.
Gemini-3-Flash-Preview
Reasoning effort: high for Mode 1 and minimal for Modes 2–6.
Temperature 0.0. Max tokens: 8,192 for all modes. The minimal setting is treated as near-off rather than strict off.
Appendix
Supplementary Table B.2: Model interfaces and decoding settings.
Result block
Models and task scope
Appendix evidence supported
Local Qwen spectrum
Qwen3-4B and Qwen3-32B on the six-benchmark suite: BoolQ, StrategyQA, MMLU, MMLU-Pro, GSM8K, and MATH, under Modes 1–6.
Global local-model summary and Qwen scaling evidence in Figure G.1 and Table H.1 .
Frontier model results
Qwen3-235B-A22B, DeepSeek-V4-Flash, GPT-5.4, and Gemini-3-Flash-Preview on the same six-benchmark suite, using each provider’s public thinking-control interface.
Frontier native no-think and suppression/forcing results in Figures H.1 and G.2 , plus the Qwen table in Table H.1 .
Full MMLU domain runs
Qwen3-4B, Qwen3-32B, DeepSeek-V4-Flash, and Qwen3-235B-A22B on MMLU, grouped by the original MMLU supercategories and, where applicable, by individual subject.
Domain and subject-level analyses in Figure H.2 , Table I.1 , and Table I.2 .
Surface and candidate-visibility interventions
Qwen3-4B and Qwen3-32B on MMLU variants with choice shuffling, numeric labels, and boolean verification with a proposed answer shown.
Question-interface controls in Tables I.3 and I.4 .
Numeric-source open-generation intervention
Qwen3-4B and Qwen3-32B on matched numeric-source items, comparing MCQ, boolean verification, and open numeric generation.
Candidate-removal stress test in Table I.5 .
Appendix
Supplementary Table C.1: Evaluation scope for appendix result blocks.
M2: Think-Off
M5: Strict
Encoder
Bool
MCQ
Open
Bool
MCQ
Open
all-MiniLM-L6-v2
0.046
0.366
0.701
0.001
0.041
0.455
Qwen3-Embedding-4B
0.047
0.406
0.854
0.001
0.045
0.546
Appendix
Supplementary Table D.1: Full-corpus Level 2 comparison. Mean relevance scores use all outputs, with empty T assigned zero. Both encoders were applied to outputs from all six models, six benchmarks, and six modes; the table reports the central M2/M5 contrast.
M2: Think-Off
M5: Strict
Encoder
Bool
MCQ
Open
Bool
MCQ
Open
all-MiniLM-L6-v2
0.119
0.460
0.677
0.052
0.132
0.439
BGE-base-en-v1.5
0.284
0.424
0.759
0.154
0.262
0.518
E5-base-v2
0.263
0.431
0.890
0.251
0.365
0.638
Appendix
Supplementary Table D.2: Level 2 sensitivity to commonly used encoders on the balanced 3,600-output sample. Absolute cosine scales are encoder-specific; the relevant comparisons are the directions within each row.
M2: Think-Off
M5: Strict
Judge
Bool
MCQ
Open
Bool
MCQ
Open
GPT-5.5
5.15%
45.57%
99.69%
0.12%
25.39%
62.00%
Claude Opus 4.8
4.24%
40.73%
98.55%
0.10%
24.21%
61.02%
Appendix
Supplementary Table D.3: Cross-family replication of the main EIR result. Values are all-output explicit-inference rates; GPT-5.5 is the primary judge and Claude Opus 4.8 is the independent cross-family judge.
Mode
Answer space
Explicit reasoning
Paraphrase-only
M2
Bool
5.15%
0.03%
MCQ
45.57%
0.11%
Open
99.69%
0.02%
M5
Bool
0.12%
0.01%
MCQ
25.39%
0.03%
Open
62.00%
0.10%
Appendix
Supplementary Table D.4: Explicit inference and paraphrase-only rates under the primary GPT-5.5 judge. Both rates use all outputs in the corresponding setting as the denominator; empty T contributes to neither numerator.
Dimension
Group
N
Five-way
Explicit vs. other
Overall
All outputs
572,645
96.92%
97.30%
Non-empty T
391,010
95.48%
96.05%
Model
Qwen3-4B
218,078
95.78%
96.23%
Qwen3-32B
218,079
98.47%
98.77%
Qwen3-235B-A22B
34,122
96.27%
96.43%
DeepSeek-V4-Flash
34,122
94.21%
94.78%
Appendix
Supplementary Table D.5 : GPT-5.5–Opus 4.8 agreement across the complete evaluation. Each group pools the other experimental dimensions. Percentages are calculated over outputs for which both judges have a usable label.
Encoder
A1
A2
A3
A4
A5
Mean
Qwen3-Embedding-4B
0.828
0.831
0.830
0.833
0.823
0.829
all-MiniLM-L6-v2
0.794
0.811
0.814
0.827
0.798
0.809
Appendix
Supplementary Table D.6 : Per-annotator relevance–QRel. calibration on the unified 600-output sample. Values are item-level Spearman correlations.
Judge / agreement
A1
A2
A3
A4
A5
Mean
GPT-5.5 / Five-way
98.50%
97.83%
98.17%
97.33%
97.50%
97.87%
GPT-5.5 / Binary
98.67%
98.83%
99.00%
98.83%
98.67%
98.80%
Opus 4.8 / Five-way
96.00%
95.67%
95.83%
95.00%
95.17%
95.53%
Opus 4.8 / Binary
96.50%
97.00%
96.83%
96.67%
96.50%
96.70%
Appendix
Supplementary Table D.7 : Per-annotator agreement with the blinded Level 3 judges on all 600 outputs. Five-way agreement requires the same rubric label; binary agreement asks only whether the judge and annotator agree on whether the output is explicit reasoning.
Annotation variable
Krippendorff’s α
Five-way reasoning label (nominal)
0.873
Explicit reasoning vs. other (nominal)
0.948
Five-point relevance score (ordinal)
0.919
Appendix
Supplementary Table D.8 : Inter-annotator reliability on the 402 non-empty pre-answer texts.
Mode
Answer space
A1
A2
A3
A4
A5
Mean
M2
Bool
4.0%
4.0%
4.0%
4.0%
4.0%
4.0%
M2
MCQ
29.6%
29.6%
33.3%
29.6%
29.6%
30.4%
M2
Open
100.0%
100.0%
100.0%
100.0%
100.0%
100.0%
M5
Bool
0.0%
0.0%
0.0%
0.0%
0.0%
0.0%
M5
MCQ
11.8%
11.8%
11.8%
11.8%
11.8%
11.8%
M5
Open
57.1%
57.1%
57.1%
57.1%
57.1%
57.1%
Appendix
Supplementary Table D.9 : Per-annotator all-output explicit-inference rates for the six sampled M2/M5 cells. The main-text human row is the column-wise mean.
Model
Responses
Copy
Answer-first
Qwen3-4B
218,100
233 (0.107%)
23/199,186 (0.012%)
Qwen3-32B
218,100
31 (0.014%)
27/205,146 (0.013%)
Qwen3-235B-A22B
34,122
2 (0.006%)
97/33,930 (0.286%)
DeepSeek-V4-Flash
34,122
16 (0.047%)
1/33,935 (0.003%)
GPT-5.4
34,122
0 (0.000%)
0/34,104 (0.000%)
Gemini-3-Flash-Preview
34,122
19 (0.056%)
1/28,907 (0.003%)
Appendix
Supplementary Table E.1 : Audits of near-verbatim question repetition and response order. Copy is the conservative component-wise four-gram detector defined in this section. Answer-first requires the first box to occur after at most three lexical tokens and to be followed by at least three lexical tokens. Percentages for Copy use all successful responses; Answer-first percentages use boxed responses.
Supplementary Figure E.1 : Independent prompt-comprehension sanity check. The figure shows the complete visible exchange in which we ask GPT to interpret the answer-only instruction used in Figure 1 . The response explicitly distinguishes a worked solution from the requested direct-answer format.
Case
Regime
Model and mode
Record
∣T∣w
QRel.
1
Answer only
Qwen3-235B-A22B, M5
MMLU, mmlu-2663
0
0.000
2
Short length, high QRel.
DeepSeek-V4-Flash, M2
MMLU, mmlu-2663
23
0.896
3
Short length, low QRel.
Qwen3-235B-A22B, M2
MMLU, mmlu-2663
4
0.310
4
Long length, high QRel.
Gemini-3-Flash-Preview, M3
MMLU, mmlu-2663
144
0.831
5
Long length, low QRel.
Gemini-3-Flash-Preview, M1
StrategyQA, strategyqa-227
137
0.276
Appendix
Supplementary Table F.1 : Real experimental counterparts of the five response regimes in Figure 3 . ∣T∣w is the visible pre-answer word count. QRel. is computed by Qwen3-Embedding-4B between the question and complete visible T under the instruction in Section 3.1 ; empty T is assigned zero. All five final answers are correct.
MCQ form is correct but has ∣T∣=209 and fails ETR. Boolean verification of the same content returns direct yes / no answers for both a true and a false proposal, and both pass ETR.
Changing the answer interface can make the same underlying content much more compressible because the candidate answer is already visible.
Supplementary Figure G.1 : Global six-mode profiles for local Qwen models. The plotted values are the corresponding dataset-level aggregates; EIR is averaged over the three answer-space families for each model–mode cell.
Supplementary Figure G.2 : Dataset-level view of the no-thinking–performance trade-off. The left panel shows the accuracy change from native no-think (M2) to strict answer-only suppression (M5); the right panel shows the resulting M5 ETR. The plotted values are the corresponding dataset-level aggregates.
Supplementary Figure H.1 : Frontier native no-think heatmaps by answer family. The three panels report the same native M2 outputs at successive levels: strict answer-only compliance (ETR), instruction-aware pre-answer relevance (QRel.), and explicit visible inference (EIR). The staircase is visible in all three panels: Bool is easiest to compress, MCQ is intermediate, and Open retains question-relevant and inferential content across providers.
Mode
Dataset
Qwen3-4B
Qwen3-32B
Qwen3-235B
Acc.
ETR
QRel.
Acc.
ETR
QRel.
Acc.
ETR
QRel.
M1 Think-On
BoolQ
87.4
0.0
0.800
89.2
0.0
0.795
89.2
0.0
0.794
StrategyQA
72.2
0.0
0.821
81.5
0.0
0.815
83.3
0.0
0.816
MMLU
75.2
0.0
0.853
85.3
0.0
0.835
88.3
0.0
0.853
MMLU-Pro
53.4
0.0
0.861
65.9
0.0
0.845
82.0
0.0
0.857
GSM8K
87.0
0.0
0.808
93.1
0.0
0.801
94.2
0.0
0.823
Appendix
Supplementary Table H.1: Qwen scaling comparison across modes and datasets. Blue, orange, and green shading encode Acc., ETR, and QRel., respectively; the corresponding Level 3 EIR summary is reported in Table H.2 .
Model
Mode
Bool
MCQ
Open
Qwen3-4B
M2 Think-Off
1.26%
49.30%
99.86%
M5 Strict
0.00%
7.05%
97.78%
Qwen3-32B
M2 Think-Off
0.00%
37.11%
99.56%
M5 Strict
0.00%
0.72%
56.73%
Qwen3-235B
M2 Think-Off
12.15%
68.70%
99.70%
M5 Strict
0.00%
27.00%
98.90%
Appendix
Supplementary Table H.2: Qwen scaling of the Level 3 signal. EIR is the all-output explicit-inference rate from the GPT-5.5 judge; empty T contributes zero. The compact table reports the two central no-thinking settings while the preceding longtable gives the corresponding Acc., ETR, and QRel. values by dataset.
Supplementary Figure H.2 : MMLU domain heatmaps under the same multiple-choice answer interface. The left panel shows answer-only success under native no-think; the right panel shows the accuracy change caused by strict answer-only suppression. The full area-level control table is given in Table I.1 .
Model
MMLU area
n
M1
M2
M5
M6
Qwen3-4B
Humanities
4705
63.7 / 0.1 / 99.9
62.5 / 36.8 / 56.6
58.0 / 100.0 / 0.0
62.8 / 34.6 / 55.7
Social Sciences
3077
82.8 / 0.5 / 100.0
78.7 / 78.2 / 16.2
76.9 / 99.8 / 0.2
78.1 / 73.6 / 17.3
STEM
3153
81.9 / 0.0 / 100.0
83.0 / 34.3 / 60.2
73.2 / 88.8 / 11.2
83.1 / 33.3 / 59.9
Other
3107
78.4 / 0.3 / 100.0
74.9 / 68.1 / 17.8
72.9 / 98.5 / 1.5
75.0 / 63.9 / 17.2
Qwen3-32B
Humanities
4705
77.3 / 4.5 / 100.0
71.6 / 95.5 / 4.1
71.3 / 100.0 / 0.0
71.6 / 94.9 / 4.5
Social Sciences
3077
90.2 / 1.5 / 100.0
88.3 / 87.7 / 10.7
87.9 / 100.0 / 0.0
88.4 / 88.2 / 10.0
Appendix
Supplementary Table I.1 : Full MMLU area-level controls for the domain analysis. Each mode cell reports Acc./ETR/EIR.
Qwen3-4B
Qwen3-32B
MMLU area / subject
n
M1 Acc.
M2 Acc./ETR
M5 Acc./ETR
M6 Acc./ETR
M1 Acc.
M2 Acc./ETR
M5 Acc./ETR
M6 Acc./ETR
Humanities
4705
63.7
62.5/36.8
58.0/100.0
62.8/34.6
77.3
71.6/95.5
71.3/100.0
71.6/94.9
formal logic
126
81.7
73.0/0.8
58.7/100.0
78.6/0.8
94.4
85.7/7.9
65.9/100.0
85.7/7.9
high school european history
165
77.0
78.8/55.2
75.2/100.0
80.6/53.3
87.9
84.2/91.5
84.2/100.0
83.6/90.9
high school us history
204
86.8
82.8/72.1
82.4/100.0
83.8/64.2
93.1
94.6/96.6
94.1/100.0
94.1/93.1
high school world history
237
85.7
82.7/64.6
84.0/100.0
83.1/58.2
91.6
90.7/93.7
90.3/100.0
90.3/92.0
Appendix
Supplementary Table I.2 : Subject-level MMLU breakdown using the original MMLU supercategories. Each mode cell reports Acc./ETR; darker orange means higher ETR. Mode 1 reports Acc. only.
Model
Mode
Variant
Acc.
ETR
QRel.
Δ Acc.
Δ ETR
Δ QRel.
Cons.
Qwen3-4B
Mode 2
base
74.2
53.0
0.369
–
–
–
–
choice shuffle
72.3
52.7
0.372
-1.9
-0.3
+0.003
82.5
numeric labels
73.4
45.0
0.426
-0.8
-8.0
+0.057
89.9
Mode 5
base
68.9
97.0
0.023
–
–
–
–
choice shuffle
67.0
96.8
0.024
-1.9
-0.1
+0.001
80.7
numeric labels
68.5
96.8
0.024
-0.4
-0.1
+0.001
93.0
Appendix
Supplementary Table I.3 : MMLU surface perturbations with the answer set fixed. Deltas are relative to the base version; darker colored cells indicate larger absolute changes.
MCQ base
Boolean verification
Change
Model
Mode
Acc.
ETR
QRel.
Acc.
ETR
QRel.
Pos.
Neg.
Δ Acc.
Δ QRel.
Qwen3-4B
Mode 1
77.6
0.1
0.852
80.0
65.6
0.833
82.4
77.6
+2.4
-0.019
Mode 2
75.0
52.8
0.370
73.5
67.2
0.235
65.5
81.6
-1.5
-0.135
Mode 5
69.4
97.0
0.023
69.1
100.0
0.000
52.9
85.2
-0.3
-0.023
Mode 6
74.6
50.4
0.382
72.9
76.8
0.179
63.9
81.8
-1.7
-0.203
Qwen3-32B
Mode 1
86.2
3.1
0.835
84.6
94.8
0.822
88.4
80.9
-1.6
-0.013
Appendix
Supplementary Table I.4 : MMLU boolean verification with a proposed answer shown. Darker cells indicate larger metric values or larger changes from MCQ.
MCQ
Bool verification
Open numeric
Drop
Model
Mode
Acc.
ETR
QRel.
Acc.
ETR
QRel.
Acc.
ETR
QRel.
Parse
Exact
Open–MCQ
Qwen3-4B
Mode 1
66.8
0.0
0.832
82.0
13.5
0.847
40.6
0.0
0.841
99.6
38.3
-26.2
Mode 2
72.1
3.7
0.760
81.2
12.1
0.722
50.3
0.0
0.811
97.0
47.2
-21.8
Mode 5
53.8
63.4
0.281
60.4
99.4
0.005
48.6
13.4
0.678
98.0
45.8
-5.2
Mode 6
72.0
1.4
0.771
81.6
12.9
0.717
50.6
0.0
0.811
97.9
47.5
-21.4
Qwen3-32B
Mode 1
72.7
0.2
0.821
81.2
44.8
0.841
43.9
0.0
0.832
99.8
41.6
-28.9
Appendix
Supplementary Table I.5 : Matched numeric-source items under MCQ, boolean verification, and open generation. Darker cells indicate larger open-generation values or larger Open–MCQ accuracy drops.
Hybrid thinking enables LLMs to switch between reasoning and direct answering, offering a balance between efficiency and reasoning capability. Yet our experiments reveal that current hybrid thinking LLMs only achieve partial mode separation: reasoning behaviors often leak into the no-think mode. To understand and mitigate this, we analyze the factors influencing controllability and identify four that matter most: (1) larger data scale, (2) using think and no-think answers from different questions rather than the same question, (3) a moderate increase in no-think data number, and (4) a two-phase strategy that first trains reasoning ability and then applies hybrid think training. Building on these findings, we propose a practical recipe that, compared to standard training, can maintain accuracy in both modes while significantly reducing no-think output length (from 1085 to 585 on MATH500) and occurrences of reasoning-supportive tokens such as "wait" (from 5917 to 522 on MATH500). Our findings highlight the limitations of current hybrid thinking and offer directions for strengthening its controllability. The code is available at: https://github.com/SR-A-W/demystifying-hybrid-thinking
While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this explicit Thinking mode from direct answer generation (NoThinking mode) remain poorly understood. To deconstruct this cognitive process, we apply Top-K Sparse Autoencoders (SAEs) to the intermediate representations of DeepSeek-R1-Distill-Qwen-7B and examine the model's divergent behaviors across math-solving tasks of three distinct difficulty levels. Observationally, we identify a clear distinction in how the model functions under two reasoning modes: Thinking mode relies on sparse and high-intensity feature activations driving verbal deduction independent of problem complexity, whereas NoThinking mode exhibits an adaptive and diffuse pattern prioritizing symbolic manipulation. Causally, suppressing the three most active sparse features by Total Activation Volume reveals three principles: (i) reasoning and syntactic structure are tightly coupled, as interventions consistently degrade \LaTeX{} and boxed-solution formatting; (ii) Thinking responds to disruption with compensatory over-generation marked by increased metacognitive cues and repetitive, low-information continuations; and (iii) coherent CoT behavior depends on a fragile coordination among specialized features, yielding distinct failure modes under perturbation but a consistently impaired output structure.
Bo Cheng, Qiaolin Lu, Yi Chang +1
School of Artificial Intelligence, Jilin University · The Hong Kong Polytechnic University · School of Artificial Intelligence, Jilin University Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China +1
Large language models (LLMs) are increasingly integrated into high-stakes decision-making. Inspired by the theory of \emph{inattentional blindness} in human cognition, we investigate whether LLMs, trained on human-preferred corpora that embed attentional biases, exhibit a similar limitation: \emph{failing to attend to subtle yet important contextual cues under explicit task instructions}. To evaluate this, we introduce the task of \textbf{explicit-implicit reasoning} and present \textbf{MixRea}, a benchmark of 2,246 multiple-choice questions across 9 reasoning types with varying distributions of explicit and implicit information. Evaluation of 21 advanced LLMs shows that even the best-performing reasoning model (Gemini 2.5 Pro) achieves only 42.8% consistency, revealing widespread inattentional blindness. To mitigate this, we propose \textbf{Potential Relation Completion Prompting (PRCP)}, a prompting method that improves reasoning by recovering overlooked causal relations. Further analysis shows that this limitation persists across diverse multi-source reasoning tasks, highlighting the need for more cognitively aligned models.
Yuanqing Cai, Ziyi Huang, Minhao Liu +3
Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China