Thinking Inertia: LLMs Keep Thinking When Told Not To
Organizations: TSINGHUA UNIVERSITY · UNIVERSITY OF OXFORD · STANFORD UNIVERSITY
Abstract
Large Language Models (LLMs) increasingly ship with explicit "thinking modes", yet their counterpart, "no-thinking", has received far less attention. We study LLMs' no-thinking behavior along two axes. a. How to measure no-thinking? Prior work typically defines no-thinking through proxies such as a disabled thinking mode or the absence of long traces. These proxies are unreliable: disabled thinking modes may still emit reasoning, while long traces may contain filler rather than genuine inference. We instead normalize each response into a pre-answer trace and final answer, and evaluate it at three levels: (i) Empty-Thinking Rate for strict answer-only compliance; (ii) instruction-aware Question-Pre-answer Relevance for similarity between the question and pre-answer trace; and (iii) LLM-as-judge Explicit Inference Rate for visible explicit inference. Together, these metrics distinguish answer-only output, relevant but non-inferential text, and explicit inference. b. How does no-thinking vary across tasks and models? We evaluate six prompting interventions on six LLMs across Boolean, multiple-choice, and open-ended questions. We find that explicit no-think controls cannot reliably eliminate visible inference. Models instead exhibit "Thinking Inertia": explicit inference persists even under strict controls and becomes more prevalent as the answer space opens. Accuracy remains stable on Boolean and multiple-choice tasks, whereas open-ended tasks reveal a trade-off between answer-only compliance and task accuracy. Rewriting the same questions across answer spaces shows that supplying candidate answers makes answer-only responses easier to produce. These findings establish no-thinking as a non-trivial capability: stopping explicit reasoning cannot be assumed from model settings or instructions alone and deserves systematic evaluation alongside reasoning ability.
Figures & tables
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Split | Size used |
| BoolQ | Validation | 3,270 |
| StrategyQA | Test | 687 |
| MMLU | Test | 14,042 |
| MMLU-Pro | Test | 12,032 |
| GSM8K | Test | 1,319 |
| MATH | Test | 5,000 |
| Model group | Thinking control | Decoding and evaluation scope |
| Qwen3-4B / Qwen3-32B | Local Qwen chat-template switch: enable_thinking=true for Mode 1 and false for Modes 2–6. | Temperature 0.0. Max tokens: 8,192 for all modes. |
| Qwen3-235B-A22B | API reasoning control: high effort for Mode 1 and no-reasoning setting for Modes 2–6. | Temperature 0.0. Max tokens: 8,192 for all modes. |
| DeepSeek-V4-Flash | Provider thinking interface: high effort for Mode 1 and disabled for Modes 2–6. | Temperature 0.0. Max tokens: 8,192 for all modes. |
| GPT-5.4 | Reasoning effort: high for Mode 1 and none for Modes 2–6. | Temperature 0.0. Max tokens: 8,192 for all modes. |
| Gemini-3-Flash-Preview | Reasoning effort: high for Mode 1 and minimal for Modes 2–6. | Temperature 0.0. Max tokens: 8,192 for all modes. The minimal setting is treated as near-off rather than strict off. |
| Result block | Models and task scope | Appendix evidence supported |
| Local Qwen spectrum | Qwen3-4B and Qwen3-32B on the six-benchmark suite: BoolQ, StrategyQA, MMLU, MMLU-Pro, GSM8K, and MATH, under Modes 1–6. | Global local-model summary and Qwen scaling evidence in Figure G.1 and Table H.1 . |
| Frontier model results | Qwen3-235B-A22B, DeepSeek-V4-Flash, GPT-5.4, and Gemini-3-Flash-Preview on the same six-benchmark suite, using each provider’s public thinking-control interface. | Frontier native no-think and suppression/forcing results in Figures H.1 and G.2 , plus the Qwen table in Table H.1 . |
| Full MMLU domain runs | Qwen3-4B, Qwen3-32B, DeepSeek-V4-Flash, and Qwen3-235B-A22B on MMLU, grouped by the original MMLU supercategories and, where applicable, by individual subject. | Domain and subject-level analyses in Figure H.2 , Table I.1 , and Table I.2 . |
| Surface and candidate-visibility interventions | Qwen3-4B and Qwen3-32B on MMLU variants with choice shuffling, numeric labels, and boolean verification with a proposed answer shown. | Question-interface controls in Tables I.3 and I.4 . |
| Numeric-source open-generation intervention | Qwen3-4B and Qwen3-32B on matched numeric-source items, comparing MCQ, boolean verification, and open numeric generation. | Candidate-removal stress test in Table I.5 . |
| M2: Think-Off | M5: Strict | |||||
| Encoder | Bool | MCQ | Open | Bool | MCQ | Open |
| all-MiniLM-L6-v2 | 0.046 | 0.366 | 0.701 | 0.001 | 0.041 | 0.455 |
| Qwen3-Embedding-4B | 0.047 | 0.406 | 0.854 | 0.001 | 0.045 | 0.546 |
| M2: Think-Off | M5: Strict | |||||
| Encoder | Bool | MCQ | Open | Bool | MCQ | Open |
| all-MiniLM-L6-v2 | 0.119 | 0.460 | 0.677 | 0.052 | 0.132 | 0.439 |
| BGE-base-en-v1.5 | 0.284 | 0.424 | 0.759 | 0.154 | 0.262 | 0.518 |
| E5-base-v2 | 0.263 | 0.431 | 0.890 | 0.251 | 0.365 | 0.638 |
| M2: Think-Off | M5: Strict | |||||
| Judge | Bool | MCQ | Open | Bool | MCQ | Open |
| GPT-5.5 | 5.15% | 45.57% | 99.69% | 0.12% | 25.39% | 62.00% |
| Claude Opus 4.8 | 4.24% | 40.73% | 98.55% | 0.10% | 24.21% | 61.02% |
| Mode | Answer space | Explicit reasoning | Paraphrase-only |
| M2 | Bool | 5.15% | 0.03% |
| MCQ | 45.57% | 0.11% | |
| Open | 99.69% | 0.02% | |
| M5 | Bool | 0.12% | 0.01% |
| MCQ | 25.39% | 0.03% | |
| Open | 62.00% | 0.10% |
| Dimension | Group | Five-way | Explicit vs. other | |
| Overall | All outputs | 572,645 | 96.92% | 97.30% |
| Non-empty | 391,010 | 95.48% | 96.05% | |
| Model | Qwen3-4B | 218,078 | 95.78% | 96.23% |
| Qwen3-32B | 218,079 | 98.47% | 98.77% | |
| Qwen3-235B-A22B | 34,122 | 96.27% | 96.43% | |
| DeepSeek-V4-Flash | 34,122 | 94.21% | 94.78% |
| Encoder | A1 | A2 | A3 | A4 | A5 | Mean |
| Qwen3-Embedding-4B | 0.828 | 0.831 | 0.830 | 0.833 | 0.823 | 0.829 |
| all-MiniLM-L6-v2 | 0.794 | 0.811 | 0.814 | 0.827 | 0.798 | 0.809 |
| Judge / agreement | A1 | A2 | A3 | A4 | A5 | Mean |
| GPT-5.5 / Five-way | 98.50% | 97.83% | 98.17% | 97.33% | 97.50% | 97.87% |
| GPT-5.5 / Binary | 98.67% | 98.83% | 99.00% | 98.83% | 98.67% | 98.80% |
| Opus 4.8 / Five-way | 96.00% | 95.67% | 95.83% | 95.00% | 95.17% | 95.53% |
| Opus 4.8 / Binary | 96.50% | 97.00% | 96.83% | 96.67% | 96.50% | 96.70% |
| Annotation variable | Krippendorff’s |
| Five-way reasoning label (nominal) | 0.873 |
| Explicit reasoning vs. other (nominal) | 0.948 |
| Five-point relevance score (ordinal) | 0.919 |
| Mode | Answer space | A1 | A2 | A3 | A4 | A5 | Mean |
| M2 | Bool | 4.0% | 4.0% | 4.0% | 4.0% | 4.0% | 4.0% |
| M2 | MCQ | 29.6% | 29.6% | 33.3% | 29.6% | 29.6% | 30.4% |
| M2 | Open | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% |
| M5 | Bool | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
| M5 | MCQ | 11.8% | 11.8% | 11.8% | 11.8% | 11.8% | 11.8% |
| M5 | Open | 57.1% | 57.1% | 57.1% | 57.1% | 57.1% | 57.1% |
| Model | Responses | Copy | Answer-first |
| Qwen3-4B | 218,100 | 233 (0.107%) | 23/199,186 (0.012%) |
| Qwen3-32B | 218,100 | 31 (0.014%) | 27/205,146 (0.013%) |
| Qwen3-235B-A22B | 34,122 | 2 (0.006%) | 97/33,930 (0.286%) |
| DeepSeek-V4-Flash | 34,122 | 16 (0.047%) | 1/33,935 (0.003%) |
| GPT-5.4 | 34,122 | 0 (0.000%) | 0/34,104 (0.000%) |
| Gemini-3-Flash-Preview | 34,122 | 19 (0.056%) | 1/28,907 (0.003%) |
| Case | Regime | Model and mode | Record | QRel. | |
| 1 | Answer only | Qwen3-235B-A22B, M5 | MMLU, mmlu-2663 | 0 | 0.000 |
| 2 | Short length, high QRel. | DeepSeek-V4-Flash, M2 | MMLU, mmlu-2663 | 23 | 0.896 |
| 3 | Short length, low QRel. | Qwen3-235B-A22B, M2 | MMLU, mmlu-2663 | 4 | 0.310 |
| 4 | Long length, high QRel. | Gemini-3-Flash-Preview, M3 | MMLU, mmlu-2663 | 144 | 0.831 |
| 5 | Long length, low QRel. | Gemini-3-Flash-Preview, M1 | StrategyQA, strategyqa-227 | 137 | 0.276 |
| Case | Record setting | Observed pattern | Interpretation |
| Native residual information | Qwen3-4B, MATH, item math-4 , Mode 2 | Correct prediction , ETR fails, words. The visible response rewrites , sets , and solves . | Disabling native thinking does not force an open-ended math item into answer-only behavior; the ordinary answer channel still carries computation. |
| Re-elicitation | Qwen3-235B-A22B, BoolQ, item boolq-0 , Modes 2–3 | Mode 2 returns only \boxed{no} and passes ETR. Mode 3 remains correct but expands to words and fails ETR. | A step-by-step instruction can recover visible question-conditioned content even when the provider no-think interface remains active. |
| Compression cost | Qwen3-4B, MATH, item math-811 , Modes 2 and 5 | Mode 2 is correct: prediction , . Mode 5 gives a short direct response but predicts while the gold answer is . | Stronger direct-answer pressure can remove the computation needed to check a subtle constraint, here the word “smallest.” |
| Candidate visibility | Qwen3-32B, MMLU abstract algebra, item mmlu-39 , Mode 2 | MCQ form is correct but has and fails ETR. Boolean verification of the same content returns direct yes / no answers for both a true and a false proposal, and both pass ETR. | Changing the answer interface can make the same underlying content much more compressible because the candidate answer is already visible. |
| Mode | Dataset | Qwen3-4B | Qwen3-32B | Qwen3-235B | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Acc. | ETR | QRel. | Acc. | ETR | QRel. | Acc. | ETR | QRel. | ||
| M1 Think-On | BoolQ | 87.4 | 0.0 | 0.800 | 89.2 | 0.0 | 0.795 | 89.2 | 0.0 | 0.794 |
| StrategyQA | 72.2 | 0.0 | 0.821 | 81.5 | 0.0 | 0.815 | 83.3 | 0.0 | 0.816 | |
| MMLU | 75.2 | 0.0 | 0.853 | 85.3 | 0.0 | 0.835 | 88.3 | 0.0 | 0.853 | |
| MMLU-Pro | 53.4 | 0.0 | 0.861 | 65.9 | 0.0 | 0.845 | 82.0 | 0.0 | 0.857 | |
| GSM8K | 87.0 | 0.0 | 0.808 | 93.1 | 0.0 | 0.801 | 94.2 | 0.0 | 0.823 | |
| Model | Mode | Bool | MCQ | Open |
| Qwen3-4B | M2 Think-Off | 1.26% | 49.30% | 99.86% |
| M5 Strict | 0.00% | 7.05% | 97.78% | |
| Qwen3-32B | M2 Think-Off | 0.00% | 37.11% | 99.56% |
| M5 Strict | 0.00% | 0.72% | 56.73% | |
| Qwen3-235B | M2 Think-Off | 12.15% | 68.70% | 99.70% |
| M5 Strict | 0.00% | 27.00% | 98.90% |
| Model | MMLU area | M1 | M2 | M5 | M6 | |
| Qwen3-4B | Humanities | 4705 | 63.7 / 0.1 / 99.9 | 62.5 / 36.8 / 56.6 | 58.0 / 100.0 / 0.0 | 62.8 / 34.6 / 55.7 |
| Social Sciences | 3077 | 82.8 / 0.5 / 100.0 | 78.7 / 78.2 / 16.2 | 76.9 / 99.8 / 0.2 | 78.1 / 73.6 / 17.3 | |
| STEM | 3153 | 81.9 / 0.0 / 100.0 | 83.0 / 34.3 / 60.2 | 73.2 / 88.8 / 11.2 | 83.1 / 33.3 / 59.9 | |
| Other | 3107 | 78.4 / 0.3 / 100.0 | 74.9 / 68.1 / 17.8 | 72.9 / 98.5 / 1.5 | 75.0 / 63.9 / 17.2 | |
| Qwen3-32B | Humanities | 4705 | 77.3 / 4.5 / 100.0 | 71.6 / 95.5 / 4.1 | 71.3 / 100.0 / 0.0 | 71.6 / 94.9 / 4.5 |
| Social Sciences | 3077 | 90.2 / 1.5 / 100.0 | 88.3 / 87.7 / 10.7 | 87.9 / 100.0 / 0.0 | 88.4 / 88.2 / 10.0 |
| Qwen3-4B | Qwen3-32B | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| MMLU area / subject | M1 Acc. | M2 Acc./ETR | M5 Acc./ETR | M6 Acc./ETR | M1 Acc. | M2 Acc./ETR | M5 Acc./ETR | M6 Acc./ETR | |
| Humanities | 4705 | 63.7 | 62.5/36.8 | 58.0/100.0 | 62.8/34.6 | 77.3 | 71.6/95.5 | 71.3/100.0 | 71.6/94.9 |
| formal logic | 126 | 81.7 | 73.0/0.8 | 58.7/100.0 | 78.6/0.8 | 94.4 | 85.7/7.9 | 65.9/100.0 | 85.7/7.9 |
| high school european history | 165 | 77.0 | 78.8/55.2 | 75.2/100.0 | 80.6/53.3 | 87.9 | 84.2/91.5 | 84.2/100.0 | 83.6/90.9 |
| high school us history | 204 | 86.8 | 82.8/72.1 | 82.4/100.0 | 83.8/64.2 | 93.1 | 94.6/96.6 | 94.1/100.0 | 94.1/93.1 |
| high school world history | 237 | 85.7 | 82.7/64.6 | 84.0/100.0 | 83.1/58.2 | 91.6 | 90.7/93.7 | 90.3/100.0 | 90.3/92.0 |
| Model | Mode | Variant | Acc. | ETR | QRel. | Acc. | ETR | QRel. | Cons. |
| Qwen3-4B | Mode 2 | base | 74.2 | 53.0 | 0.369 | – | – | – | – |
| choice shuffle | 72.3 | 52.7 | 0.372 | -1.9 | -0.3 | +0.003 | 82.5 | ||
| numeric labels | 73.4 | 45.0 | 0.426 | -0.8 | -8.0 | +0.057 | 89.9 | ||
| Mode 5 | base | 68.9 | 97.0 | 0.023 | – | – | – | – | |
| choice shuffle | 67.0 | 96.8 | 0.024 | -1.9 | -0.1 | +0.001 | 80.7 | ||
| numeric labels | 68.5 | 96.8 | 0.024 | -0.4 | -0.1 | +0.001 | 93.0 |
| MCQ base | Boolean verification | Change | |||||||||
| Model | Mode | Acc. | ETR | QRel. | Acc. | ETR | QRel. | Pos. | Neg. | Acc. | QRel. |
| Qwen3-4B | Mode 1 | 77.6 | 0.1 | 0.852 | 80.0 | 65.6 | 0.833 | 82.4 | 77.6 | +2.4 | -0.019 |
| Mode 2 | 75.0 | 52.8 | 0.370 | 73.5 | 67.2 | 0.235 | 65.5 | 81.6 | -1.5 | -0.135 | |
| Mode 5 | 69.4 | 97.0 | 0.023 | 69.1 | 100.0 | 0.000 | 52.9 | 85.2 | -0.3 | -0.023 | |
| Mode 6 | 74.6 | 50.4 | 0.382 | 72.9 | 76.8 | 0.179 | 63.9 | 81.8 | -1.7 | -0.203 | |
| Qwen3-32B | Mode 1 | 86.2 | 3.1 | 0.835 | 84.6 | 94.8 | 0.822 | 88.4 | 80.9 | -1.6 | -0.013 |
| MCQ | Bool verification | Open numeric | Drop | ||||||||||
| Model | Mode | Acc. | ETR | QRel. | Acc. | ETR | QRel. | Acc. | ETR | QRel. | Parse | Exact | Open–MCQ |
| Qwen3-4B | Mode 1 | 66.8 | 0.0 | 0.832 | 82.0 | 13.5 | 0.847 | 40.6 | 0.0 | 0.841 | 99.6 | 38.3 | -26.2 |
| Mode 2 | 72.1 | 3.7 | 0.760 | 81.2 | 12.1 | 0.722 | 50.3 | 0.0 | 0.811 | 97.0 | 47.2 | -21.8 | |
| Mode 5 | 53.8 | 63.4 | 0.281 | 60.4 | 99.4 | 0.005 | 48.6 | 13.4 | 0.678 | 98.0 | 45.8 | -5.2 | |
| Mode 6 | 72.0 | 1.4 | 0.771 | 81.6 | 12.9 | 0.717 | 50.6 | 0.0 | 0.811 | 97.9 | 47.5 | -21.4 | |
| Qwen3-32B | Mode 1 | 72.7 | 0.2 | 0.821 | 81.2 | 44.8 | 0.841 | 43.9 | 0.0 | 0.832 | 99.8 | 41.6 | -28.9 |