Thinking Leakage: A Causal Audit of NoThink Post-Training in Hybrid Reasoning Models
Organizations: College of Information Sciences and Technology Pennsylvania State University
Abstract
Post-training hybrid reasoning models in NoThink mode has attracted growing interest as a way to improve performance while keeping inference fast. However, these gains may draw on thinking behavior already accessible through the base model's Think mode. We formulate this thinking leakage in a causal mediation framework and audit its contribution using bidirectional interventions along a simple base-derived activation direction. Across three models and three post-training methods on competition math benchmarks, we find that leakage is real, causal, and substantial: behavioral and representational analyses reveal shifts toward Think, steering the base model along this direction reproduces most of the post-training accuracy gain, and counter-steering a checkpoint removes a substantial share of what it gains. Across nine aligned checkpoints with positive NoThink gains, the resulting leakage ratio ranges from 42% to 79%. These interventions support a substantial causal contribution of thinking leakage. Our findings show that a post-training method's apparent advantage can therefore reflect greater drift toward Think, obscuring whether it improves capability within NoThink or more effectively re-invokes existing Think behavior.
Figures & tables
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | GRPO | SFT | OPSD |
|---|---|---|---|
| General | |||
| Backbones | Qwen3-8B, Qwen3-4B, MiniCPM4.1-8B | ||
| Training data | DAPO-Math | OpenThoughts-Math | OpenThoughts-Math |
| Data | |||
| Max. prompt length | 2048 | 2048 | 2048 |
| Max. response length | 8192 | 16000 | 1024 |
| Quantity | Checkpoint | True | Drawn | Abs. err. | Rel. err. |
|---|---|---|---|---|---|
| base | |||||
| GRPO@200 | |||||
| GRPO@500 | |||||
| rotation vs. base | GRPO@200 | ||||
| GRPO@500 | |||||
| base |
| Corpus | Ratio | ||
|---|---|---|---|
| OpenThoughts-Math(SFT, OPSD) | |||
| Qwen3-8B Think | |||
| Qwen3-4B Think | |||
| MiniCPM4.1-8B Think |
| Model | Configuration | Acc. | Density | Trunc. | Loop | Len. |
|---|---|---|---|---|---|---|
| ( ) | (%) | (%) | (tok) | |||
| Qwen3-8B | base, NoThink | 0.178 | 3.08 | 3.1 | 7.9 | 4,268 |
| base, Think | 0.557 | 18.30 | 12.4 | 1.1 | 17,782 | |
| GRPO@500 (best ckpt) | 0.472 | 11.85 | 1.6 | 2.2 | 9,293 | |
| steering, at | ||||||
| 0.231 | 5.57 | 1.3 | 6.0 | 4,940 | ||
| Model | Configuration | Acc. | Density | Trunc. | Loop | Len. |
|---|---|---|---|---|---|---|
| ( ) | (%) | (%) | (tok) | |||
| Qwen3-8B | base, NoThink | 0.178 | 3.08 | 3.1 | 7.9 | 4,268 |
| steer along | 0.335 | 10.26 | 0.8 | 2.9 | 6,850 | |
| steer along | 0.159 | 3.10 | 0.1 | 1.9 | 2,057 | |
| Qwen3-4B | base, NoThink | 0.160 | 3.23 | 2.3 | 7.8 | 4,069 |
| steer along | 0.299 | 11.62 | 0.3 | 2.9 | 6,570 |
| Injection layer | Acc. | Density | Trunc. | Loop | Len. |
|---|---|---|---|---|---|
| ( ) | (%) | (%) | (tok) | ||
| base, NoThink | 0.178 | 3.08 | 3.1 | 7.9 | 4,268 |
| 0.254 | 4.25 | 2.1 | 5.6 | 6,527 | |
| 0.233 | 4.27 | 1.3 | 7.3 | 5,358 | |
| † | 0.335 | 10.26 | 0.8 | 2.9 | 6,850 |
| 0.224 | 6.68 | 0.5 | 4.2 | 4,260 |
| Configuration | Acc. | Density | Trunc. | Loop | Len. |
|---|---|---|---|---|---|
| ( ) | (%) | (%) | (tok) | ||
| base, NoThink | 0.178 | 3.08 | 3.1 | 7.9 | 4,268 |
| base, Think | 0.557 | 18.30 | 12.4 | 1.1 | 17,782 |
| GRPO@500 | 0.472 | 11.85 | 1.6 | 2.2 | 9,293 |
| steering, | 0.545 | 16.42 | 1.3 | 1.8 | 12,179 |
| steering, | 0.553 | 18.31 | 3.8 | 1.0 | 15,033 |
| Accuracy | Density | Loop (%) | Len. (tok) | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | Checkpoint | bare | c.-s. | bare | c.-s. | bare | c.-s. | bare | c.-s. | |
| Qwen3-8B | GRPO@500 | +0.95 | 0.472 | 0.349 | 11.85 | 6.51 | 2.2 | 4.5 | 9,293 | 6,991 |
| GRPO@200 | +0.95 | 0.447 | 0.315 | 9.47 | 4.55 | 4.3 | 13.9 | 12,583 | 9,936 | |
| SFT@50 | +0.91 | 0.417 | 0.267 | 13.91 | 9.72 | 2.4 | 14.2 | 13,745 | 12,337 | |
| OPSD@25 | +0.77 | 0.331 | 0.210 | 6.59 | 2.70 | 4.8 | 28.2 | 13,847 | 9,244 | |
| OPSD@100 | +0.46 | 0.150 | 0.131 | 2.33 | 1.90 | 23.3 | 28.0 | 8,613 | 6,088 | |
| Model | Configuration | Acc. | Density | Trunc. | Loop | Len. |
|---|---|---|---|---|---|---|
| ( ) | (%) | (%) | (tok) | |||
| Qwen3-8B | GRPO@500 | 0.472 | 11.85 | 1.6 | 2.2 | 9,293 |
| counter-steering | ||||||
| 0.349 | 6.51 | 1.2 | 4.5 | 6,991 | ||
| 0.459 | 12.06 | 0.5 | 2.1 | 9,090 | ||
| random | 0.455 | 12.05 | 0.9 | 1.9 | 9,325 | |
| Accuracy | Density | Loop (%) | Len. (tok) | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Checkpoint | ||||||||
| Qwen3-8B | GRPO@500 | 0.349 | 0.212 | 6.51 | 3.43 | 4.5 | 13.3 | 6,991 | 6,078 |
| GRPO@200 | 0.315 | 0.191 | 4.55 | 2.40 | 13.9 | 22.6 | 9,936 | 7,696 | |
| SFT@50 | 0.267 | 0.131 | 9.72 | 4.79 | 14.2 | 30.1 | 12,337 | 9,587 | |
| OPSD@25 | 0.210 | 0.087 | 2.70 | 1.21 | 28.2 | 59.6 | 9,244 | 6,154 | |
| OPSD@100 | 0.131 | 0.076 | 1.90 | 1.53 | 28.0 | 41.4 | 6,088 | 5,353 | |
| Configuration | Acc. | Density | Trunc. | Loop | Len. |
|---|---|---|---|---|---|
| ( ) | (%) | (%) | (tok) | ||
| GRPO@500 | 0.472 | 11.85 | 1.6 | 2.2 | 9,293 |
| 0.349 | 6.51 | 1.2 | 4.5 | 6,991 | |
| 0.325 | 6.18 | 1.6 | 5.6 | 6,571 | |
| GRPO@200 | 0.447 | 9.47 | 6.2 | 4.3 | 12,583 |
| 0.315 | 4.55 | 5.2 | 13.9 | 9,936 |
| Configuration | Acc. | Density | Trunc. | Loop | Len. |
|---|---|---|---|---|---|
| ( ) | (%) | (%) | (tok) | ||
| base, NoThink | 0.173 | 3.36 | 3.5 | 15.8 | 6,202 |
| post-trained checkpoint | 0.304 | 7.22 | 4.9 | 14.7 | 10,391 |
| + counter-steering | 0.226 | 4.77 | 3.7 | 21.7 | 8,963 |
| + projection clamp | 0.240 | 5.84 | 3.6 | 24.5 | 9,220 |
| Model | Configuration | Acc. | Density | Trunc. | Loop | Len. |
|---|---|---|---|---|---|---|
| ( ) | (%) | (%) | (tok) | |||
| Qwen3-8B | base, NoThink | 0.178 | 3.08 | 3.1 | 7.9 | 4,268 |
| 0.151 | 1.99 | 1.3 | 8.6 | 3,322 | ||
| 0.111 | 1.33 | 1.9 | 9.5 | 3,233 | ||
| Qwen3-4B | base, NoThink | 0.160 | 3.24 | 2.3 | 7.8 | 4,069 |
| 0.129 | 2.02 | 1.0 | 10.4 | 3,270 |
| via the base sweep | via counter-steering | ||||||
|---|---|---|---|---|---|---|---|
| Model | Checkpoint | ||||||
| Qwen3-8B | GRPO@500 | +0.095 | +0.199 | +0.123 | +0.171 | +0.028 | 0.418 |
| GRPO@200 | +0.084 | +0.185 | +0.132 | +0.136 | +0.049 | 0.492 | |
| SFT@50 | +0.132 | +0.107 | +0.150 | +0.089 | +0.018 | 0.627 | |
| OPSD@25 | +0.112 | +0.041 | +0.121 | +0.032 | +0.009 | 0.789 | |
| OPSD@100 | +0.070 | -0.098 | +0.019 | -0.047 | -0.051 | -0.667 | |
| Model | Checkpoint | ||||
|---|---|---|---|---|---|
| Qwen3-8B | GRPO@500 | +0.772 | +0.766 | +0.836 | +0.677 |
| GRPO@200 | +0.792 | +0.806 | +0.852 | ||
| SFT@50 | +0.754 | +0.752 | +0.882 | ||
| OPSD@25 | +0.572 | +0.643 | +0.677 | ||
| OPSD@100 | +0.341 | +0.510 | +0.349 | ||
| Qwen3-4B | GRPO@450 | +0.741 | +0.820 | +0.874 | +0.613 |