Masked Self-Distillation: Internalizing the Chain-of-Thought in Language Models
Organizations: SCAI, Arizona State University
Abstract
Large Reasoning Models produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate traces dominate latency, memory usage, and serving cost, even though final answer correctness is not causally related to the trace correctness and the trace length is not a reliable indicator of the problem complexity. This raises an obvious question: can the computation expressed in these intermediate tokens be internalized into the parameters of a language model, enabling it to produce answers with much shorter intermediate traces? We propose masked self-distillation, a knowledge-distillation based post-training framework in which copies of the same model are instantiated as teacher and student, and the student model is trained to internalize all or part of the intermediate trace, thus becoming more efficient at inference. We vary the fraction of intermediate trace the student is trained to internalize, interpolating between full internalization and no internalization. We conduct controlled experiments on two reasoning domains: math and graph coloring. We use the masked self-distillation framework to post-train Qwen3-4B & 8B models. Our results demonstrate that this method can be used to improve task performance while increasing inference efficiency across various domains and model sizes. We systematically analyze whether improved efficiency gain in the post-trained models generalize to OOD problems. We find that masked self-distillation models generalize well for in-domain OOD problems, and the masked self-distillation training does not induce catastrophic forgetting in the student model on out-of-domain problems. Furthermore, our ablation study shows that supervised fine-tuning can train models to produce shorter traces, but at the cost of generalization, highlighting the importance of on-policy training in masked self-distillation.
Figures & tables
| - | - | - | GSM8K | |||||
| Masked | SFT | Reverse KL | SFT | Reverse KL | SFT | Reverse KL | SFT | Reverse KL |
| 100% | 88.7 | 11.5 | 11.4 | 10.5 | 26.2 | 10.6 | 0.0 | 87.0 |
| 68 | 67 | 68 | 66 | 68 | 115 | 32k ∗ | 299 | |
| 70% | 94.2 | 97.5 | 52.7 | 79.5 | 72.1 | 88.3 | 68.5 | 87.7 |
| 2631 | 3274 | 5899 | 6906 | 4681 | 5185 | 4866 | 549 | |
| 50% | 95.1 | 97.0 | 60.2 | 75.5 | 82.3 | 89.9 | 77.0 | 90.8 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Masked fraction | Base | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | 100% | 70% | 50% | 30% | 0% | think | no-think |
| GSM8K | 91.4 ±0.5 | 92.9 ±0.4 | 92.8 ±0.2 | 92.7 ±0.5 | 93.5 ±0.5 | 94.2 ±0.1 | 86.1 ±0.4 |
| 302 ±3 | 1124 ±12 | 1536 ±71 | 1961 ±78 | 1061 ±79 | 2332 ±27 | 258 ±2 | |
| MATH500 | 74.1 ±1.0 | 75.9 ±1.0 | 75.7 ±0.1 | 74.9 ±1.1 | 78.9 ±0.1 | 76.9 ±0.3 | 75.2 ±0.7 |
| 763 ±29 | 3080 ±125 | 3671 ±61 | 4722 ±557 | 3340 ±146 | 5137 ±214 | 1051 ±34 | |
| AIME25 | 17.8 ±1.9 | 44.4 ±1.9 | 51.1 ±1.9 | 54.4 ±6.9 | 46.7 ±3.3 | 61.1 ±5.1 | 18.9 ±1.9 |
| Masked fraction | Base | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | 100% | 70% | 50% | 30% | 0% | think | no-think |
| - | 11.5 ±0.3 | 97.5 ±1.0 | 97.0 ±0.7 | 95.8 ±1.1 | 92.3 ±0.6 | 91.9 ±0.6 | 29.2 ±1.4 |
| 67 ±0 | 3274 ±129 | 3772 ±115 | 4112 ±129 | 6369 ±120 | 6185 ±237 | 1576 ±85 | |
| - | 10.5 ±0.2 | 79.5 ±1.5 | 75.5 ±0.7 | 71.1 ±1.5 | 46.7 ±1.6 | 45.4 ±2.5 | 13.7 ±2.1 |
| 66 ±3 | 6906 ±196 | 7986 ±389 | 9384 ±344 | 12439 ±487 | 15489 ±643 | 1254 ±166 | |
| - | 10.6 ±0.3 | 88.3 ±1.0 | 89.9 ±1.3 | 87.3 ±1.9 | 65.7 ±1.8 | 62.1 ±2.6 | 18.4 ±1.9 |
| Masked fraction | Base | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | 100% | 70% | 50% | 30% | 0% | think | no-think |
| GSM8K | 92.8 ±0.1 | 94.0 ±0.2 | 94.5 ±0.1 | 93.9 ±0.5 | 95.0 ±0.3 | 94.6 ±0.2 | 80.7 ±0.4 |
| 302 ±1 | 1429 ±30 | 1884 ±45 | 2141 ±91 | 1109 ±12 | 2402 ±6 | 266 ±2 | |
| MATH500 | 74.9 ±0.2 | 73.5 ±0.8 | 76.3 ±0.4 | 76.3 ±0.1 | 73.4 ±1.4 | 77.5 ±0.1 | 73.8 ±2.5 |
| 805 ±37 | 4472 ±365 | 4249 ±77 | 4619 ±42 | 3348 ±64 | 5355 ±116 | 1181 ±120 | |
| AIME25 | 21.1 ±1.9 | 45.6 ±8.4 | 56.7 ±3.3 | 62.2 ±3.8 | 41.1 ±13.5 | 61.1 ±1.9 | 18.9 ±1.9 |
| Masked fraction | Base | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | 100% | 70% | 50% | 30% | 0% | think | no-think |
| - | 22.1 ±0.5 | 98.5 ±0.6 | 98.5 ±0.8 | 97.6 ±0.5 | 94.3 ±1.0 | 95.5 ±0.3 | 33.2 ±1.8 |
| 62 ±0 | 3650 ±123 | 3766 ±233 | 4277 ±196 | 5253 ±69 | 5752 ±131 | 1149 ±28 | |
| - | 13.0 ±0.1 | 80.4 ±1.0 | 85.6 ±0.6 | 78.3 ±2.2 | 54.4 ±1.9 | 64.4 ±2.5 | 13.5 ±0.5 |
| 211 ±25 | 9598 ±405 | 8289 ±181 | 10224 ±691 | 11855 ±266 | 14434 ±12 | 1217 ±23 | |
| - | 15.2 ±0.9 | 91.4 ±0.9 | 93.3 ±1.0 | 91.3 ±0.4 | 76.5 ±1.9 | 81.5 ±0.4 | 18.4 ±1.9 |
| Metric | Base (think) | 0% masked | 30% masked | 50% masked | 70% masked |
|---|---|---|---|---|---|
| Abrupt-start rate | 0.00 | 0.00 | 0.94 | 1.00 | 1.00 |
| Mean coherence score (1–5) | 5.00 | 5.00 | 2.09 | 1.42 | 1.29 |
| Dangling-reference rate | 0.00 | 0.00 | 0.24 | 0.91 | 0.90 |