Large Reasoning Models produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate traces dominate latency, memory usage, and serving cost, even though final answer correctness is not causally related to the trace correctness and the trace length is not a reliable indicator of the problem complexity. This raises an obvious question: can the computation expressed in these intermediate tokens be internalized into the parameters of a language model, enabling it to produce answers with much shorter intermediate traces? We propose masked self-distillation, a knowledge-distillation based post-training framework in which copies of the same model are instantiated as teacher and student, and the student model is trained to internalize all or part of the intermediate trace, thus becoming more efficient at inference. We vary the fraction of intermediate trace the student is trained to internalize, interpolating between full internalization and no internalization. We conduct controlled experiments on two reasoning domains: math and graph coloring. We use the masked self-distillation framework to post-train Qwen3-4B & 8B models. Our results demonstrate that this method can be used to improve task performance while increasing inference efficiency across various domains and model sizes. We systematically analyze whether improved efficiency gain in the post-trained models generalize to OOD problems. We find that masked self-distillation models generalize well for in-domain OOD problems, and the masked self-distillation training does not induce catastrophic forgetting in the student model on out-of-domain problems. Furthermore, our ablation study shows that supervised fine-tuning can train models to produce shorter traces, but at the cost of generalization, highlighting the importance of on-policy training in masked self-distillation.
Figures & tables
Figure 1: Overview of the masked self-distillation framework. a. For each problem x∼D , we sample a response from the teacher model. From each response, we extract the intermediate tokens z and the solution tokens s . b. The teacher πT and the student πθS are copies of the same model. The frozen teacher is conditioned on x and the first α fraction of z , while the student is conditioned only on x , so it has to internalize this α fraction into its parameters. The student’s distribution Q is matched to the teacher’s distribution P with the on-policy reverse-KL loss Ey∼Q[logQ/P] , computed on student-sampled responses; only the student is updated.
Figure 2: Qwen3-4B Pareto plot : Accuracy vs. mean response length comparison on math (top row) and graph-coloring (bottom row) domain test-datasets for α -masked variants, teacher (base think) and student (base no-think). Note: Accuracy and mean response are averaged over 3 random seeds.
Figure 3: Qwen3-8B Pareto plot : Accuracy vs. mean response length comparison on math (top-row) and graph-coloring (bottom-row) domain test-datasets for α -masked variants, teacher (base think) and student (base no-think). Note: Accuracy and mean response are averaged over 3 random seeds.
Figure 4: Out-of-domain Qwen3-8B Pareto plots: (left) graph coloring models evaluated on the GSM8K test set and (right) GSM8K models evaluated on GC - ERχ3−4 test set.
Figure 5: Qwen3-4B : Accuracy curve over max-token budget.
GC - ERχ3−4
GC - ERχ6
GC - BAχ6
GSM8K
Masked
SFT
Reverse KL
SFT
Reverse KL
SFT
Reverse KL
SFT
Reverse KL
100%
88.7
11.5
11.4
10.5
26.2
10.6
0.0
87.0
68
67
68
66
68
115
32k ∗
299
70%
94.2
97.5
52.7
79.5
72.1
88.3
68.5
87.7
2631
3274
5899
6906
4681
5185
4866
549
50%
95.1
97.0
60.2
75.5
82.3
89.9
77.0
90.8
Table 1: Qwen3-4B variants : Accuracy and mean token length comparison between SFT vs. on-policy trained α -masked variants. Each cell reports accuracy (%) on the first line and mean token length on the second. Bold marks the highest accuracy; underline marks the shortest mean response. "*" denotes the model collapses and exhausts its token generation budget.
Figure 6: Dual-model setting Pareto plot : accuracy vs. average response length comparison on graph-coloring domain test datasets for α -masked variants, base think, student (base no-think) and teacher (8B think).
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Qwen3-4B Pareto plot : accuracy vs. mean response length comparison on countdown domain test datasets for α -masked variants, teacher (base think) and student (base no-think).
Masked fraction
Base
Dataset
100%
70%
50%
30%
0%
think
no-think
GSM8K
91.4 ±0.5
92.9 ±0.4
92.8 ±0.2
92.7 ±0.5
93.5 ±0.5
94.2 ±0.1
86.1 ±0.4
302 ±3
1124 ±12
1536 ±71
1961 ±78
1061 ±79
2332 ±27
258 ±2
MATH500
74.1 ±1.0
75.9 ±1.0
75.7 ±0.1
74.9 ±1.1
78.9 ±0.1
76.9 ±0.3
75.2 ±0.7
763 ±29
3080 ±125
3671 ±61
4722 ±557
3340 ±146
5137 ±214
1051 ±34
AIME25
17.8 ±1.9
44.4 ±1.9
51.1 ±1.9
54.4 ±6.9
46.7 ±3.3
61.1 ±5.1
18.9 ±1.9
Appendix
Table 2: Qwen3-4B, math domain : accuracy and mean token length for each masked fraction, trained on GSM8K. Each entry is the mean over 3 random seeds, with the sample standard deviation in the subscript. Accuracy (%) is on the first line; the smaller line is mean response length in tokens. Bold marks the highest accuracy in each row; underline marks the shortest mean response.
Masked fraction
Base
Dataset
100%
70%
50%
30%
0%
think
no-think
GC - ERχ3-4
11.5 ±0.3
97.5 ±1.0
97.0 ±0.7
95.8 ±1.1
92.3 ±0.6
91.9 ±0.6
29.2 ±1.4
67 ±0
3274 ±129
3772 ±115
4112 ±129
6369 ±120
6185 ±237
1576 ±85
GC - ERχ6
10.5 ±0.2
79.5 ±1.5
75.5 ±0.7
71.1 ±1.5
46.7 ±1.6
45.4 ±2.5
13.7 ±2.1
66 ±3
6906 ±196
7986 ±389
9384 ±344
12439 ±487
15489 ±643
1254 ±166
GC - BAχ6
10.6 ±0.3
88.3 ±1.0
89.9 ±1.3
87.3 ±1.9
65.7 ±1.8
62.1 ±2.6
18.4 ±1.9
Appendix
Table 3: Qwen3-4B, graph coloring domain : accuracy and mean token length for each masked fraction, trained on GC - ERχ3-4 . Each entry is the mean over 3 random seeds, with the sample standard deviation in the subscript. Accuracy (%) is on the first line; the smaller line is mean response length in tokens. Bold marks the highest accuracy in each row; underline marks the shortest mean response.
Masked fraction
Base
Dataset
100%
70%
50%
30%
0%
think
no-think
GSM8K
92.8 ±0.1
94.0 ±0.2
94.5 ±0.1
93.9 ±0.5
95.0 ±0.3
94.6 ±0.2
80.7 ±0.4
302 ±1
1429 ±30
1884 ±45
2141 ±91
1109 ±12
2402 ±6
266 ±2
MATH500
74.9 ±0.2
73.5 ±0.8
76.3 ±0.4
76.3 ±0.1
73.4 ±1.4
77.5 ±0.1
73.8 ±2.5
805 ±37
4472 ±365
4249 ±77
4619 ±42
3348 ±64
5355 ±116
1181 ±120
AIME25
21.1 ±1.9
45.6 ±8.4
56.7 ±3.3
62.2 ±3.8
41.1 ±13.5
61.1 ±1.9
18.9 ±1.9
Appendix
Table 4: Qwen3-8B, math domain : accuracy and mean token length for each masked fraction, trained on GSM8K. Each entry is the mean over 3 random seeds, with the sample standard deviation in the subscript. Accuracy (%) is on the first line; the smaller line is mean response length in tokens. Bold marks the highest accuracy in each row; underline marks the shortest mean response.
Masked fraction
Base
Dataset
100%
70%
50%
30%
0%
think
no-think
GC - ERχ3-4
22.1 ±0.5
98.5 ±0.6
98.5 ±0.8
97.6 ±0.5
94.3 ±1.0
95.5 ±0.3
33.2 ±1.8
62 ±0
3650 ±123
3766 ±233
4277 ±196
5253 ±69
5752 ±131
1149 ±28
GC - ERχ6
13.0 ±0.1
80.4 ±1.0
85.6 ±0.6
78.3 ±2.2
54.4 ±1.9
64.4 ±2.5
13.5 ±0.5
211 ±25
9598 ±405
8289 ±181
10224 ±691
11855 ±266
14434 ±12
1217 ±23
GC - BAχ6
15.2 ±0.9
91.4 ±0.9
93.3 ±1.0
91.3 ±0.4
76.5 ±1.9
81.5 ±0.4
18.4 ±1.9
Appendix
Table 5: Qwen3-8B, graph coloring domain : accuracy and mean token length for each masked fraction, trained on GC - ERχ3-4 . Each entry is the mean over 3 random seeds, with the sample standard deviation in the subscript. Accuracy (%) is on the first line; the smaller line is mean response length in tokens. Bold marks the highest accuracy in each row; underline marks the shortest mean response.
Figure 8: Out-of-domain Qwen3-4B Pareto plots: (left) graph coloring models evaluated on the GSM8K test set and (right) GSM8K models evaluated on GC - ERχ3−4 test set.
Figure 9: Violin Plot
Figure 10: Qwen3-4B Violin plot : Token distribution comparison on math (top row), graph-coloring (middle row) and countdown (bottom row) domain test datasets for α -masked variants, teacher (base think) and student (base no-think).
Figure 11: Qwen3-8B Violin plot : Token distribution comparison on math (top row), graph-coloring (middle row) and countdown (bottom row) domain test datasets for α -masked variants, teacher (base think) and student (base no-think).
Figure 12: Illustration of standard generation before distillation and shorter generation after masked self-distillation.
Metric
Base (think)
0% masked
30% masked
50% masked
70% masked
Abrupt-start rate
0.00
0.00
0.94
1.00
1.00
Mean coherence score (1–5)
5.00
5.00
2.09
1.42
1.29
Dangling-reference rate
0.00
0.00
0.24
0.91
0.90
Appendix
Table 6: LLM-judge coherence analysis of traces generated by the Qwen3-8B model (teacher) and its α masked variants trained on the graph coloring ID test set
Large reasoning models often solve problems through long chain-of-thought (CoT) traces, yet much of this computation is spent on redundant derivations, repeated self-verification, and detours that do not improve the final answer. Existing on-policy self-distillation methods reduce this cost by matching a student model to a concise copy of itself on prefixes sampled from the student's own rollouts. We show that this objective has an initialization bottleneck. Since supervision is applied only to visited prefixes, training from a verbose base model places the KL loss on contexts that are often noisy, redundant, or already off track. In such regions, a concise teacher can provide only local corrections, while the student continues to explore trajectories that an efficient reasoner should avoid. In this paper, we propose BIRD(Bootstrapped Iterative Self-Reasoning Distillation), a two-stage self-reasoning distillation method that improves the rollout distribution before on-policy training. BIRD first samples concise solutions from the base model under a brevity instruction, keeps only answer-correct traces, and performs a lightweight prompt-switch SFT step. The traces are generated with the brevity instruction but learned under the original task prompt, turning instruction-induced conciseness into a default reasoning behavior. Starting from this warm model, BIRD then applies on-policy reverse-KL distillation with a concise self-teacher, now on cleaner and more informative prefixes. Across Qwen3 series models, BIRD achieves a stronger accuracy-efficiency trade-off than prompting and cold-start on-policy distillation on MATH-500 and AIME benchmarks. On Qwen3-8B, it improves MATH-500 accuracy from 86.2% to 92.0% while reducing the average response length from 3,099 to 1,115 tokens. These results highlight prefix support as a central factor in efficient reasoning distillation.
Can post-trained large language models (LLMs) further improve themselves using only unlabeled prompts, without external teachers or feedback from tools? We study this setting starting only from unlabeled seed questions with no ground-truth solutions, across three reasoning domains: math, science, and coding. We propose Self-Verified Distillation, a simple post-training refinement algorithm in which the model generates candidate solutions to these seed questions, filters them using prompt-based self-verification, and trains on the resulting self-curated dataset. Inspired by the UQ benchmark's use of multiple validators to screen candidate answers to hard unsolved questions, we adapt this validation-based filtering idea to self-training: the model filters its own generated solutions through a three-stage cascade of cycle-consistency, factuality, and correctness checks, accepting a solution only if it passes all stages with unanimous judge votes. We find that sampling more candidate generations and using a larger verification budget during training data construction produces higher-quality self-curated data and, in turn, better reasoning models. We then train Qwen3 models at multiple scales with Self-Verified Distillation and obtain gains across all three domains. For Qwen3-4B, our method improves aggregate held-out pass@1 by +16.7 points in math (AIME26 and HMMT), +11.1 points in science (GPQA Diamond and HLE), and +8.3 points in coding (LCBv5 and LCBv6), with gains also extending to 0.6B and 8B models. Compared to our test-time-only baseline (UQ-TTC), which improves performance by spending extra compute at inference time, Self-Verified Distillation achieves better performance in most settings while requiring only a single inference call at test time.
Distilling reasoning traces from strong large language models into smaller ones is a promising route to improve intelligence in resource-constrained settings. Existing approaches face a fundamental trade-off: offline distillation from teacher-generated traces provides high-quality, sample-efficient supervision but suffers from distributional drift: during training, the student model conditions on teacher-generated prefixes, whereas during inference the student autoregresses on self-generated prefixes, leading to compounding errors over long reasoning trajectories. Meanwhile, on-policy or self-distillation methods better match the student's inference-time distribution, but require costly online sampling and often produce low-quality traces in early training. We propose a principled offline reasoning distillation framework that preserves the efficiency and supervision quality of offline teacher-generated data while correcting teacher-student distribution drift. It adaptively emphasizes teacher supervision that is better aligned with the student's on-policy distribution. Evaluations on mathematical reasoning benchmarks of GSM8K, MATH, MATH500, and harder held-out competition-style tasks, including AMC, AIME, and OlympiadBench, show that our method improves reasoning accuracy over prior offline distillation algorithms and yields more stable reasoning traces while preserving instruction-following capabilities. Our work shows that lightweight, distribution-correction-aware training can substantially strengthen offline reasoning distillation without online rollouts.