Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out challenging examples through data selection or introduce weaker intermediate assistant models, inherently compromising supervision coverage or quality. We propose Teacher Alignment, which directly adapts the teacher toward the student's distribution without discarding data or degrading reasoning quality. However, naive alignment through standard knowledge distillation triggers catastrophic collapse of the teacher's reasoning capabilities. To address this, we reformulate teacher alignment as reinforcement learning and introduce TeacherGRPO, built on Group Relative Policy Optimization with two key innovations: (i) Curriculum Selective Alignment applies dual token- and distribution-level curricula to focus rewards on high-signal reasoning gaps while filtering noise from trivial tokens and uncertain tail distributions, and (ii) Importance-Adaptive Length Regularization selectively penalizes verbose redundancy while preserving pedagogically critical reasoning steps. The aligned teacher then distills knowledge to students via standard pipelines. Extensive experiments show TeacherGRPO significantly outperforms baselines across diverse reasoning benchmarks and distillation methods. Our code is available at https://github.com/LzyFischer/TeacherGRPO.
Figures & tables
Figure 1: Approaches to address the Gap Curse. Data-Centric Selection discards potentially valuable examples while Intermediate Supervision dilutes supervision quality. Teacher Alignment adapts the teacher to student capacity, preserving both data and teaching quality.
Figure 2: Teacher and student performance during alignment. Stars indicate peak student performance. KD-based alignment initially improves distillation but causes catastrophic teacher collapse. GRPO maintains teacher capability but achieves limited alignment effectiveness, motivating TeacherGRPO.
Figure 3: TeacherGRPO framework overview. The Alignment phase optimizes the teacher via GRPO with two novel reward CSA and IALR, plus a correctness verification. The aligned teacher then supervises student distillation.
Method
Qwen2.5 ( 3B→0.5B )
Gemma3 ( 1B→270M )
Date
SQA
ARC
CQA
Avg
Date
SQA
ARC
CQA
Avg
Teacher
Base
58.6
59.8
81.7
71.5
67.9
37.9
51.1
47.9
47.3
46.1
+ TeacherKD
53.9
55.5
79.2
61.2
62.5
31.5
47.2
46.3
44.3
43.8
+ TeacherGRPO
59.2
64.6
81.2
72.7
69.4
40.2
51.5
48.3
47.1
46.8
Student
Table 1: Performance comparison across four reasoning benchmarks with two model families: Qwen and Gemma. Best average values are in bold.
Method
Date
ARC
Avg
KD
TeacherGRPO
40.2
45.2
42.7
w/o CSA
34.9
43.6
39.3
w/o IALR
36.4
42.8
39.6
SeqKD
TeacherGRPO
38.8
41.3
40.1
Table 2: Ablation results on Date and ARC with Qwen. Best values are in bold.
Method
Date
ARC
Avg
KD
+ Data-Centric Selection
32.5
43.9
38.2
+ Intermediate Supervision
34.9
41.2
38.1
+ TeacherGRPO
40.2
45.2
42.7
SeqKD
+ Data-Centric Selection
36.1
40.6
38.4
Table 3: Performance comparison of different paradigms on Date and ARC with Qwen.
Figure 4: KL divergence visualization for a sample path. Darker tokens correspond to high-divergence.
Method
Token
Distribution
Curriculum
40.2
40.2
Top 30%
36.1
35.5
Bottom 30%
30.2
27.2
Table 4: Student performance using top 30%, bottom 30%, or curriculum-based token and distribution selection on Date Understanding with Qwen.
Figure 5: Impact of length penalty. (a) Student performance and (b) generation length during Alignment.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Qwen2.5 (3B → 0.5B)
Gemma3 (1B → 270M)
Base
20.9
3.7
KD
24.8
4.1
+ TeacherKD
22.1
3.9
+ TeacherGRPO
25.5
4.2
SeqKD
25.6
3.5
+ TeacherKD
30.0
3.8
Appendix
Table 5: Accuracy on MATH , a multi-step mathematical reasoning benchmark with longer and more interdependent reasoning chains than the main benchmarks. TeacherGRPO consistently improves over vanilla distillation across both model families and all three distillation pipelines, while TeacherKD remains unstable.
Method
Date
SQA
ARC
CQA
Base (Qwen2.5-0.5B)
20.12
35.81
37.37
39.14
SpeculativeKD
24.26
44.98
39.68
47.42
+ TeacherKD
22.49
50.22
35.58
39.23
+ TeacherGRPO
25.44
50.66
41.72
47.50
Appendix
Table 6: Results on SpeculativeKD ( Xu et al., 2025b ) with the Qwen model family. TeacherGRPO improves SpeculativeKD across all four benchmarks, while TeacherKD is unstable and degrades performance on ARC and CommonsenseQA.
β
0.001
0.005
0.01
0.02
0.05
Accuracy
36.4
37.0
40.2
39.4
35.8
Appendix
Table 7: Sensitivity of student accuracy to the IALR penalty strength β on Date Understanding with Qwen (KD distillation). The default value ( β=0.01 , bolded) performs best, and performance degrades gracefully away from it rather than collapsing.
Figure 6: Qualitative comparison of teacher generations with different length penalties. We compare reasoning outputs from four teacher configurations on a Date Understanding example: Vanilla (no alignment), TeacherGRPO, Uniform Penalty, and w/o IALR.
Figure 7: Training dynamics during Teacher Alignment on the Date Understanding dataset. Both curves are smoothed using exponential moving average (EMA weight = 0.9).
On-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this signal is beneficial and under which it is detrimental. Which teacher model should be used, and in the case of self-distillation, which specific context should serve as the supervisory signal? Does the optimal choice vary from one token to the next? At present, addressing these questions typically requires costly training runs whose aggregate performance metrics obscure the dynamics at the level of individual tokens. We introduce a training-free diagnostic framework that operates at the highest resolution: per token, per question, and per teacher. We derive an ideal per-node gradient defined as the parameter update that maximally increases the student's probability of success. We then develop a scalable targeted-rollout algorithm to estimate this gradient efficiently, even for long chains of intermediate thoughts. The gradient alignment score, defined as the cosine similarity between this ideal gradient and any given distillation gradient, quantifies the extent to which a particular configuration approximates the ideal signal. Across a range of self-distillation settings and external teacher models, we observe that distillation guidance exhibits substantially higher alignment with the ideal on incorrect rollouts than on correct ones, where the student already performs well and the teacher's signal tends to become noisy. Furthermore, we find that the optimal distillation context depends jointly on the student model's capacity and the target task, and that no single universally effective configuration emerges. These findings motivate the use of per-task, per-token diagnostic analyses for distillation.
Mohammadreza Armandpour, Fatih Ilhan, David Harrison +6
When distilling reasoning from large language models (LLMs) into smaller ones, teacher rationales for similar problems often vary wildly in structure and strategy. Like a chef who makes the same dish differently each time, this inconsistency burdens the student with noisy supervision that is hard to internalize. We propose Distillation through Reasoning Path Compression (D-RPC), which constrains the teacher to follow a compact, dynamically maintained bank of reusable high-level reasoning paths. For each training question, D-RPC retrieves the most relevant path and conditions the teacher to follow it, producing rationales that are consistent across similar problems yet diverse enough to cover different problem types. A PAC-Bayes analysis formalizes the resulting trade-off between bank size and coverage: smaller banks reduce supervision entropy but risk coverage gaps, and the generalization bound identifies an optimal intermediate size confirmed by our ablations. Across five math and commonsense reasoning benchmarks with two student models, D-RPC consistently outperforms chain-of-thought distillation, freeform rationale generation, direct distillation, and structured-supervision baselines, while using fewer tokens than template-heavy alternatives.
Jialin Yang, Jiankun Wang, Jiajun Wu +3
University of Calgary · Calgary, Canada · University of Michigan +1
Recent advances in large reasoning models (LRMs) have shown strong performance on complex problems through long chain-of-thought (Long CoT) reasoning. However, distilling such trajectories into smaller student models remains challenging: direct Long CoT supervision often provides limited gains and can be less effective than concise Short CoT rationales. In this work, we investigate this phenomenon from a gradient-centric perspective. Our analysis shows that Long CoT induces larger gradient magnitudes and more concentrated update directions than Short CoT, with this effect becoming more pronounced as student model capacity increases. These findings suggest that effective Long CoT distillation requires balancing the reasoning information density of reasoning trajectories with their distributional alignment to the student model. Motivated by this insight, we propose \textbf{M}odel \textbf{I}nterporlation \textbf{Distillation} (\textbf{MI-Distillation}), a framework that constructs a continuous Instruct-Reasoning data spectrum through model interpolation. To select suitable trajectories from this spectrum, we further introduce \textbf{Seq}uential \textbf{L}earnable \textbf{S}urprisal \textbf{S}core (\textbf{SeqLSS}), which favors reasoning paths that are both informative and learnable for the student. Extensive experiments on reasoning benchmarks show that MI-Distillation consistently improves small model CoT distillation over strong Long CoT baselines.
Yangsong Lan, Renkai Hu, HongKai Zheng +4
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, Nanjing, China · 2The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China