Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out challenging examples through data selection or introduce weaker intermediate assistant models, inherently compromising supervision coverage or quality. We propose Teacher Alignment, which directly adapts the teacher toward the student's distribution without discarding data or degrading reasoning quality. However, naive alignment through standard knowledge distillation triggers catastrophic collapse of the teacher's reasoning capabilities. To address this, we reformulate teacher alignment as reinforcement learning and introduce TeacherGRPO, built on Group Relative Policy Optimization with two key innovations: (i) Curriculum Selective Alignment applies dual token- and distribution-level curricula to focus rewards on high-signal reasoning gaps while filtering noise from trivial tokens and uncertain tail distributions, and (ii) Importance-Adaptive Length Regularization selectively penalizes verbose redundancy while preserving pedagogically critical reasoning steps. The aligned teacher then distills knowledge to students via standard pipelines. Extensive experiments show TeacherGRPO significantly outperforms baselines across diverse reasoning benchmarks and distillation methods. Our code is available at https://github.com/LzyFischer/TeacherGRPO.
Figures & tables
Figure 1: Approaches to address the Gap Curse. Data-Centric Selection discards potentially valuable examples while Intermediate Supervision dilutes supervision quality. Teacher Alignment adapts the teacher to student capacity, preserving both data and teaching quality.
Figure 2: Teacher and student performance during alignment. Stars indicate peak student performance. KD-based alignment initially improves distillation but causes catastrophic teacher collapse. GRPO maintains teacher capability but achieves limited alignment effectiveness, motivating TeacherGRPO.
Figure 3: TeacherGRPO framework overview. The Alignment phase optimizes the teacher via GRPO with two novel reward CSA and IALR, plus a correctness verification. The aligned teacher then supervises student distillation.
Method
Qwen2.5 ( 3B→0.5B )
Gemma3 ( 1B→270M )
Date
SQA
ARC
CQA
Avg
Date
SQA
ARC
CQA
Avg
Teacher
Base
58.6
59.8
81.7
71.5
67.9
37.9
51.1
47.9
47.3
46.1
+ TeacherKD
53.9
55.5
79.2
61.2
62.5
31.5
47.2
46.3
44.3
43.8
+ TeacherGRPO
59.2
64.6
81.2
72.7
69.4
40.2
51.5
48.3
47.1
46.8
Student
Table 1: Performance comparison across four reasoning benchmarks with two model families: Qwen and Gemma. Best average values are in bold.
Method
Date
ARC
Avg
KD
TeacherGRPO
40.2
45.2
42.7
w/o CSA
34.9
43.6
39.3
w/o IALR
36.4
42.8
39.6
SeqKD
TeacherGRPO
38.8
41.3
40.1
Table 2: Ablation results on Date and ARC with Qwen. Best values are in bold.
Method
Date
ARC
Avg
KD
+ Data-Centric Selection
32.5
43.9
38.2
+ Intermediate Supervision
34.9
41.2
38.1
+ TeacherGRPO
40.2
45.2
42.7
SeqKD
+ Data-Centric Selection
36.1
40.6
38.4
Table 3: Performance comparison of different paradigms on Date and ARC with Qwen.
Figure 4: KL divergence visualization for a sample path. Darker tokens correspond to high-divergence.
Method
Token
Distribution
Curriculum
40.2
40.2
Top 30%
36.1
35.5
Bottom 30%
30.2
27.2
Table 4: Student performance using top 30%, bottom 30%, or curriculum-based token and distribution selection on Date Understanding with Qwen.
Figure 5: Impact of length penalty. (a) Student performance and (b) generation length during Alignment.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Qwen2.5 (3B → 0.5B)
Gemma3 (1B → 270M)
Base
20.9
3.7
KD
24.8
4.1
+ TeacherKD
22.1
3.9
+ TeacherGRPO
25.5
4.2
SeqKD
25.6
3.5
+ TeacherKD
30.0
3.8
Appendix
Table 5: Accuracy on MATH , a multi-step mathematical reasoning benchmark with longer and more interdependent reasoning chains than the main benchmarks. TeacherGRPO consistently improves over vanilla distillation across both model families and all three distillation pipelines, while TeacherKD remains unstable.
Method
Date
SQA
ARC
CQA
Base (Qwen2.5-0.5B)
20.12
35.81
37.37
39.14
SpeculativeKD
24.26
44.98
39.68
47.42
+ TeacherKD
22.49
50.22
35.58
39.23
+ TeacherGRPO
25.44
50.66
41.72
47.50
Appendix
Table 6: Results on SpeculativeKD ( Xu et al., 2025b ) with the Qwen model family. TeacherGRPO improves SpeculativeKD across all four benchmarks, while TeacherKD is unstable and degrades performance on ARC and CommonsenseQA.
β
0.001
0.005
0.01
0.02
0.05
Accuracy
36.4
37.0
40.2
39.4
35.8
Appendix
Table 7: Sensitivity of student accuracy to the IALR penalty strength β on Date Understanding with Qwen (KD distillation). The default value ( β=0.01 , bolded) performs best, and performance degrades gracefully away from it rather than collapsing.
Figure 6: Qualitative comparison of teacher generations with different length penalties. We compare reasoning outputs from four teacher configurations on a Date Understanding example: Vanilla (no alignment), TeacherGRPO, Uniform Penalty, and w/o IALR.
Figure 7: Training dynamics during Teacher Alignment on the Date Understanding dataset. Both curves are smoothed using exponential moving average (EMA weight = 0.9).
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, Nanjing, China · 2The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China