Organizations: Institute of Software, Chinese Academy of Sciences · Microsoft · Work done during an internship at Microsoft · Northwestern Polytechnical University
Reinforcement learning can substantially improve a reasoning teacher, but it is unclear which of those improvements survive when the teacher supervises a smaller on-policy student. We study this question in mathematical reasoning by comparing teacher lineages before and after GRPO, multiple student scales, direct GRPO, and several on-policy distillation objectives. The central finding is that transfer is structured rather than scalar: teacher strength alone does not make dense distillation competitive, while an RL-improved teacher creates useful but metric-dependent student gains. This motivates Transfer-Stratified On-Policy Distillation (TS-OPD), which screens training problems by the joint sampled success of the student and teacher, routes acquisition problems to gated forward KL, routes consolidation problems to gated reverse KL, and adds an entropy brake to protect sampled coverage. Across the main comparison, TS-OPD is the strongest student objective for macro average correctness with the GRPO-improved teacher, while pass@K remains more mixed. Ablations show that the gains come from routing and token gating rather than skipping problems. These results support a transfer-aware view of OPD: stronger teachers help when the supervision direction and token budget match the student's observed ability, not merely because the teacher endpoint is stronger.
Figures & tables
Figure 1: Overview of the transfer gap and its problem-level decomposition. Left: GRPO yields a teacher improvement ΔT , but standard OPD transfers only part of it into a student gain ΔS , leaving a transfer gap of ΔT−ΔS . Middle: Joint teacher–student pass rates identify knowledge gaps (teacher >0 , student =0 ), precision gaps (teacher > student >0 ), and problems with no transferable teacher advantage (teacher ≤ student). Right: TS-OPD routes these three strata to matched actions—discover, refine, and skip, respectively.
Table 1: Training strata and test gain shortfalls. (a) Fractions of all 7,496 MATH training problems, classified using eight responses per model before training. A/C denote knowledge/precision gaps; P/N denote student advantage/no observed gap. (b) Test avg@16 accuracy on AIME24/25 and HMMT-Feb/Nov25 (120 problems), with gains and transfer gaps in pp. Both students share the Qwen3-8B teacher checkpoints; SOPD uses standard reverse-KL OPD from S0 with T1 . Values are rounded independently to two decimal places.
Figure 2: Overview of TS-OPD. Before training, each problem is labeled by the sampled success of the teacher and the student. Problem level: P and N problems are skipped, A problems use forward KL, and C problems use reverse KL. Token level: gates select positions within a response, and C adds an entropy brake.
base-8B teacher
GRPO-8B teacher
TS-OPD
TS-OPD
Direct
Student
Metric
OPD-r
OPD-f
routing only
full
OPD-r
OPD-f
routing only
full
GRPO
0.6B
avg@ K↑
19.61
19.15
20.60
21.27
27.00
26.57
27.62
28.88
26.04
pass@ K↑
32.80
32.91
34.14
36.22
44.11
41.96
44.83
44.57
40.46
Gtransfer↓
15.10
15.51
13.84
13.06
7.39
8.22
6.93
5.68
8.71
1.7B
avg@ K↑
29.22
27.57
29.29
29.71
39.04
39.25
39.84
40.50
39.53
Table 2: Main results. avg@ K and pass@ K are macro averages over the eight benchmarks (%). Gtransfer is the transfer gap in percentage points, computed on seven benchmarks. ↑ / ↓ : higher/lower is better. Direct GRPO uses no teacher. In each row, bold marks the best value and underline the second best.
Figure 3: Outcome categories of test problems in each student. (a) The 567 problems GRPO improves for the teacher. (b) Knowledge-gap (upper bar) and precision-gap (lower bar) problems. (c, d) Problems where T1 is not better than S0 . Numbers: Tables 9 , 11 , and 10 .
Step
Method
0.6B
1.7B
4B
No teacher
Direct GRPO
26.04 / 8.71
39.53 / 6.94
51.57 / 3.45
+ A ∪ C only
Direct GRPO (A ∪ C)
24.91 / 9.34
37.82 / 8.16
47.48 / 4.21
Start
OPD-r
27.00 / 7.39
39.04 / 7.28
52.27 / 2.65
+ filtering
OPD-r (A ∪ C)
25.66 / 9.02
37.51 / 9.12
47.70 / 7.87
Start
OPD-f
26.57 / 8.22
39.25 / 7.15
52.83 / 1.98
+ filtering
OPD-f (A ∪ C)
24.87 / 10.17
36.52 / 10.34
47.19 / 8.45
Table 3: Adding one component at a time: macro avg@ K (%) / Gtransfer (points), GRPO-8B teacher. All A ∪ C rows use the same problems and steps.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Student
base-8B teacher
GRPO-8B teacher
0.6B
0.078
0.079
1.7B
0.058
0.063
4B
0.053
0.054
Appendix
Table 4: Entropy-brake coefficient β for each teacher–student pair.
Shared
Optimizer
Adam ( β1=0.9 , β2=0.98 ), weight decay 0.1
Learning rate
1×10−6 , constant
Responses per update
64
Rollout
temperature 1.0, max response length 8,192
Policy clipping
ϵlow=0.2 , ϵhigh=0.28
KL / entropy regularizer
none (coefficients 0)
Appendix
Table 5: Training hyperparameters. Shared settings apply to all runs.
Arm
Training problems
Loss
Direct GRPO
all
correctness reward, no teacher
Direct GRPO (A ∪ C)
A ∪ C
same
OPD-r
all
reverse KL
OPD-f
all
forward KL
OPD-r (A ∪ C)
A ∪ C
reverse KL
OPD-f (A ∪ C)
A ∪ C
forward KL
Appendix
Table 6: Training arms. All arms except the two Direct GRPO rows use the GRPO-8B teacher; Table 2 also reports the four OPD objectives with the base-8B teacher.
Setting
avg@ K
pass@ K
τ=0.60
27.17
41.45
τ=0.95
24.71
36.41
20% of positions by student entropy
26.28
40.54
20% of positions at random
26.36
41.34
mass@n>0.0
26.70
39.93
mass@n>0.5
26.75
42.59
Appendix
Table 7: Ablations with the 0.6B student and the GRPO-8B teacher (macro averages over eight benchmarks, %). Each row changes one setting of the default TS-OPD.
Figure 4: Outcome categories of all 1,566 test problems after training. Students are compared with S0 ; the first row shows the teacher from T0 to T1 .
0.6B
1.7B
4B
Method
A
C
P
N
A
C
P
N
A
C
P
N
Teacher T0→T1
8.4
27.8
5.8
58.0
8.4
27.8
5.8
58.0
8.4
27.8
5.8
58.0
Direct GRPO
14.3
27.3
7.9
50.6
10.7
28.9
6.9
53.6
7.5
22.8
9.8
59.8
Direct GRPO (A ∪ C)
13.5
27.0
8.0
51.5
10.2
28.3
7.2
54.3
7.0
22.3
9.8
60.9
OPD-r
14.6
26.8
8.6
50.0
9.6
25.0
11.7
53.7
7.9
24.3
7.2
60.7
OPD-r (A ∪ C)
15.8
28.7
7.3
48.1
10.0
27.3
9.1
53.6
7.0
24.1
7.9
61.0
Appendix
Table 8: Outcome categories of all 1,566 benchmark problems after training (%), relative to the raw student S0 ; the teacher row is T0→T1 .
0.6B
1.7B
4B
Method
A
C
P
N
A
C
P
N
A
C
P
N
Direct GRPO
24.3
16.9
8.8
49.9
23.3
41.4
6.7
28.6
17.3
45.1
10.8
26.8
Direct GRPO (A ∪ C)
21.0
16.1
9.0
54.0
20.5
39.0
7.6
33.0
14.5
43.6
11.0
31.0
OPD-r
25.0
16.2
7.9
50.8
21.7
38.9
8.6
30.8
18.7
47.6
7.4
26.3
OPD-r (A ∪ C)
28.0
18.0
7.9
46.0
23.3
39.9
8.6
28.2
15.9
47.4
8.5
28.2
OPD-f
29.1
20.3
6.5
44.1
22.4
35.6
11.3
30.7
18.2
46.9
8.8
26.1
Appendix
Table 9: Outcome categories, in the student, of the 567 problems that GRPO improves for the teacher (%).
0.6B
1.7B
4B
Method
A
C
P
N
A
C
P
N
A
C
P
N
P stratum: S0 more accurate than T1
Direct GRPO
-
12.2
65.9
22.0
-
20.0
49.2
30.8
-
6.8
72.8
20.4
Direct GRPO (A ∪ C)
-
14.6
63.4
22.0
-
18.5
50.8
30.8
-
7.8
71.8
20.4
OPD-r
-
14.6
78.0
7.3
-
16.9
70.8
12.3
-
13.6
58.3
28.2
OPD-r (A ∪ C)
-
17.1
73.2
9.8
-
13.8
64.6
21.5
-
12.6
68.9
18.4
Appendix
Table 10: Outcome categories (%) of test problems where the teacher is not better (Figure 3 (c, d)). P stratum: S0 more accurate than T1 (41, 65, and 103 problems for 0.6B, 1.7B, and 4B). N stratum: equally accurate (453, 711, and 909 problems). Each row sums to 100 within a size; “-” marks a category that cannot occur. All arms except the two Direct GRPO rows use the GRPO-8B teacher.
0.6B
1.7B
4B
Method
A
C
P
N
A
C
P
N
A
C
P
N
Knowledge-gap problems: S0 at 0, T1 above 0
Direct GRPO
41.0
-
-
59.0
54.9
-
-
45.1
67.8
-
-
32.2
Direct GRPO (A ∪ C)
38.1
-
-
61.9
51.9
-
-
48.1
59.9
-
-
40.1
OPD-r
41.6
-
-
58.4
52.7
-
-
47.3
71.1
-
-
28.9
OPD-r (A ∪ C)
40.2
-
-
59.8
50.4
-
-
49.6
59.2
-
-
40.8
Appendix
Table 11: Outcome categories (%) of knowledge-gap and precision-gap test problems (Figure 3 (b)). All arms except the two Direct GRPO rows use the GRPO-8B teacher; the A ∪ C rows share the same problems and steps. Knowledge-gap problems start at 0, so they can only be A or N ; precision-gap problems start above 0, so they can only be C , P , or N . There are 517 and 555 such problems for 0.6B, 264 and 526 for 1.7B, and 152 and 402 for 4B.
base-8B teacher
GRPO-8B teacher
No teacher
Student
Benchmark
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
Direct GRPO
0.6B
MATH500
-3.05
-1.45
-5.92
-6.64
-14.42
-12.17
-15.22
-18.07
-12.55
Minerva
6.20
5.87
5.32
4.31
-1.52
-0.74
-0.47
-0.79
-0.38
OlympiadBench
4.63
5.58
3.11
2.26
-4.62
-0.99
-4.14
-4.94
-1.12
AIME24
30.62
28.75
29.58
27.08
24.16
23.54
21.46
22.50
24.58
AIME25
26.05
26.46
23.96
22.92
15.00
14.17
13.55
12.09
14.79
Appendix
Table 12: Transfer gap Gtransfer (percentage points) on each benchmark with K≥8 . The last row per student is the macro average used in Table 2 .
Teacher: base-8B
Teacher: GRPO-8B
Benchmark
Metric
Direct GRPO
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
MATH500
Avg@8
62.58
53.08
51.48
55.95
56.67
64.45
62.20
65.25
68.10
Pass@8
82.00
78.80
78.00
82.80
81.60
85.00
82.60
84.40
85.40
Minerva
Avg@8
16.87
10.29
10.62
11.17
12.18
18.01
17.23
16.96
17.28
Pass@8
30.88
25.74
27.21
28.31
26.84
33.46
34.93
32.72
32.35
OlympiadBench
Avg@8
28.86
23.11
22.16
24.63
25.48
32.36
28.73
31.88
32.68
Appendix
Table 13: Per-benchmark results for the 0.6B student (%). Direct GRPO uses no teacher. Macro averages weight each benchmark equally.
Teacher: base-8B
Teacher: GRPO-8B
Benchmark
Metric
Direct GRPO
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
MATH500
Avg@8
81.38
71.38
68.17
71.17
72.10
79.92
78.80
79.75
80.25
Pass@8
93.00
90.40
89.60
90.40
89.80
91.80
91.20
92.00
92.00
Minerva
Avg@8
29.00
20.36
18.29
21.00
22.01
27.34
27.85
27.67
28.31
Pass@8
44.49
37.50
35.66
37.50
37.13
42.28
42.65
44.12
44.12
OlympiadBench
Avg@8
48.94
36.20
33.77
37.11
37.22
45.72
44.25
46.07
46.42
Appendix
Table 14: Per-benchmark results for the 1.7B student (%). Direct GRPO uses no teacher. Macro averages weight each benchmark equally.
Teacher: base-8B
Teacher: GRPO-8B
Benchmark
Metric
Direct GRPO
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
MATH500
Avg@8
88.47
80.12
78.95
80.15
80.88
88.97
88.47
89.40
90.03
Pass@8
94.80
93.60
93.40
93.40
93.20
95.40
94.60
95.60
95.20
Minerva
Avg@8
35.80
29.32
29.73
29.55
29.41
36.67
35.39
36.12
39.20
Pass@8
48.53
43.75
44.49
43.38
41.54
48.53
51.84
47.79
52.21
OlympiadBench
Avg@8
55.10
45.16
43.90
46.07
46.88
56.84
56.66
56.94
57.14
Appendix
Table 15: Per-benchmark results for the 4B student (%). Direct GRPO uses no teacher. Macro averages weight each benchmark equally.
Benchmark
Metric
Qwen3-0.6B
Qwen3-1.7B
Qwen3-4B
base-8B
GRPO-8B
MATH500
Avg@8
42.23
70.00
81.45
82.30
90.10
Pass@8
72.60
89.40
93.40
93.60
95.40
Minerva
Avg@8
8.64
20.96
29.46
30.06
37.91
Pass@8
25.74
35.66
41.54
42.65
50.00
OlympiadBench
Avg@8
15.37
34.92
47.94
47.59
59.96
Pass@8
37.09
56.82
66.02
67.21
72.26
Appendix
Table 16: Per-benchmark results of the raw students and the two teachers (%).
On-policy distillation is a powerful way to transfer reasoning ability from a strong teacher to a smaller student: the student samples trajectories from its own policy, and the teacher provides dense token-level supervision on the states the student actually visits. However, this supervision is not always reliable: a teacher can assign high likelihood to plausible but incorrect solutions, or low likelihood to correct student solutions that follow different reasoning paths. Unconditionally distilling the teacher can therefore reinforce bad modes or erase useful student behavior. To address these limitations, we introduce RG-OPD: Reward-Gated On-Policy Distillation that uses verifier feedback to decide when teacher logits should be trusted. RG-OPD bridges sparse verifier rewards and dense teacher logits, preserving token-level supervision while filtering misleading teacher signals. Across reasoning and coding benchmarks, RG-OPD produces stronger distilled students, outperforming both vanilla reverse-KL distillation and the recent TSD-KD baseline. At 1K generation length, RG-OPD improves over reverse-KL by 2.9 points and over TSD-KD by 4.9 points; in the long-generation setting, it improves over the untuned student by 8.2 points. Our code is available at https://github.com/UoC-tail/RG-OPD.
Mohammad Sadegh Akhondzadeh, Vijay Lingam, Atula Tejaswi +3
Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond improvements in the distilled model's initial accuracy. Under shared RL settings, students initialized with OPD reach higher final performance than those trained with direct RL or supervised fine-tuning followed by RL. This advantage can emerge even when OPD produces little immediate improvement in accuracy. Pre-RL Pass@k does not fully explain the benefit: similar or even higher values do not necessarily lead to better performance after RL. Behavioral analyses point to alignment with the teacher's distribution beyond top-1 agreement as a possible explanation. Such alignment may favor higher-quality reasoning paths while retaining alternatives that RL can further refine using outcome feedback. We further examine how trajectory sources and divergence objectives affect the value of distillation for subsequent RL. Standard reverse-KL OPD performs better before RL, but forward-KL OPD overtakes it afterward; with teacher-generated distillation trajectories, reverse KL remains ahead at both stages. These findings suggest that the preferred distillation objective depends on both the trajectory source and the training that follows. Our results support evaluating OPD as preparation for RL and selecting distillation choices by the performance achieved after subsequent training.
Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out challenging examples through data selection or introduce weaker intermediate assistant models, inherently compromising supervision coverage or quality. We propose Teacher Alignment, which directly adapts the teacher toward the student's distribution without discarding data or degrading reasoning quality. However, naive alignment through standard knowledge distillation triggers catastrophic collapse of the teacher's reasoning capabilities. To address this, we reformulate teacher alignment as reinforcement learning and introduce TeacherGRPO, built on Group Relative Policy Optimization with two key innovations: (i) Curriculum Selective Alignment applies dual token- and distribution-level curricula to focus rewards on high-signal reasoning gaps while filtering noise from trivial tokens and uncertain tail distributions, and (ii) Importance-Adaptive Length Regularization selectively penalizes verbose redundancy while preserving pedagogically critical reasoning steps. The aligned teacher then distills knowledge to students via standard pipelines. Extensive experiments show TeacherGRPO significantly outperforms baselines across diverse reasoning benchmarks and distillation methods. Our code is available at https://github.com/LzyFischer/TeacherGRPO.
Zhenyu Lei, Zihan Chen, Yaochen Zhu +5
University of Virginia · Netflix · University of Washington +2