Organizations: Institute of Software, Chinese Academy of Sciences · Microsoft · Work done during an internship at Microsoft · Northwestern Polytechnical University
Reinforcement learning can substantially improve a reasoning teacher, but it is unclear which of those improvements survive when the teacher supervises a smaller on-policy student. We study this question in mathematical reasoning by comparing teacher lineages before and after GRPO, multiple student scales, direct GRPO, and several on-policy distillation objectives. The central finding is that transfer is structured rather than scalar: teacher strength alone does not make dense distillation competitive, while an RL-improved teacher creates useful but metric-dependent student gains. This motivates Transfer-Stratified On-Policy Distillation (TS-OPD), which screens training problems by the joint sampled success of the student and teacher, routes acquisition problems to gated forward KL, routes consolidation problems to gated reverse KL, and adds an entropy brake to protect sampled coverage. Across the main comparison, TS-OPD is the strongest student objective for macro average correctness with the GRPO-improved teacher, while pass@K remains more mixed. Ablations show that the gains come from routing and token gating rather than skipping problems. These results support a transfer-aware view of OPD: stronger teachers help when the supervision direction and token budget match the student's observed ability, not merely because the teacher endpoint is stronger.
Figures & tables
Figure 1: Overview of the transfer gap and its problem-level decomposition. Left: GRPO yields a teacher improvement ΔT , but standard OPD transfers only part of it into a student gain ΔS , leaving a transfer gap of ΔT−ΔS . Middle: Joint teacher–student pass rates identify knowledge gaps (teacher >0 , student =0 ), precision gaps (teacher > student >0 ), and problems with no transferable teacher advantage (teacher ≤ student). Right: TS-OPD routes these three strata to matched actions—discover, refine, and skip, respectively.
Table 1: Training strata and test gain shortfalls. (a) Fractions of all 7,496 MATH training problems, classified using eight responses per model before training. A/C denote knowledge/precision gaps; P/N denote student advantage/no observed gap. (b) Test avg@16 accuracy on AIME24/25 and HMMT-Feb/Nov25 (120 problems), with gains and transfer gaps in pp. Both students share the Qwen3-8B teacher checkpoints; SOPD uses standard reverse-KL OPD from S0 with T1 . Values are rounded independently to two decimal places.
Figure 2: Overview of TS-OPD. Before training, each problem is labeled by the sampled success of the teacher and the student. Problem level: P and N problems are skipped, A problems use forward KL, and C problems use reverse KL. Token level: gates select positions within a response, and C adds an entropy brake.
base-8B teacher
GRPO-8B teacher
TS-OPD
TS-OPD
Direct
Student
Metric
OPD-r
OPD-f
routing only
full
OPD-r
OPD-f
routing only
full
GRPO
0.6B
avg@ K↑
19.61
19.15
20.60
21.27
27.00
26.57
27.62
28.88
26.04
pass@ K↑
32.80
32.91
34.14
36.22
44.11
41.96
44.83
44.57
40.46
Gtransfer↓
15.10
15.51
13.84
13.06
7.39
8.22
6.93
5.68
8.71
1.7B
avg@ K↑
29.22
27.57
29.29
29.71
39.04
39.25
39.84
40.50
39.53
Table 2: Main results. avg@ K and pass@ K are macro averages over the eight benchmarks (%). Gtransfer is the transfer gap in percentage points, computed on seven benchmarks. ↑ / ↓ : higher/lower is better. Direct GRPO uses no teacher. In each row, bold marks the best value and underline the second best.
Figure 3: Outcome categories of test problems in each student. (a) The 567 problems GRPO improves for the teacher. (b) Knowledge-gap (upper bar) and precision-gap (lower bar) problems. (c, d) Problems where T1 is not better than S0 . Numbers: Tables 9 , 11 , and 10 .
Step
Method
0.6B
1.7B
4B
No teacher
Direct GRPO
26.04 / 8.71
39.53 / 6.94
51.57 / 3.45
+ A ∪ C only
Direct GRPO (A ∪ C)
24.91 / 9.34
37.82 / 8.16
47.48 / 4.21
Start
OPD-r
27.00 / 7.39
39.04 / 7.28
52.27 / 2.65
+ filtering
OPD-r (A ∪ C)
25.66 / 9.02
37.51 / 9.12
47.70 / 7.87
Start
OPD-f
26.57 / 8.22
39.25 / 7.15
52.83 / 1.98
+ filtering
OPD-f (A ∪ C)
24.87 / 10.17
36.52 / 10.34
47.19 / 8.45
Table 3: Adding one component at a time: macro avg@ K (%) / Gtransfer (points), GRPO-8B teacher. All A ∪ C rows use the same problems and steps.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Student
base-8B teacher
GRPO-8B teacher
0.6B
0.078
0.079
1.7B
0.058
0.063
4B
0.053
0.054
Appendix
Table 4: Entropy-brake coefficient β for each teacher–student pair.
Shared
Optimizer
Adam ( β1=0.9 , β2=0.98 ), weight decay 0.1
Learning rate
1×10−6 , constant
Responses per update
64
Rollout
temperature 1.0, max response length 8,192
Policy clipping
ϵlow=0.2 , ϵhigh=0.28
KL / entropy regularizer
none (coefficients 0)
Appendix
Table 5: Training hyperparameters. Shared settings apply to all runs.
Arm
Training problems
Loss
Direct GRPO
all
correctness reward, no teacher
Direct GRPO (A ∪ C)
A ∪ C
same
OPD-r
all
reverse KL
OPD-f
all
forward KL
OPD-r (A ∪ C)
A ∪ C
reverse KL
OPD-f (A ∪ C)
A ∪ C
forward KL
Appendix
Table 6: Training arms. All arms except the two Direct GRPO rows use the GRPO-8B teacher; Table 2 also reports the four OPD objectives with the base-8B teacher.
Setting
avg@ K
pass@ K
τ=0.60
27.17
41.45
τ=0.95
24.71
36.41
20% of positions by student entropy
26.28
40.54
20% of positions at random
26.36
41.34
mass@n>0.0
26.70
39.93
mass@n>0.5
26.75
42.59
Appendix
Table 7: Ablations with the 0.6B student and the GRPO-8B teacher (macro averages over eight benchmarks, %). Each row changes one setting of the default TS-OPD.
Figure 4: Outcome categories of all 1,566 test problems after training. Students are compared with S0 ; the first row shows the teacher from T0 to T1 .
0.6B
1.7B
4B
Method
A
C
P
N
A
C
P
N
A
C
P
N
Teacher T0→T1
8.4
27.8
5.8
58.0
8.4
27.8
5.8
58.0
8.4
27.8
5.8
58.0
Direct GRPO
14.3
27.3
7.9
50.6
10.7
28.9
6.9
53.6
7.5
22.8
9.8
59.8
Direct GRPO (A ∪ C)
13.5
27.0
8.0
51.5
10.2
28.3
7.2
54.3
7.0
22.3
9.8
60.9
OPD-r
14.6
26.8
8.6
50.0
9.6
25.0
11.7
53.7
7.9
24.3
7.2
60.7
OPD-r (A ∪ C)
15.8
28.7
7.3
48.1
10.0
27.3
9.1
53.6
7.0
24.1
7.9
61.0
Appendix
Table 8: Outcome categories of all 1,566 benchmark problems after training (%), relative to the raw student S0 ; the teacher row is T0→T1 .
0.6B
1.7B
4B
Method
A
C
P
N
A
C
P
N
A
C
P
N
Direct GRPO
24.3
16.9
8.8
49.9
23.3
41.4
6.7
28.6
17.3
45.1
10.8
26.8
Direct GRPO (A ∪ C)
21.0
16.1
9.0
54.0
20.5
39.0
7.6
33.0
14.5
43.6
11.0
31.0
OPD-r
25.0
16.2
7.9
50.8
21.7
38.9
8.6
30.8
18.7
47.6
7.4
26.3
OPD-r (A ∪ C)
28.0
18.0
7.9
46.0
23.3
39.9
8.6
28.2
15.9
47.4
8.5
28.2
OPD-f
29.1
20.3
6.5
44.1
22.4
35.6
11.3
30.7
18.2
46.9
8.8
26.1
Appendix
Table 9: Outcome categories, in the student, of the 567 problems that GRPO improves for the teacher (%).
0.6B
1.7B
4B
Method
A
C
P
N
A
C
P
N
A
C
P
N
P stratum: S0 more accurate than T1
Direct GRPO
-
12.2
65.9
22.0
-
20.0
49.2
30.8
-
6.8
72.8
20.4
Direct GRPO (A ∪ C)
-
14.6
63.4
22.0
-
18.5
50.8
30.8
-
7.8
71.8
20.4
OPD-r
-
14.6
78.0
7.3
-
16.9
70.8
12.3
-
13.6
58.3
28.2
OPD-r (A ∪ C)
-
17.1
73.2
9.8
-
13.8
64.6
21.5
-
12.6
68.9
18.4
Appendix
Table 10: Outcome categories (%) of test problems where the teacher is not better (Figure 3 (c, d)). P stratum: S0 more accurate than T1 (41, 65, and 103 problems for 0.6B, 1.7B, and 4B). N stratum: equally accurate (453, 711, and 909 problems). Each row sums to 100 within a size; “-” marks a category that cannot occur. All arms except the two Direct GRPO rows use the GRPO-8B teacher.
0.6B
1.7B
4B
Method
A
C
P
N
A
C
P
N
A
C
P
N
Knowledge-gap problems: S0 at 0, T1 above 0
Direct GRPO
41.0
-
-
59.0
54.9
-
-
45.1
67.8
-
-
32.2
Direct GRPO (A ∪ C)
38.1
-
-
61.9
51.9
-
-
48.1
59.9
-
-
40.1
OPD-r
41.6
-
-
58.4
52.7
-
-
47.3
71.1
-
-
28.9
OPD-r (A ∪ C)
40.2
-
-
59.8
50.4
-
-
49.6
59.2
-
-
40.8
Appendix
Table 11: Outcome categories (%) of knowledge-gap and precision-gap test problems (Figure 3 (b)). All arms except the two Direct GRPO rows use the GRPO-8B teacher; the A ∪ C rows share the same problems and steps. Knowledge-gap problems start at 0, so they can only be A or N ; precision-gap problems start above 0, so they can only be C , P , or N . There are 517 and 555 such problems for 0.6B, 264 and 526 for 1.7B, and 152 and 402 for 4B.
base-8B teacher
GRPO-8B teacher
No teacher
Student
Benchmark
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
Direct GRPO
0.6B
MATH500
-3.05
-1.45
-5.92
-6.64
-14.42
-12.17
-15.22
-18.07
-12.55
Minerva
6.20
5.87
5.32
4.31
-1.52
-0.74
-0.47
-0.79
-0.38
OlympiadBench
4.63
5.58
3.11
2.26
-4.62
-0.99
-4.14
-4.94
-1.12
AIME24
30.62
28.75
29.58
27.08
24.16
23.54
21.46
22.50
24.58
AIME25
26.05
26.46
23.96
22.92
15.00
14.17
13.55
12.09
14.79
Appendix
Table 12: Transfer gap Gtransfer (percentage points) on each benchmark with K≥8 . The last row per student is the macro average used in Table 2 .
Teacher: base-8B
Teacher: GRPO-8B
Benchmark
Metric
Direct GRPO
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
MATH500
Avg@8
62.58
53.08
51.48
55.95
56.67
64.45
62.20
65.25
68.10
Pass@8
82.00
78.80
78.00
82.80
81.60
85.00
82.60
84.40
85.40
Minerva
Avg@8
16.87
10.29
10.62
11.17
12.18
18.01
17.23
16.96
17.28
Pass@8
30.88
25.74
27.21
28.31
26.84
33.46
34.93
32.72
32.35
OlympiadBench
Avg@8
28.86
23.11
22.16
24.63
25.48
32.36
28.73
31.88
32.68
Appendix
Table 13: Per-benchmark results for the 0.6B student (%). Direct GRPO uses no teacher. Macro averages weight each benchmark equally.
Teacher: base-8B
Teacher: GRPO-8B
Benchmark
Metric
Direct GRPO
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
MATH500
Avg@8
81.38
71.38
68.17
71.17
72.10
79.92
78.80
79.75
80.25
Pass@8
93.00
90.40
89.60
90.40
89.80
91.80
91.20
92.00
92.00
Minerva
Avg@8
29.00
20.36
18.29
21.00
22.01
27.34
27.85
27.67
28.31
Pass@8
44.49
37.50
35.66
37.50
37.13
42.28
42.65
44.12
44.12
OlympiadBench
Avg@8
48.94
36.20
33.77
37.11
37.22
45.72
44.25
46.07
46.42
Appendix
Table 14: Per-benchmark results for the 1.7B student (%). Direct GRPO uses no teacher. Macro averages weight each benchmark equally.
Teacher: base-8B
Teacher: GRPO-8B
Benchmark
Metric
Direct GRPO
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
OPD-r
OPD-f
TS-OPD (routing only)
TS-OPD
MATH500
Avg@8
88.47
80.12
78.95
80.15
80.88
88.97
88.47
89.40
90.03
Pass@8
94.80
93.60
93.40
93.40
93.20
95.40
94.60
95.60
95.20
Minerva
Avg@8
35.80
29.32
29.73
29.55
29.41
36.67
35.39
36.12
39.20
Pass@8
48.53
43.75
44.49
43.38
41.54
48.53
51.84
47.79
52.21
OlympiadBench
Avg@8
55.10
45.16
43.90
46.07
46.88
56.84
56.66
56.94
57.14
Appendix
Table 15: Per-benchmark results for the 4B student (%). Direct GRPO uses no teacher. Macro averages weight each benchmark equally.
Benchmark
Metric
Qwen3-0.6B
Qwen3-1.7B
Qwen3-4B
base-8B
GRPO-8B
MATH500
Avg@8
42.23
70.00
81.45
82.30
90.10
Pass@8
72.60
89.40
93.40
93.60
95.40
Minerva
Avg@8
8.64
20.96
29.46
30.06
37.91
Pass@8
25.74
35.66
41.54
42.65
50.00
OlympiadBench
Avg@8
15.37
34.92
47.94
47.59
59.96
Pass@8
37.09
56.82
66.02
67.21
72.26
Appendix
Table 16: Per-benchmark results of the raw students and the two teachers (%).