On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the student, such positive guidance can provide insufficient learning signals. In this work, we introduce Negative-Policy OPD (NP-OPD), which complements teacher supervision with rollouts from a lower-performing, lower-capability negative policy that serves as a negative reference for the student. Rather than modifying the distillation reward formulation, NP-OPD introduces the negative policy at the rollout stage, continuously supplying tokens preferred by the negative policy over the teacher so that they remain exposed to teacher supervision throughout training. This provides an explicit negative signal through negative-policy rollouts while preserving the positive teacher supervision used in OPD. Through extensive experiments, we show that NP-OPD improves OPD across model scales, generation modes, reasoning domains, and different OPD variants. Furthermore, our analyses show that NP-OPD effectively suppresses tokens preferred by the negative policy over the teacher and moves the student away from the negative policy. These results support our design of introducing negative signals through negative-policy rollouts and provide new insight into the role of the rollout policy in OPD. Code will be available at https://github.com/naver-ai/np-opd.
Figures & tables
Figure 1: Conceptual comparison of preference optimization, on-policy distillation (OPD), and Negative-Policy OPD (NP-OPD). (a) Preference optimization provides both toward and away directions through preferred and non-preferred responses, whereas (b) OPD provides an explicit positive direction through teacher supervision but no explicit negative reference indicating what the student should move away from. (c) NP-OPD complements the positive signal from teacher supervision with rollouts from a lower-performing, lower-capability negative policy that serves as a negative reference. Here, “negative” refers to the role of the negative policy as a reference from which the student is intended to move away. In our main setting, the negative policy is selected from the same model family as the student with lower model capacity and overall reasoning performance.
Math
Model
AIME24
AIME25
AMC23
HMMT25
MATH500
Olympiad
RGMath
Avg
Qwen3-1.7B
11.5
10.0
45.9
5.4
69.0
25.9
82.5
35.7
+ OPD
31.7
21.7
65.3
13.3
80.4
36.0
89.9
48.3
+ NP-OPD
43.1
33.3
78.4
20.0
86.0
40.6
92.3
56.3
Qwen3-4B
24.2
20.0
69.4
11.7
79.0
34.7
88.9
46.8
+ OPD
53.8
42.3
90.0
26.2
88.2
43.7
90.9
62.2
Table 1: Results of Qwen3 models on math, code, and science benchmarks (non-thinking mode). The best and second-best results are shown in bold and underlined, respectively.
Math
Code
Science
All
-
NP-OPD
Δ
-
NP-OPD
Δ
-
NP-OPD
Δ
-
NP-OPD
Δ
Qwen3-1.7B
OPD
60.0
61.8
+1.83
34.3
37.6
+3.24
40.9
41.9
+1.06
49.7
51.6
+1.98
ExOPD
63.1
63.4
+0.27
38.2
39.0
+0.77
43.4
43.2
-0.12
52.8
53.1
+0.30
OPD 2
64.3
65.3
+1.00
38.5
43.0
+4.48
44.2
45.0
+0.83
53.7
55.5
+1.76
Qwen3-4B
Table 2: Effect of NP-OPD across OPD variants (thinking mode). Results where NP-OPD improves over the on-policy baseline are shown in bold. “All” denotes the average across all 13 benchmarks.
Math
Model
AIME24
AIME25
AMC23
HMMT25
MATH500
Olympiad
RGMath
Avg
Gemma-4-E4B
55.0
35.0
88.8
28.7
88.0
42.6
89.7
61.1
+ OPD
57.9
42.1
81.2
35.8
87.4
42.1
94.1
63.0
+ NP-OPD
61.9
45.6
91.6
37.5
89.8
44.7
93.9
66.4
Table 3: Results of Gemma-4-E4B models on math, code, and science benchmarks (thinking mode). The best and second-best results are shown in bold and underlined, respectively.
Figure 5Figure 6
Rollout policy
AIME24
AIME25
GPQA
LCB
AVG
On-policy
50.0
35.6
49.1
36.9
42.9
+ high temp.
52.7
37.1
48.3
37.1
43.8
+ persona prompt
52.1
36.5
51.2
35.9
43.9
+ layer drop
49.0
35.8
51.0
32.5
42.1
NP-OPD
54.2
38.5
51.4
41.0
46.3
Table 4: Performance with other degraded rollouts.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Qwen3-1.7B
Qwen3-4B
Qwen3-8B
Model
Math
Code
Sci.
Math
Code
Sci.
Math
Code
Sci.
Baseline
35.7
12.3
30.8
46.8
24.9
40.4
45.7
29.8
43.0
OPD
48.3
27.2
37.2
62.2
40.2
47.8
65.1
44.0
50.1
+ NP-OPD
56.3
31.1
41.5
70.5
51.1
51.2
69.4
44.3
52.6
ExOPD
53.1
30.4
38.2
65.8
44.5
48.6
68.5
45.6
51.6
+ NP-OPD
55.5
30.2
41.5
71.3
51.2
51.5
71.7
47.2
54.2
Appendix
Table C.1: Summary of Qwen3 results in non-thinking mode. The table reports average performance across Math, Code, and Science benchmarks. The best and second-best results are shown in bold and underlined, respectively.
Model
AIME24
AIME25
AMC23
HMMT25
MATH500
Olympiad
RGMath
Avg
Qwen3-1.7B
11.5
10.0
45.9
5.4
69.0
25.9
82.5
35.7
OPD
31.7
21.7
65.3
13.3
80.4
36.0
89.9
48.3
+ NP-OPD ( α =0.25)
32.3
25.6
72.8
15.4
82.8
36.8
91.1
51.0
+ NP-OPD ( α =0.5)
34.6
26.0
70.0
15.8
83.2
37.6
90.3
51.1
+ NP-OPD ( α =0.75)
38.5
26.7
74.1
17.9
83.4
38.6
92.0
53.0
+ NP-OPD ( α =1)
43.1
33.3
78.4
20.0
86.0
40.6
92.3
56.3
Appendix
Table C.2: Full results of Qwen3 models on mathematical reasoning benchmarks with OPD (non-thinking mode). The best and second-best results among the OPD variants are shown in bold and underlined, respectively.
Code
Science
Model
Code Forces
LCBv5
RG Algo
Avg
GPQA
Super GPQA
Sci Bench
Avg
Qwen3-1.7B
9.2
8.3
19.5
12.3
40.9
23.8
27.6
30.8
OPD
19.1
29.9
32.7
27.2
48.9
26.2
36.5
37.2
+ NP-OPD ( α =0.25)
19.4
30.5
34.1
28.0
49.6
26.5
37.5
37.8
+ NP-OPD ( α =0.5)
16.8
30.5
34.0
27.1
51.3
27.9
36.9
38.7
+ NP-OPD ( α =0.75)
18.8
31.9
37.7
29.4
51.8
27.8
39.4
39.7
Appendix
Table C.3: Full results of Qwen3 models on code and science benchmarks with OPD (non-thinking mode). The best and second-best results among the OPD variants are shown in bold and underlined, respectively. Avg is the per-domain average.
Model
AIME24
AIME25
AMC23
HMMT25
MATH500
Olympiad
RGMath
Avg
Qwen3-1.7B
11.5
10.0
45.9
5.4
69.0
25.9
82.5
35.7
ExOPD
41.2
26.7
75.0
18.3
82.0
37.7
90.8
53.1
+ NP-OPD ( α =0.5)
36.9
27.5
73.8
17.5
82.8
37.8
90.1
52.3
+ NP-OPD ( α =1)
41.2
31.5
80.9
19.6
83.2
39.6
92.2
55.5
Qwen3-4B
24.2
20.0
69.4
11.7
79.0
34.7
88.9
46.8
ExOPD
59.6
51.0
93.8
30.4
89.6
44.9
91.4
65.8
Appendix
Table C.4: Full results of Qwen3 models on mathematical reasoning benchmarks with ExOPD (non-thinking mode). The best and second-best results among the OPD variants are shown in bold and underlined, respectively.
Code
Science
Model
Code Forces
LCBv5
RG Algo
Avg
GPQA
Super GPQA
Sci Bench
Avg
Qwen3-1.7B
9.2
8.3
19.5
12.3
40.9
23.8
27.6
30.8
ExOPD
21.0
33.3
36.9
30.4
47.9
29.4
37.5
38.2
+ NP-OPD ( α =0.5)
15.2
34.7
37.0
29.0
52.5
27.9
38.0
39.4
+ NP-OPD ( α =1)
13.5
32.5
44.6
30.2
56.4
29.0
39.2
41.5
Qwen3-4B
16.1
26.6
32.0
24.9
43.1
31.5
46.6
40.4
Appendix
Table C.5: Full results of Qwen3 models on code and science benchmarks with ExOPD (non-thinking mode). The best and second-best results among the OPD variants are shown in bold and underlined, respectively. Avg is the per-domain average.
Model
AIME24
AIME25
AMC23
HMMT25
MATH500
Olympiad
RGMath
Avg
Qwen3-1.7B
11.5
10.0
45.9
5.4
69.0
25.9
82.5
35.7
OPD 2
39.2
28.1
76.9
13.8
84.2
38.7
90.9
53.1
+ NP-OPD ( α =0.5)
39.2
30.0
77.2
19.2
85.2
39.8
92.3
54.7
+ NP-OPD ( α =1)
38.1
31.0
76.6
18.8
85.4
39.5
91.1
54.4
Qwen3-4B
24.2
20.0
69.4
11.7
79.0
34.7
88.9
46.8
OPD 2
60.8
50.4
92.8
31.7
91.0
45.4
92.3
66.3
Appendix
Table C.6: Full results of Qwen3 models on mathematical reasoning benchmarks with OPD 2 (non-thinking mode). The best and second-best results among the OPD variants are shown in bold and underlined, respectively.
Code
Science
Model
Code Forces
LCBv5
RG Algo
Avg
GPQA
Super GPQA
Sci Bench
Avg
Qwen3-1.7B
9.2
8.3
19.5
12.3
40.9
23.8
27.6
30.8
OPD 2
21.3
33.5
40.3
31.7
51.0
28.8
38.6
39.5
+ NP-OPD ( α =0.5)
19.6
36.0
40.2
31.9
60.3
28.7
40.4
43.1
+ NP-OPD ( α =1)
15.3
30.9
40.6
28.9
55.1
28.2
38.9
40.7
Qwen3-4B
16.1
26.6
32.0
24.9
43.1
31.5
46.6
40.4
Appendix
Table C.7: Full results of Qwen3 models on code and science benchmarks with OPD 2 (non-thinking mode). The best and second-best results among the OPD variants are shown in bold and underlined, respectively. Avg is the per-domain average.
Qwen3-1.7B
Qwen3-4B
Qwen3-8B
Model
Math
Code
Sci.
Math
Code
Sci.
Math
Code
Sci.
Baseline
60.1
33.5
39.5
72.9
53.2
51.8
73.3
52.7
53.6
OPD
60.0
34.3
40.9
70.7
52.9
49.8
71.7
56.4
52.2
+ NP-OPD
61.8
37.6
41.9
72.2
53.4
50.6
74.3
54.0
53.1
ExOPD
63.1
38.2
43.4
72.8
58.7
51.5
74.2
60.3
53.6
+ NP-OPD
63.4
39.0
43.2
76.0
57.6
53.2
76.1
59.6
55.9
Appendix
Table C.8: Summary of Qwen3 results in thinking mode. The table reports average performance across Math, Code, and Science benchmarks. The best and second-best results are shown in bold and underlined, respectively.
Model
AIME24
AIME25
AMC23
HMMT25
MATH500
Olympiad
RGMath
Avg
Qwen3-1.7B
53.1
35.0
85.0
22.9
87.2
39.7
97.8
60.1
OPD
50.0
35.6
84.4
25.4
87.4
39.8
97.5
60.0
+ NP-OPD ( α =0.25)
51.0
36.5
84.1
25.4
86.4
40.5
98.1
60.3
+ NP-OPD ( α =0.5)
51.5
39.2
82.8
25.4
86.6
40.8
96.1
60.3
+ NP-OPD ( α =0.75)
52.9
36.5
84.4
26.7
87.0
40.7
97.0
60.7
+ NP-OPD ( α =1)
54.2
38.5
86.6
27.1
87.4
41.3
97.9
61.8
Appendix
Table C.9: Full results of Qwen3 models on mathematical reasoning benchmarks with OPD (thinking mode). The best and second-best results among the OPD variants are shown in bold and underlined, respectively.
Code
Science
Model
Code Forces
LCBv5
RG Algo
Avg
GPQA
Super GPQA
Sci Bench
Avg
Qwen3-1.7B
22.5
33.5
44.6
33.5
47.5
28.7
42.4
39.5
OPD
22.5
36.9
43.6
34.3
49.1
29.8
43.8
40.9
+ NP-OPD ( α =0.25)
25.3
38.5
42.9
35.6
49.2
27.6
43.4
40.1
+ NP-OPD ( α =0.5)
25.6
39.2
45.8
36.9
49.2
28.7
43.7
40.5
+ NP-OPD ( α =0.75)
24.7
37.3
46.2
36.1
50.7
28.8
43.9
41.1
Appendix
Table C.10: Full results of Qwen3 models on code and science benchmarks with OPD (thinking mode). The best and second-best results among the OPD variants are shown in bold and underlined, respectively.
Model
AIME24
AIME25
AMC23
HMMT25
MATH500
Olympiad
RGMath
Avg
Qwen3-1.7B
53.1
35.0
85.0
22.9
87.2
39.7
97.8
60.1
ExOPD
59.0
40.6
88.4
25.8
87.6
42.0
98.5
63.1
+ NP-OPD ( α =0.5)
56.7
40.8
87.5
32.1
87.4
41.6
98.2
63.5
+ NP-OPD ( α =1)
57.9
41.5
88.8
27.5
87.8
41.8
98.6
63.4
Qwen3-4B
72.9
61.9
95.9
44.2
91.2
45.1
99.1
72.9
ExOPD
69.6
63.3
97.8
41.7
90.6
47.2
99.7
72.8
Appendix
Table C.11: Full results of Qwen3 models on mathematical reasoning benchmarks with ExOPD (thinking mode). The best and second-best results among the OPD variants are shown in bold and underlined, respectively.
Code
Science
Model
Code Forces
LCBv5
RG Algo
Avg
GPQA
Super GPQA
Sci Bench
Avg
Qwen3-1.7B
22.5
33.5
44.6
33.5
47.5
28.7
42.4
39.5
ExOPD
26.2
40.6
47.9
38.2
54.5
30.4
45.2
43.4
+ NP-OPD ( α =0.5)
27.7
39.7
50.8
39.4
52.8
30.6
44.5
42.6
+ NP-OPD ( α =1)
27.5
40.3
49.2
39.0
55.7
29.3
44.7
43.2
Qwen3-4B
39.7
58.9
60.8
53.2
57.0
38.5
60.0
51.8
Appendix
Table C.12: Full results of Qwen3 models on code and science benchmarks with ExOPD (thinking mode). The best and second-best results among the OPD variants are shown in bold and underlined, respectively.
Model
AIME24
AIME25
AMC23
HMMT25
MATH500
Olympiad
RGMath
Avg
Qwen3-1.7B
53.1
35.0
85.0
22.9
87.2
39.7
97.8
60.1
OPD 2
62.7
40.0
88.8
29.2
87.8
42.6
99.0
64.3
+ NP-OPD ( α =0.5)
61.9
43.8
88.8
32.1
88.6
42.5
99.5
65.3
+ NP-OPD ( α =1)
60.8
42.9
90.6
30.4
88.2
42.8
98.7
64.9
Qwen3-4B
72.9
61.9
95.9
44.2
91.2
45.1
99.1
72.9
OPD 2
70.6
64.2
96.2
48.3
91.6
47.8
99.2
74.0
Appendix
Table C.13: Full results of Qwen3 models on mathematical reasoning benchmarks with OPD 2 (thinking mode). The best and second-best results among the OPD variants are shown in bold and underlined, respectively.
Code
Science
Model
Code Forces
LCBv5
RG Algo
Avg
GPQA
Super GPQA
Sci Bench
Avg
Qwen3-1.7B
22.5
33.5
44.6
33.5
47.5
28.7
42.4
39.5
OPD 2
26.3
37.7
51.5
38.5
58.5
29.6
44.5
44.2
+ NP-OPD ( α =0.5)
30.2
43.9
54.8
43.0
60.2
30.2
44.7
45.0
+ NP-OPD ( α =1)
28.0
42.5
52.8
41.1
58.3
29.4
45.9
44.6
Qwen3-4B
39.7
58.9
60.8
53.2
57.0
38.5
60.0
51.8
Appendix
Table C.14: Full results of Qwen3 models on code and science benchmarks with OPD 2 (thinking mode). The best and second-best results among the OPD variants are shown in bold and underlined, respectively.
Proportion of negative-policy rollouts ( α )
Mode
-
0
0.25
0.5
0.75
1
Thinking
OPD
9h 19m
8h 09m (1.14 × )
6h 57m (1.34 × )
6h 04m (1.54 × )
3h 34m (2.61 × )
Non-thinking
OPD
5h 15m
5h 11m (1.01 × )
4h 34m (1.15 × )
3h 50m (1.37 × )
1h 16m (4.14 × )
Appendix
Table C.15: Wall-clock training time of Qwen3-1.7B on a single node with 8 × H100 GPUs. α denotes the proportion of negative-policy rollouts. Gray numbers show the speedup relative to α=0 .
Benchmark
Method
Qwen3-1.7B
Qwen3-4B
Qwen3-8B
AIME24
Base
17.0
14.2
15.1
OPD
20.9→21.5
18.5→15.3
17.0→15.1
ExOPD
20.7→21.1
18.7→16.0
18.1→15.9
OPD 2
20.1→21.1
19.0→16.6
18.6→16.1
AIME25
Base
18.0
18.1
18.3
OPD
22.9→23.5
20.5→18.3
19.9→18.0
Appendix
Table C.16: Mean generation length (in thousands of tokens) in thinking mode. Entries for OPD variants show training with on-policy rollouts ( α =0.0) → NP-OPD with α =1.0.
Model
AIME24
AIME25
AMC23
HMMT25
MATH500
Olympiad
RGMath
Avg
Qwen3-8B
75.6
65.6
94.1
46.7
91.8
39.7
99.4
73.3
ExOPD
76.5
63.3
96.6
44.2
92.2
47.4
99.3
74.2
+ NP-OPD ( α =1)
78.5
69.8
97.8
50.0
90.2
47.7
99.0
76.1
+ merge ( α =0 ⊕ 1)
80.0
67.9
97.8
49.6
91.8
47.9
99.3
76.3
Appendix
Table C.17: Model merging with ExOPD on Qwen3-8B (thinking mode). merge averages the weights of the ExOPD model ( α =0) and the NP-OPD ( α =1) model with equal weights. The best and second-best results among the ExOPD variants are shown in bold and underlined, respectively.
Code
Science
Model
Code Forces
LCBv5
RG Algo
Avg
GPQA
Super GPQA
Sci Bench
Avg
Qwen3-8B
40.1
50.9
67.1
52.7
63.2
39.5
58.0
53.6
ExOPD
47.2
63.6
70.1
60.3
63.9
41.0
56.0
53.6
+ NP-OPD ( α =1)
44.8
63.9
70.1
59.6
64.1
42.8
60.7
55.9
+ merge ( α =0 ⊕ 1)
46.6
65.0
70.7
60.7
63.6
41.4
59.3
54.7
Appendix
Table C.18: Model merging with ExOPD on Qwen3-8B: code and science benchmarks (thinking mode). merge averages the weights of the ExOPD model ( α =0) and the NP-OPD ( α =1) model with equal weights. The best and second-best results among the ExOPD variants are shown in bold and underlined, respectively.
Model
AIME24
AIME25
AMC23
HMMT25
MATH500
Olympiad
RGMath
Avg
Qwen3-8B
75.6
65.6
94.1
46.7
91.8
39.7
99.4
73.3
OPD 2
73.8
66.5
97.8
47.9
91.8
47.6
99.7
75.0
+ NP-OPD ( α =1)
81.5
71.2
97.8
52.9
90.6
48.5
99.3
77.4
+ merge ( α =0 ⊕ 1)
82.1
70.0
98.4
51.7
91.6
48.0
99.7
77.4
Appendix
Table C.19: Model merging with OPD 2 on Qwen3-8B: mathematical reasoning benchmarks (thinking mode). merge averages the weights of the OPD 2 model ( α =0) and the NP-OPD ( α =1) model with equal weights. The best and second-best results among the OPD 2 variants are shown in bold and underlined, respectively.
Code
Science
Model
Code Forces
LCBv5
RG Algo
Avg
GPQA
Super GPQA
Sci Bench
Avg
Qwen3-8B
40.1
50.9
67.1
52.7
63.2
39.5
58.0
53.6
OPD 2
46.5
65.5
72.1
61.4
62.3
41.3
53.8
52.5
+ NP-OPD ( α =1)
46.7
62.8
71.4
60.3
64.2
42.1
60.6
55.6
+ merge ( α =0 ⊕ 1)
47.8
65.0
74.0
62.3
63.4
41.7
57.4
54.2
Appendix
Table C.20: Model merging with OPD 2 on Qwen3-8B: code and science benchmarks (thinking mode). merge averages the weights of the OPD 2 model ( α =0) and the NP-OPD ( α =1) model with equal weights. The best and second-best results among the OPD 2 variants are shown in bold and underlined, respectively.
Figure C.1: Reduction in overlap across rollout policies. Same quantity as Fig. 5 , with the rollout policy varied over the Qwen3 family: (a) 0.6B (negative policy), (b) 1.7B, (c) 4B, and (d) 8B (teacher). Each point compares trajectory-level NOR under OPD (y-axis) with NOR under the corresponding rollouts (x-axis). Only 0.6B rollouts place most trajectories below the diagonal and the paired gap decreases monotonically with the size of the rollout policy.
Figure C.2: Reduction in overlap with the negative policy on held-out prompts. Same quantity as Fig. 5 , measured on 300 held-out prompts drawn from nine benchmarks spanning mathematics, science, and code, none of which appear in training. Each point is one trajectory, comparing NOR under OPD (y-axis) with (a) negative-policy or (b) teacher rollouts (x-axis).
On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, remains unclear. We quantitatively analyze teacher supervision during OPD training and find substantial noise whose prevalence increases with teacher scale. Surprisingly, the student policy is insensitive to such noise, converging to comparable performance regardless of whether noisy supervision is retained or removed. Does OPD distill at all? By analyzing what drives its gains, we find that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones. This suggests that OPD works largely by suppressing low log-probability tokens, which requires no teacher. These findings motivate On-Policy Self-Adaptation (OPSA), a supervision-free method using entropy-adaptive negative advantages. It assigns stronger learning signals to high-entropy positions, suppressing tail tokens, and evenly redistributing probability mass among head tokens. Compared with the base \texttt{Qwen3-1.7B}, OPSA improves Avg@32 by 35.41 points on AIME24, corresponding to a 263% relative gain, and more than doubles Pass@32 across all three benchmarks. It also outperforms OPD by 16.77 points in Avg@32 on AIME24. Extensive experiments and analyses across model families and tasks further demonstrate its effectiveness and generalizability.
Yi Ding, Ruqi Zhang
Department of Computer Science, Purdue University, USA
On-policy distillation is a powerful way to transfer reasoning ability from a strong teacher to a smaller student: the student samples trajectories from its own policy, and the teacher provides dense token-level supervision on the states the student actually visits. However, this supervision is not always reliable: a teacher can assign high likelihood to plausible but incorrect solutions, or low likelihood to correct student solutions that follow different reasoning paths. Unconditionally distilling the teacher can therefore reinforce bad modes or erase useful student behavior. To address these limitations, we introduce RG-OPD: Reward-Gated On-Policy Distillation that uses verifier feedback to decide when teacher logits should be trusted. RG-OPD bridges sparse verifier rewards and dense teacher logits, preserving token-level supervision while filtering misleading teacher signals. Across reasoning and coding benchmarks, RG-OPD produces stronger distilled students, outperforming both vanilla reverse-KL distillation and the recent TSD-KD baseline. At 1K generation length, RG-OPD improves over reverse-KL by 2.9 points and over TSD-KD by 4.9 points; in the long-generation setting, it improves over the untuned student by 8.2 points. Our code is available at https://github.com/UoC-tail/RG-OPD.
Mohammad Sadegh Akhondzadeh, Vijay Lingam, Atula Tejaswi +3
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.
Langlin Huang, Hao Liu, Mononito Goswami +5
Washington University in St. Louis · AWS AI Labs · Carnegie Mellon University +1