DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation
Authors: Ao Yu, Weibo Gao, Heng Zhou, Linan Yue, Rui Li, Suyi Liu, Yu Yan, Yizhong Zhang, +1 more
Organizations: University of Science and Technology of China · The Hong Kong Polytechnic University · The University of Hong Kong · Southeast University
On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and, on average, pushes down even the student's correct responses; gating feedback by student correctness fixes the direction but uses the teacher in the same way whether or not it succeeded. We introduce DuoOPD, in which the student's outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it: when only the teacher succeeds, its verified answer becomes context for scoring the student's failed response, and when only the student succeeds, a weight shared within the task reinforces the whole response. A single rule covers all four outcome combinations without task-specific settings. Across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation. Ablations show that outcome-based direction alone stays near the gated baseline, while the joint-outcome designs supply most of the gain.
Figures & tables
Qwen3 4B → 0.6B
Llama 3.1-8B → 3.2-3B
Method
Bio.
Chem.
Phys.
Macro
Bio.
Chem.
Phys.
Macro
Student (initial)
37.38
38.88
33.38
36.54
36.63
42.44
44.94
41.33
Teacher
79.56
81.38
76.88
79.27
73.25
75.06
69.88
72.73
OPD
67.38
67.42
58.38
64.39
69.19
70.90
61.75
67.28
EOPD
68.06
68.44
60.92
65.81
68.73
69.54
61.21
66.49
OPDVR
66.92
66.69
60.90
64.83
69.75
72.40
63.90
68.68
Table 1: Results across model families. Test avg@8 accuracy (%); Llama models are Instruct variants. Bold and underlined scores mark the best and second-best distillation methods per column.
Science knowledge, understanding, calculation
Heterogeneous answers, instructions, code
Method
Mat.
Chem.
Phys.
Macro
Phys.
IFEval
MBPP
Macro
Student (initial)
32.63
56.44
24.38
37.81
32.25
54.99
20.69
35.98
Teacher
67.44
96.94
56.00
73.46
77.25
78.70
57.69
71.21
OPD
55.50
91.71
39.77
62.33
59.29
50.95
31.41
47.22
EOPD
55.00
92.75
36.94
61.56
59.56
50.25
30.88
46.90
OPDVR
53.98
92.83
39.67
62.16
59.35
52.31
31.83
47.83
Table 2: Results on two additional task mixtures with Qwen3-4B → Qwen3-0.6B. Left: materials knowledge, chemistry understanding, and physics calculation. Right: physics knowledge, IFEval (prompt-level strict), and MBPP. Protocol and marking follow Table 1 .
Teacher
Student
Feedback replacement
Share
Bio.
Chem.
Phys.
Macro
Δ
DuoOPD
—
100
68.60
70.19
62.10
66.97
0.00
✓
✓
sp(d0)→d0
56.9
67.40
68.63
61.73
65.92
− 1.05
✓
×
−sp(−dT)→d0
21.5
68.21
67.81
60.29
65.44
− 1.53
×
✓
mˉk→d0
6.9
68.38
68.81
61.44
66.21
− 0.76
×
×
−sp(−d0)→d0
14.7
68.27
68.79
60.67
65.91
− 1.06
Table 3: Replacing one outcome’s feedback with OPD. Share is the outcome’s fraction of DuoOPD training responses (%). Accuracy in %; in Tables 3 – 5 , bold and underlined values mark the best and second-best means. Δ is relative to DuoOPD.
Variant
T ✓ S ×
T × S ✓
Macro
Δ
DuoOPD
−sp(−dT)
mˉk
66.97
0.00
w/o sharing
−sp(−dT)
sp(d0)
66.26
− 0.71
w/o reference
−sp(−d0)
mˉk
65.48
− 1.49
w/o both
−sp(−d0)
sp(d0)
64.56
− 2.41
Table 4: Removing the two disagreement designs. Feedback used for each disagreement outcome; the agreement rules are unchanged.
Variant
T ✓ S ×
T × S ✓
Macro
Δ
DuoOPD
−sp(−dT)
mˉk
66.97
0.00
w/o sharing
−sp(−dT)
sp(d0)
66.26
− 0.71
w/o reference
−sp(−d0)
mˉk
65.48
− 1.49
w/o both
−sp(−d0)
sp(d0)
64.56
− 2.41
Table 4: Removing the two disagreement designs. Feedback used for each disagreement outcome; the agreement rules are unchanged.
Shared
B/C/P
Sci.
Het.
Across tasks, mˉ
66.74
63.22
48.89
Within task
66.97
63.48
49.24
Table 5: Sharing scope. Macro (%) with the student-success weight shared across or within tasks.
One response per question; temperature 1; top- p=1 ; unrestricted top- k
Prompt / response limits
4,096 / 1,024 tokens for the student
Appendix
Table 6: Settings for the initial biology, chemistry, and physics comparison. Additional-mixture differences are specified in Appendices A.2 and A.3 .
Quantity
Qwen
Llama
Cache generated / student-sampled tokens per run
1.01–1.10 ×
0.76–0.90 ×
Responses scored with a reference per run
20.7–22.6%
13.3–14.6%
Extra teacher scoring input per run
14.16–15.45%
6.21–6.39%
Scoring/verification share of OPD training loop
5.9%
4.7%
Training-loop change, cache excluded
− 0.7%
+7.3%
Cache generation / one OPD training loop
29.7%
23.0%
Appendix
Table 7: DuoOPD training and teacher-cache cost relative to OPD. The first three rows give per-run ranges; rows with cache compare combined GPU-time against OPD’s training-loop GPU-time, amortizing one cache over R runs.
Method
Qwen3, B/C/P
Llama, B/C/P
Qwen3, science
Qwen3, heterogeneous
OPD
64.39 ± 0.45
67.28 ± 1.25
62.33 ± 0.39
47.22 ± 0.36
EOPD
65.81 ± 0.54
66.49 ± 0.52
61.56 ± 0.58
46.90 ± 0.36
OPDVR
64.83 ± 0.89
68.68 ± 1.37
62.16 ± 0.17
47.83 ± 0.35
ExOPD
64.42 ± 0.37
66.34 ± 0.58
62.62 ± 0.01
46.83 ± 0.08
FiRe-OPD
64.58 ± 0.39
67.64 ± 0.79
61.72 ± 0.22
47.53 ± 0.02
DuoOPD
66.97 ± 0.26
73.26 ± 1.04
63.48 ± 0.37
49.24 ± 0.64
Appendix
Table 8: Macro test avg@8 (%), mean ± standard deviation.
Mean d0
Gate
Softplus
Joint outcome
Qwen3
Llama
Qwen3
Llama
Qwen3
Llama
Both succeed
−1.31
−0.22
+0.09
+0.08
+0.57
+0.65
Only the student succeeds
−1.24
−0.25
+0.09
+0.09
+0.57
+0.64
Both fail
−1.26
−0.24
−1.34
−0.33
−1.82
−0.88
Appendix
Table 9: Mean token weight from d0 under OPDVR’s gate, σmax{σd0,0} , and under Softplus, σsp(σd0) , where σ=2rS−1 . The gate keeps little positive feedback on verified successes.
On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however, depends on how the student completes the subsequent reasoning. This mismatch can cause imitation to suppress viable student strategies or reinforce paths the student cannot reliably execute. Verified trajectory outcomes provide complementary evidence about continuation quality, but do not directly identify the utility of individual decisions. We introduce Reward-Aligned Reweighting for On-Policy Distillation (R2-OPD), which uses outcome agreement and the magnitude of teacher--student disagreement to continuously reallocate teacher supervision. It gives reward-aligned corrections greater relative influence while retaining dense feedback, moving beyond uniform imitation and hard filtering. Our analysis formalizes the mismatch between local teacher preference and student continuation value and establishes sufficient conditions for reallocation to improve first-order task progress over uniform OPD. Across seven mathematical reasoning benchmarks, R2-OPD achieves the highest average accuracy among the compared training methods in both cross-size and same-size distillation. It outperforms standard OPD on all seven benchmarks, with average gains of 3.5 and 2.4 percentage points for 1.7B and 4B students, respectively. An extension to code generation yields an average gain of 1.6 percentage points over standard OPD. These results highlight outcome-guided supervision allocation as an effective way to translate dense teacher feedback into stronger student performance across model scales and task domains.
Haofeng Xu, Junwei Su, Lansong Diao +2
The University of Hong Kong · Alibaba Group · University of Science and Technology of China
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student's performance: some correct important reasoning errors, while others have little effect on the final answer. Motivated by this observation, we introduce Dr. OPD (OPD Done Right), which defines the optimal weighted OPD to maximize the student's performance. We formulate Dr. OPD as a bilevel optimization problem in which the student learns from weighted teacher supervision, while the weights are selected to maximize the expected reward of the resulting student. To solve Dr. OPD, we develop an efficient iterative solver that updates the token weights and student policy alternatively. At each round, it updates weights in closed form and then takes one gradient step on the resulting weighted OPD objective. Under regularity conditions, we show that this weighted update achieves a higher expected reward than a vanilla OPD update. Empirically, across strong-to-weak and same-size distillation on math and code, Dr. OPD consistently outperforms all evaluated baselines. In particular, in the strong-to-weak distillation setting, Dr. OPD improves average math performance by 9.7 points over vanilla OPD, and enables the smaller student to surpass its larger teacher.
On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher is reliable for each prompt. Because reverse KL is mode-seeking, a confidently wrong teacher can induce a strong yet misleading update. Distributional proxies, such as entropy or teacher-student likelihood agreement, measure uncertainty or agreement but do not directly verify outcome correctness. We introduce Teacher-Gated On-Policy Distillation (TGOPD), built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted. TGOPD estimates reliability from a small set of verifier-scored teacher probes and routes each prompt exclusively to dense OPD when the reliability check passes or to verifier-grounded GRPO otherwise. Across 4B and 35B students in mathematics, code, and instruction following, TGOPD outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training. By using otherwise-idle teacher capacity for reliability estimation, TGOPD also reduces teacher-side compute waste in asynchronous OPD, increasing teacher-node GPU utilization from 9.8% to 78.9% in the measured 4B single-domain run.