DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation
Authors: Ao Yu, Weibo Gao, Heng Zhou, Linan Yue, Rui Li, Suyi Liu, Yu Yan, Yizhong Zhang, +1 more
Organizations: University of Science and Technology of China · The Hong Kong Polytechnic University · The University of Hong Kong · Southeast University
On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and, on average, pushes down even the student's correct responses; gating feedback by student correctness fixes the direction but uses the teacher in the same way whether or not it succeeded. We introduce DuoOPD, in which the student's outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it: when only the teacher succeeds, its verified answer becomes context for scoring the student's failed response, and when only the student succeeds, a weight shared within the task reinforces the whole response. A single rule covers all four outcome combinations without task-specific settings. Across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation. Ablations show that outcome-based direction alone stays near the gated baseline, while the joint-outcome designs supply most of the gain.
Figures & tables
Qwen3 4B → 0.6B
Llama 3.1-8B → 3.2-3B
Method
Bio.
Chem.
Phys.
Macro
Bio.
Chem.
Phys.
Macro
Student (initial)
37.38
38.88
33.38
36.54
36.63
42.44
44.94
41.33
Teacher
79.56
81.38
76.88
79.27
73.25
75.06
69.88
72.73
OPD
67.38
67.42
58.38
64.39
69.19
70.90
61.75
67.28
EOPD
68.06
68.44
60.92
65.81
68.73
69.54
61.21
66.49
OPDVR
66.92
66.69
60.90
64.83
69.75
72.40
63.90
68.68
Table 1: Results across model families. Test avg@8 accuracy (%); Llama models are Instruct variants. Bold and underlined scores mark the best and second-best distillation methods per column.
Science knowledge, understanding, calculation
Heterogeneous answers, instructions, code
Method
Mat.
Chem.
Phys.
Macro
Phys.
IFEval
MBPP
Macro
Student (initial)
32.63
56.44
24.38
37.81
32.25
54.99
20.69
35.98
Teacher
67.44
96.94
56.00
73.46
77.25
78.70
57.69
71.21
OPD
55.50
91.71
39.77
62.33
59.29
50.95
31.41
47.22
EOPD
55.00
92.75
36.94
61.56
59.56
50.25
30.88
46.90
OPDVR
53.98
92.83
39.67
62.16
59.35
52.31
31.83
47.83
Table 2: Results on two additional task mixtures with Qwen3-4B → Qwen3-0.6B. Left: materials knowledge, chemistry understanding, and physics calculation. Right: physics knowledge, IFEval (prompt-level strict), and MBPP. Protocol and marking follow Table 1 .
Teacher
Student
Feedback replacement
Share
Bio.
Chem.
Phys.
Macro
Δ
DuoOPD
—
100
68.60
70.19
62.10
66.97
0.00
✓
✓
sp(d0)→d0
56.9
67.40
68.63
61.73
65.92
− 1.05
✓
×
−sp(−dT)→d0
21.5
68.21
67.81
60.29
65.44
− 1.53
×
✓
mˉk→d0
6.9
68.38
68.81
61.44
66.21
− 0.76
×
×
−sp(−d0)→d0
14.7
68.27
68.79
60.67
65.91
− 1.06
Table 3: Replacing one outcome’s feedback with OPD. Share is the outcome’s fraction of DuoOPD training responses (%). Accuracy in %; in Tables 3 – 5 , bold and underlined values mark the best and second-best means. Δ is relative to DuoOPD.
Variant
T ✓ S ×
T × S ✓
Macro
Δ
DuoOPD
−sp(−dT)
mˉk
66.97
0.00
w/o sharing
−sp(−dT)
sp(d0)
66.26
− 0.71
w/o reference
−sp(−d0)
mˉk
65.48
− 1.49
w/o both
−sp(−d0)
sp(d0)
64.56
− 2.41
Table 4: Removing the two disagreement designs. Feedback used for each disagreement outcome; the agreement rules are unchanged.
Variant
T ✓ S ×
T × S ✓
Macro
Δ
DuoOPD
−sp(−dT)
mˉk
66.97
0.00
w/o sharing
−sp(−dT)
sp(d0)
66.26
− 0.71
w/o reference
−sp(−d0)
mˉk
65.48
− 1.49
w/o both
−sp(−d0)
sp(d0)
64.56
− 2.41
Table 4: Removing the two disagreement designs. Feedback used for each disagreement outcome; the agreement rules are unchanged.
Shared
B/C/P
Sci.
Het.
Across tasks, mˉ
66.74
63.22
48.89
Within task
66.97
63.48
49.24
Table 5: Sharing scope. Macro (%) with the student-success weight shared across or within tasks.
One response per question; temperature 1; top- p=1 ; unrestricted top- k
Prompt / response limits
4,096 / 1,024 tokens for the student
Appendix
Table 6: Settings for the initial biology, chemistry, and physics comparison. Additional-mixture differences are specified in Appendices A.2 and A.3 .
Quantity
Qwen
Llama
Cache generated / student-sampled tokens per run
1.01–1.10 ×
0.76–0.90 ×
Responses scored with a reference per run
20.7–22.6%
13.3–14.6%
Extra teacher scoring input per run
14.16–15.45%
6.21–6.39%
Scoring/verification share of OPD training loop
5.9%
4.7%
Training-loop change, cache excluded
− 0.7%
+7.3%
Cache generation / one OPD training loop
29.7%
23.0%
Appendix
Table 7: DuoOPD training and teacher-cache cost relative to OPD. The first three rows give per-run ranges; rows with cache compare combined GPU-time against OPD’s training-loop GPU-time, amortizing one cache over R runs.
Method
Qwen3, B/C/P
Llama, B/C/P
Qwen3, science
Qwen3, heterogeneous
OPD
64.39 ± 0.45
67.28 ± 1.25
62.33 ± 0.39
47.22 ± 0.36
EOPD
65.81 ± 0.54
66.49 ± 0.52
61.56 ± 0.58
46.90 ± 0.36
OPDVR
64.83 ± 0.89
68.68 ± 1.37
62.16 ± 0.17
47.83 ± 0.35
ExOPD
64.42 ± 0.37
66.34 ± 0.58
62.62 ± 0.01
46.83 ± 0.08
FiRe-OPD
64.58 ± 0.39
67.64 ± 0.79
61.72 ± 0.22
47.53 ± 0.02
DuoOPD
66.97 ± 0.26
73.26 ± 1.04
63.48 ± 0.37
49.24 ± 0.64
Appendix
Table 8: Macro test avg@8 (%), mean ± standard deviation.
Mean d0
Gate
Softplus
Joint outcome
Qwen3
Llama
Qwen3
Llama
Qwen3
Llama
Both succeed
−1.31
−0.22
+0.09
+0.08
+0.57
+0.65
Only the student succeeds
−1.24
−0.25
+0.09
+0.09
+0.57
+0.64
Both fail
−1.26
−0.24
−1.34
−0.33
−1.82
−0.88
Appendix
Table 9: Mean token weight from d0 under OPDVR’s gate, σmax{σd0,0} , and under Softplus, σsp(σd0) , where σ=2rS−1 . The gate keeps little positive feedback on verified successes.