On-policy distillation from multiple teachers combines expertise from different domains in a single student, but conflicting gradients can hinder this integration. Gradient corrections directly constrain parameter updates under plain SGD. With optimizers such as AdamW, however, momentum, adaptive scaling, and weight decay can turn a corrected gradient into an update that increases a domain loss to first order. To address this gap, we propose Update Projection for Multi-Teacher On-Policy Distillation (UP-MOPD). UP-MOPD lets the original mixed gradient update the optimizer state and generate a candidate displacement, then projects only violating candidates before they are committed to the parameters. The projection gives the unique feasible update closest to the candidate in Euclidean distance. In experiments combining medical and general domains, UP-MOPD improves IFEval-loose accuracy late in training by 2.96 points over vanilla M-OPD. It achieves an average score of 60.03 across eight metrics, compared with 59.00 for gradient projection and 59.15 for update rejection. On a public benchmark covering mathematics, code, and instruction following, it achieves the best average across six tasks (32.67), leads on LiveCodeBench v5, and ties for the best IFEval result.These results support projecting optimizer updates to reduce interference between domains.
Figures & tables
Figure 1: Overview of UP-MOPD. AdamW forms a candidate parameter displacement from the mixed distillation gradient. UP-MOPD evaluates its first-order effect on each active domain and minimally corrects a violating candidate before commit, while retaining the current-step optimizer state produced by the original mixture.
Figure 2: Medical performance and instruction following. Student points average the last three saved checkpoints of one training epoch. The arrow from M-OPD to UP-MOPD shows improved IFEval-loose accuracy at a similar HealthBench-Hard score.
Medical
General
Math
Method
HealthBench Hard
Med- MCQA
PubMed- QA
MMLU- Med
GPQA
IFEval
AIME24
AIME25
Mean
Qwen3-4B teacher
6.38
57.26
76.60
78.32
42.42
86.88
54.00
43.33
55.65
Medical Teacher
39.06
59.07
70.00
80.09
57.68
73.38
47.33
42.00
58.58
M-OPD
38.73
59.77
71.67
79.54
47.81
77.94
53.78
43.78
59.13
GP-MOPD
38.79
59.80
73.00
79.85
48.28
78.93
52.89
40.44
59.00
UP-MOPD
38.57
59.37
74.33
79.54
48.38
80.90
53.33
45.78
60.03
Table 1: Medical and general capability integration. Student results are averaged over the final three checkpoints. Mean averages the eight metrics. Bold indicates the best student result in each column.
Math
Code
Instruction Following
Overall
Method
AIME24
AIME25
Avg.
LCB5
LCB6
Avg.
IFEval
IFBench
Avg.
Total
MixSFT
13.33
17.78
15.56
15.57
17.71
16.64
69.44
16.00
42.72
24.97
RFT ∗
22.97
23.91
23.44
18.98
19.43
19.21
55.08
18.67
36.87
26.51
MixRL ∗
21.15
22.14
21.64
16.59
20.97
18.78
70.24
22.67
46.45
28.96
ParamMerge-Avg ∗
18.91
20.99
19.95
18.38
21.20
19.79
70.06
17.67
43.86
27.87
ParamMerge-TA ∗
21.93
22.76
22.34
21.74
23.14
22.44
71.53
21.67
46.60
30.46
Table 2: Results on the public three-domain benchmark under the Open-MOPD evaluation protocol. Domain averages cover two tasks, and Total averages all six tasks. ∗ marks results reported by Gao et al. (2026) ; unmarked rows are our reproductions or proposed methods. RouteOPD uses three independently distilled students. Bold indicates the best result in each column.
Medical
General
Math
Method
HealthBench Hard
Med- MCQA
PubMed- QA
MMLU- Med
GPQA
IFEval
AIME24
AIME25
Mean
M-OPD
38.73
59.77
71.67
79.54
47.81
77.94
53.78
43.78
59.13
M-OPD + Grad. Clip
38.31
59.42
73.47
80.04
48.35
73.20
53.78
45.78
59.04
Reweighting w/ rejection
39.11
60.00
73.27
79.94
45.25
76.77
52.89
44.89
59.02
Reweighting w/o rejection
38.16
59.49
72.87
79.77
47.44
78.99
54.67
45.11
59.56
Update Rejection
38.68
58.07
73.87
79.11
45.99
80.77
52.00
44.67
59.15
Table 3: Effect of intervention strategies in the medical–general setting, using the same evaluation protocol as Table 1 . Mean averages the eight metrics. Bold indicates the best result in each column.
Figure 3: Gradient alignment and update conflict. Each point is a training step, with g1⊤g2 on the horizontal axis and h1h2 on the vertical axis. A negative h1h2 means the candidate update is predicted to lower one domain’s loss and raise the other’s. Gradient alignment alone does not determine this outcome.
Figure 4: Capability and distillation dynamics. (a,b) HealthBench-Hard and IFEval-loose checkpoint scores; lines denote teacher or initial-student references and shading the late-training window. (c,d) Trailing 50-step means of domain-conditioned sampled reverse KL on each method’s responses; steps lacking a domain and incomplete initial windows are omitted.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Method
HB-Hard
IFEval
Mean
Tokens
Trunc. (%)
M-OPD
38.73
77.94
59.13
4,067
9.20
Update Rejection
38.68
80.77
59.15
4,071
9.48
GP-MOPD
38.79
78.93
59.00
4,056
9.63
UP-MOPD
38.57
80.90
60.03
4,089
9.20
Appendix
Table 4: Capability scores and generation behavior. HB-Hard denotes HealthBench-Hard; IFEval uses prompt-level loose accuracy. Mean averages eight benchmarks. Tokens and Trunc. are mean response length and truncation rate over the last 200 plotted training steps.
Figure 5: Offline projection of 645 rejected candidates. Cumulative distributions show relative correction size (left) and retained predicted medical loss decrease (right). Dashed lines mark medians.
Domain
Teacher suffix
Step
Prompt source
Prompts
Mathematics
RL-Math
100
DAPO-Math-17k
17,917
Code
RL-Code
180
DeepCoder-Preview-Dataset
23,667
Instruction following
RL-IF
2,440
Nemotron-IF-RL-46k
45,347
Total
86,931
Appendix
Table 5: Released teachers and training data in the public setting. Model suffixes follow the shared repository prefix given in the text. Teacher steps identify the checkpoints described in their model cards. The released prompt pool excludes code records identified as LiveCodeBench.
Hyperparameter
Value
Student
Qwen3-4B-Instruct-2507
Teachers
Qwen3-4B-Instruct-2507 and Medical Teacher; both frozen, with supervision routed by prompt domain
Training framework
Megatron
Training data
22,564 prompts: 17,398 mathematical prompts and 5,166 prompts assigned to the medical teacher; combined and randomly shuffled
Training duration
1 epoch; 1,410 steps
Batch size
16 prompts per step; 4 responses per prompt (64 samples)
Appendix
Table 6: Training hyperparameters for medical and general capability integration.
Hyperparameter
Value
Student
SmolLM3-3B initialized from MixSFT
Teachers
Frozen RL-Math, RL-Code, and RL-IF releases (Table 5 )
Training framework
verl; public data and checkpoints from Open-MOPD ( Gao et al., 2026 )
Training data
86,931 prompts mixed across the three domains
Training duration
1 epoch; 679 full batches
Batch size
128 prompts per step; mini-batch size 128; micro-batch size 1 per GPU; 8 GPUs
Appendix
Table 7: Training hyperparameters for the public mathematics, code, and instruction-following benchmark.
Benchmark
Token limit
Repetitions
Demonstrations
HealthBench-Hard
16,384
1
0
GPQA-Diamond, AIME24/25
8,192
5
0
IFEval, MMLU-Med
8,192
1
0
MedMCQA
32,768
1
5
PubMedQA
32,768
1
0
Appendix
Table 8: Generation settings for the medical evaluation. Limits count generated tokens. Repetitions are completions per question.
Method
Logged step time (s)
Actor training (s)
M-OPD
66.55±6.58
11.57±1.65
Update Rejection
92.58±9.45
37.49±5.71
GP-MOPD
92.94±9.69
38.00±5.77
UP-MOPD
94.76±10.08
39.69±5.82
Appendix
Table 9: Runtime in the medical–general setting on eight NVIDIA H200 GPUs. Values are means ± standard deviations in seconds per step, using the same 200-step late-training window for all four methods. Actor training is included in the logged step time.
Step
Math
Code
IF
Total
100
18.89
21.24
50.58
30.24
200
26.67
20.44
49.48
32.20
300
23.33
19.87
48.53
30.57
400
19.45
21.33
48.49
29.75
500
23.89
21.93
47.54
31.12
600
22.22
21.94
49.12
31.09
Appendix
Table 10: UP-MOPD training trajectory in the public three-domain setting.
Figure 6: UP-MOPD training trajectory in the public three-domain setting. Each panel shows a domain mean over two benchmarks. The vertical scales differ across panels.
Method
Retained
Lost
Newly passed
Accuracy (%)
M-OPD
400.7
69.3
21.0
77.94
UP-MOPD
419.7
50.3
18.0
80.90
Appendix
Table 11: IFEval retention across three late-training checkpoints. The initial student passes 470 of 541 prompts under loose scoring. Retained and lost prompts partition these 470 initial passes; newly passed prompts come from the 71 initial failures. Counts and accuracies are means over the three checkpoints.
Metric
M-OPD
UP-MOPD
Difference
95% CI
BH-adjusted p
HealthBench-Hard
38.51
38.63
+0.12
[−1.34,1.60]
1.0000
GPQA
49.80
48.18
−1.62
[−4.65,1.31]
0.5279
IFEval
77.82
80.96
+3.14
[0.37,5.91]
0.1454
MedMCQA
59.67
58.71
−0.96
[−1.77,−0.12]
0.1454
PubMedQA
72.00
74.40
+2.40
[−0.20,5.00]
0.2564
MMLU-Med
79.41
79.44
+0.03
[−1.34,1.44]
1.0000
Appendix
Table 12: Paired comparison at the middle late-training checkpoint. IFEval uses prompt-level loose accuracy. Differences are UP-MOPD minus M-OPD, in score points. Intervals are pointwise 95% paired bootstrap confidence intervals; p -values are corrected across all eight metrics.
Figure 7: Instruction-following behavior in the medical setting. (a) UP-MOPD minus M-OPD instruction-level loose accuracy for all ten evaluator-defined constraint categories; parentheses give instruction counts. (b) Prompt-level loose accuracy grouped by the number of constraints in a prompt; labels above the groups give the mean gain in percentage points. Faint points show three matched late-training checkpoints from one run per method; solid points show their arithmetic mean.
Constraint
M-OPD
UP-MOPD
Exactly three Markdown bullet points
Repeats the three-item list after “Final concise version with desert emphasis:”. Loose: fail.
Produces one three-item list. Loose: pass.
At most 100 words and at least two highlighted sections
115 words and two highlighted spans. Highlighting passes, but the length constraint fails. Loose: fail.
99 words and three highlighted spans. Both constraints pass. Loose: pass.
Appendix
Table 13: Examples of constraint compliance. The initial student passes both prompts; distilled outputs are from the middle late-training checkpoint. Counts are computed from complete responses; word counts follow IFEval’s regular-expression tokenizer.
Multi-teacher on-policy distillation (MOPD) combines independently developed domain teachers into a single student by distilling their predictions on student-generated samples. We study a setting where teachers share a reference model but undergo different post-training procedures, and find that MOPD can struggle to recover some teacher capabilities. Because distillation occurs on student-generated prefixes, the student initialization can strongly affect subsequent recovery. However, initial benchmark performance is not a reliable predictor of a good MOPD initialization. For example, merge initialization can start below SFT warm-up yet finish higher after MOPD. We further find that effective merging depends on both the relative teacher contributions and the overall merge scale, with some strong configurations lying outside the simplex of convex parameter averaging. Thus, selecting a good merge initialization requires evaluating not only its immediate performance but also the learning it enables under MOPD, making one-shot coefficient search difficult. We propose Iterative Merging for MOPD (IM-MOPD), which starts from a uniform merge and progressively adds task-vector increments for under-recovered domains during distillation. In a 5-domain setting, IM-MOPD achieves higher average normalized recovery than MOPD with either uniform merge initialization or SFT warm-up, showing that effective teacher contributions can be determined progressively during training.
Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one domain suppresses capabilities acquired from another. Inspired by the distinctive update geometry of OPD, we find that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. We therefore propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions. We further develop a lightweight conflict probe to characterize task interactions and guide task ordering, together with a cycling strategy that balances subspace estimation and timely task revisitation. Experiments on representative Code, Reason, and Math tasks show that PMOPD improves every evaluated capability over MOPD, raising the average score across the three tasks by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. These consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.
Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97% of FP32 master weights differ from initialization, but only 7--11% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.
Siqi Zhu, Suozhi Huang, Kaixuan Zhang +4
University of Illinois Urbana-Champaign · Princeton University · Westlake University