On-policy distillation from multiple teachers combines expertise from different domains in a single student, but conflicting gradients can hinder this integration. Gradient corrections directly constrain parameter updates under plain SGD. With optimizers such as AdamW, however, momentum, adaptive scaling, and weight decay can turn a corrected gradient into an update that increases a domain loss to first order. To address this gap, we propose Update Projection for Multi-Teacher On-Policy Distillation (UP-MOPD). UP-MOPD lets the original mixed gradient update the optimizer state and generate a candidate displacement, then projects only violating candidates before they are committed to the parameters. The projection gives the unique feasible update closest to the candidate in Euclidean distance. In experiments combining medical and general domains, UP-MOPD improves IFEval-loose accuracy late in training by 2.96 points over vanilla M-OPD. It achieves an average score of 60.03 across eight metrics, compared with 59.00 for gradient projection and 59.15 for update rejection. On a public benchmark covering mathematics, code, and instruction following, it achieves the best average across six tasks (32.67), leads on LiveCodeBench v5, and ties for the best IFEval result.These results support projecting optimizer updates to reduce interference between domains.
Figures & tables
Figure 1: Overview of UP-MOPD. AdamW forms a candidate parameter displacement from the mixed distillation gradient. UP-MOPD evaluates its first-order effect on each active domain and minimally corrects a violating candidate before commit, while retaining the current-step optimizer state produced by the original mixture.
Figure 2: Medical performance and instruction following. Student points average the last three saved checkpoints of one training epoch. The arrow from M-OPD to UP-MOPD shows improved IFEval-loose accuracy at a similar HealthBench-Hard score.
Medical
General
Math
Method
HealthBench Hard
Med- MCQA
PubMed- QA
MMLU- Med
GPQA
IFEval
AIME24
AIME25
Mean
Qwen3-4B teacher
6.38
57.26
76.60
78.32
42.42
86.88
54.00
43.33
55.65
Medical Teacher
39.06
59.07
70.00
80.09
57.68
73.38
47.33
42.00
58.58
M-OPD
38.73
59.77
71.67
79.54
47.81
77.94
53.78
43.78
59.13
GP-MOPD
38.79
59.80
73.00
79.85
48.28
78.93
52.89
40.44
59.00
UP-MOPD
38.57
59.37
74.33
79.54
48.38
80.90
53.33
45.78
60.03
Table 1: Medical and general capability integration. Student results are averaged over the final three checkpoints. Mean averages the eight metrics. Bold indicates the best student result in each column.
Math
Code
Instruction Following
Overall
Method
AIME24
AIME25
Avg.
LCB5
LCB6
Avg.
IFEval
IFBench
Avg.
Total
MixSFT
13.33
17.78
15.56
15.57
17.71
16.64
69.44
16.00
42.72
24.97
RFT ∗
22.97
23.91
23.44
18.98
19.43
19.21
55.08
18.67
36.87
26.51
MixRL ∗
21.15
22.14
21.64
16.59
20.97
18.78
70.24
22.67
46.45
28.96
ParamMerge-Avg ∗
18.91
20.99
19.95
18.38
21.20
19.79
70.06
17.67
43.86
27.87
ParamMerge-TA ∗
21.93
22.76
22.34
21.74
23.14
22.44
71.53
21.67
46.60
30.46
Table 2: Results on the public three-domain benchmark under the Open-MOPD evaluation protocol. Domain averages cover two tasks, and Total averages all six tasks. ∗ marks results reported by Gao et al. (2026) ; unmarked rows are our reproductions or proposed methods. RouteOPD uses three independently distilled students. Bold indicates the best result in each column.
Medical
General
Math
Method
HealthBench Hard
Med- MCQA
PubMed- QA
MMLU- Med
GPQA
IFEval
AIME24
AIME25
Mean
M-OPD
38.73
59.77
71.67
79.54
47.81
77.94
53.78
43.78
59.13
M-OPD + Grad. Clip
38.31
59.42
73.47
80.04
48.35
73.20
53.78
45.78
59.04
Reweighting w/ rejection
39.11
60.00
73.27
79.94
45.25
76.77
52.89
44.89
59.02
Reweighting w/o rejection
38.16
59.49
72.87
79.77
47.44
78.99
54.67
45.11
59.56
Update Rejection
38.68
58.07
73.87
79.11
45.99
80.77
52.00
44.67
59.15
Table 3: Effect of intervention strategies in the medical–general setting, using the same evaluation protocol as Table 1 . Mean averages the eight metrics. Bold indicates the best result in each column.
Figure 3: Gradient alignment and update conflict. Each point is a training step, with g1⊤g2 on the horizontal axis and h1h2 on the vertical axis. A negative h1h2 means the candidate update is predicted to lower one domain’s loss and raise the other’s. Gradient alignment alone does not determine this outcome.
Figure 4: Capability and distillation dynamics. (a,b) HealthBench-Hard and IFEval-loose checkpoint scores; lines denote teacher or initial-student references and shading the late-training window. (c,d) Trailing 50-step means of domain-conditioned sampled reverse KL on each method’s responses; steps lacking a domain and incomplete initial windows are omitted.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Method
HB-Hard
IFEval
Mean
Tokens
Trunc. (%)
M-OPD
38.73
77.94
59.13
4,067
9.20
Update Rejection
38.68
80.77
59.15
4,071
9.48
GP-MOPD
38.79
78.93
59.00
4,056
9.63
UP-MOPD
38.57
80.90
60.03
4,089
9.20
Appendix
Table 4: Capability scores and generation behavior. HB-Hard denotes HealthBench-Hard; IFEval uses prompt-level loose accuracy. Mean averages eight benchmarks. Tokens and Trunc. are mean response length and truncation rate over the last 200 plotted training steps.
Figure 5: Offline projection of 645 rejected candidates. Cumulative distributions show relative correction size (left) and retained predicted medical loss decrease (right). Dashed lines mark medians.
Domain
Teacher suffix
Step
Prompt source
Prompts
Mathematics
RL-Math
100
DAPO-Math-17k
17,917
Code
RL-Code
180
DeepCoder-Preview-Dataset
23,667
Instruction following
RL-IF
2,440
Nemotron-IF-RL-46k
45,347
Total
86,931
Appendix
Table 5: Released teachers and training data in the public setting. Model suffixes follow the shared repository prefix given in the text. Teacher steps identify the checkpoints described in their model cards. The released prompt pool excludes code records identified as LiveCodeBench.
Hyperparameter
Value
Student
Qwen3-4B-Instruct-2507
Teachers
Qwen3-4B-Instruct-2507 and Medical Teacher; both frozen, with supervision routed by prompt domain
Training framework
Megatron
Training data
22,564 prompts: 17,398 mathematical prompts and 5,166 prompts assigned to the medical teacher; combined and randomly shuffled
Training duration
1 epoch; 1,410 steps
Batch size
16 prompts per step; 4 responses per prompt (64 samples)
Appendix
Table 6: Training hyperparameters for medical and general capability integration.
Hyperparameter
Value
Student
SmolLM3-3B initialized from MixSFT
Teachers
Frozen RL-Math, RL-Code, and RL-IF releases (Table 5 )
Training framework
verl; public data and checkpoints from Open-MOPD ( Gao et al., 2026 )
Training data
86,931 prompts mixed across the three domains
Training duration
1 epoch; 679 full batches
Batch size
128 prompts per step; mini-batch size 128; micro-batch size 1 per GPU; 8 GPUs
Appendix
Table 7: Training hyperparameters for the public mathematics, code, and instruction-following benchmark.
Benchmark
Token limit
Repetitions
Demonstrations
HealthBench-Hard
16,384
1
0
GPQA-Diamond, AIME24/25
8,192
5
0
IFEval, MMLU-Med
8,192
1
0
MedMCQA
32,768
1
5
PubMedQA
32,768
1
0
Appendix
Table 8: Generation settings for the medical evaluation. Limits count generated tokens. Repetitions are completions per question.
Method
Logged step time (s)
Actor training (s)
M-OPD
66.55±6.58
11.57±1.65
Update Rejection
92.58±9.45
37.49±5.71
GP-MOPD
92.94±9.69
38.00±5.77
UP-MOPD
94.76±10.08
39.69±5.82
Appendix
Table 9: Runtime in the medical–general setting on eight NVIDIA H200 GPUs. Values are means ± standard deviations in seconds per step, using the same 200-step late-training window for all four methods. Actor training is included in the logged step time.
Step
Math
Code
IF
Total
100
18.89
21.24
50.58
30.24
200
26.67
20.44
49.48
32.20
300
23.33
19.87
48.53
30.57
400
19.45
21.33
48.49
29.75
500
23.89
21.93
47.54
31.12
600
22.22
21.94
49.12
31.09
Appendix
Table 10: UP-MOPD training trajectory in the public three-domain setting.
Figure 6: UP-MOPD training trajectory in the public three-domain setting. Each panel shows a domain mean over two benchmarks. The vertical scales differ across panels.
Method
Retained
Lost
Newly passed
Accuracy (%)
M-OPD
400.7
69.3
21.0
77.94
UP-MOPD
419.7
50.3
18.0
80.90
Appendix
Table 11: IFEval retention across three late-training checkpoints. The initial student passes 470 of 541 prompts under loose scoring. Retained and lost prompts partition these 470 initial passes; newly passed prompts come from the 71 initial failures. Counts and accuracies are means over the three checkpoints.
Metric
M-OPD
UP-MOPD
Difference
95% CI
BH-adjusted p
HealthBench-Hard
38.51
38.63
+0.12
[−1.34,1.60]
1.0000
GPQA
49.80
48.18
−1.62
[−4.65,1.31]
0.5279
IFEval
77.82
80.96
+3.14
[0.37,5.91]
0.1454
MedMCQA
59.67
58.71
−0.96
[−1.77,−0.12]
0.1454
PubMedQA
72.00
74.40
+2.40
[−0.20,5.00]
0.2564
MMLU-Med
79.41
79.44
+0.03
[−1.34,1.44]
1.0000
Appendix
Table 12: Paired comparison at the middle late-training checkpoint. IFEval uses prompt-level loose accuracy. Differences are UP-MOPD minus M-OPD, in score points. Intervals are pointwise 95% paired bootstrap confidence intervals; p -values are corrected across all eight metrics.
Figure 7: Instruction-following behavior in the medical setting. (a) UP-MOPD minus M-OPD instruction-level loose accuracy for all ten evaluator-defined constraint categories; parentheses give instruction counts. (b) Prompt-level loose accuracy grouped by the number of constraints in a prompt; labels above the groups give the mean gain in percentage points. Faint points show three matched late-training checkpoints from one run per method; solid points show their arithmetic mean.
Constraint
M-OPD
UP-MOPD
Exactly three Markdown bullet points
Repeats the three-item list after “Final concise version with desert emphasis:”. Loose: fail.
Produces one three-item list. Loose: pass.
At most 100 words and at least two highlighted sections
115 words and two highlighted spans. Highlighting passes, but the length constraint fails. Loose: fail.
99 words and three highlighted spans. Both constraints pass. Loose: pass.
Appendix
Table 13: Examples of constraint compliance. The initial student passes both prompts; distilled outputs are from the middle late-training checkpoint. Counts are computed from complete responses; word counts follow IFEval’s regular-expression tokenizer.