Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillation, prompts from different domains are assigned to the corresponding specialist. Both settings usually transfer each teacher's endpoint policy, which mixes what post-training changed with preferences inherited from the teacher's base. We introduce Δ-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed. We first expose the mechanism that impedes endpoint transfer: inherited base pull can exceed the post-training shift. Removing it reduces the teacher-term norm ratio and target--student KL. Across our experiments, the results suggest that shift targets are particularly useful when teacher signals are combined at a state. With three composed teachers, Δ-MOPD exceeds endpoint composition by 4.11 Math and 1.95 five-benchmark points; with two, it matches endpoint accuracy. Under phased routing, it achieves higher mean performance in both phase orders and reduces the observed order gap from 10.50 to 6.42 points. Under interleaved routing, where each update involves one teacher, the two targets perform comparably. The phased results provide supporting evidence that the benefit may extend to signals accumulated across training phases. Target construction is thus an independent design axis in MOPD, complementary to teacher selection.
Figures & tables
Figure 1: Endpoint versus shift targets. Each endpoint term zTi−zA contains a post-training shift zTi−zBi and an inherited base difference zBi−zA . Δ -MOPD keeps only the shift and re-anchors it at A , giving the effective teacher Ti ; the two composite targets differ by the summed base differences. Teacher 1 is drawn with B1=A , so T1=T1 . Positions are schematic.
Role
Model
Base
Origin
Specialization
Use
Anchor / student init
DS-R1-Distill-Qwen-1.5B
self
self
–
all runs
Math teacher
JustRL-DeepSeek-1.5B
DS-R1-Distill-Qwen-1.5B
same
Math RL
composition ( M=3 )
Science/IF teacher
Nemotron-Reasoning-Qwen-1.5B
DS-R1-Distill-Qwen-1.5B
same
multi-domain RL
composition, routing
Math teacher
Polaris-7B-Preview
DS-R1-Distill-Qwen-7B
cross
Math RL
all runs
Math teacher
Qwen3-DAPO-449
Qwen3-8B
cross
DAPO Math RL
extension (App. F )
Table 1: Public checkpoints. “Base” is the exact precursor used to compute each shift. Nemotron is pinned to revision v1 .
Figure 2: Composition diagnostics in the mechanism run. (a) For Polaris, the base-reference pull has a larger gradient norm than the shift. (b) The endpoint composite has a 5.2:1 teacher-term norm ratio; the shift composite has 1.44:1 . (c) The shift target is about five times closer to the student in target–student KL. Values are late-training averages (Table 7 ).
Teachers
Target
Math
Sci/IF
Five-suite
M=2
Endpoint composite
28.96
15.70
23.66
M=2
Δ -MOPD
29.10
17.38
24.41
M=3
Endpoint composite
29.39
16.51
24.24
M=3
Δ -MOPD
33.50
15.22
26.19
Δ3−2
Endpoint composite
+0.43
+0.81
+0.58
Δ3−2
Δ -MOPD
+4.40
−2.16
+1.78
Table 2: Common-domain composition in the scaling run at step 100 . Greedy pass@ 1 macros (%). All teachers score every BigMath rollout. Δ3−2 is the change from adding same-origin JustRL. Per-benchmark results are in Table 9 .
Schedule
Macro
Endpoint
Δ -MOPD
Δ
Interleaved
Math
30.61
31.52
+0.91
Sci/IF
18.26
18.45
+0.19
Five-suite
25.67
26.29
+0.62
Phased: Sci/IF → Math
Math
28.02
34.39
+6.37
Sci/IF
18.25
22.89
+4.64
Five-suite
24.11
29.79
+5.68
Table 3: Routed-domain distillation. Interleaved routing mixes domains in each batch ( 50/25/25 Math/Science/IF); phased routing trains one teacher–domain phase at a time. Greedy pass@ 1 macros (%); Δ is Δ -MOPD minus Endpoint in pp.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Acquisition ( M=1 )
Mechanism ( M=2 )
Scaling ( M=2,3 ; proj. M=4 )
Phased routing
Interleaved routing
Phase structure
single
single
single
2 phases, optimizer resumed, dataloader reset
single, domains interleaved from step 1
On-policy prompts
BigMath, 25 K Math
BigMath, 25 K Math
BigMath Math
BigMath; TextbookReasoning + Llama-Nemotron IF
Math + Science/IF mixture ( 50/25/25 )
Temperature / top- p
1.0 / 1.0
1.0 / 1.0
1.0 / 1.0
1.0 / 1.0
1.0 / 1.0, top- k off
Maximum response
10,000
10,000
10,000
10,000
10,000
Learning rate
10−6
10−6
10−6
10−6
10−6
Hardware
4 × H100 80GB
4 × H100 80GB
8 × H200
4 × H100 80GB per arm
8 × H200 (Endpoint) / 8 × H100 ( Δ -MOPD)
Appendix
Table 4: Training protocols. Values are shared by both arms of a run unless an arm-specific value is shown. Fields marked “–” were not recorded in the run’s frozen configuration summary.
Group
Benchmark
Test items
Avg@ K
Generation / scoring
English Math
AMC 2023
40
16
sampled, official answer checker
English Math
MATH-500
500
1
greedy, official answer checker
English Math
AIME 2025
30
16
sampled, official answer checker
Science
GPQA-Diamond
198
4
sampled, final option-letter scorer
Instruction following
IF-Eval
541
4
sampled, vendored rule checker
Appendix
Table 5: Held-out evaluation suite. Avg@ K applies to the acquisition and mechanism runs; other runs decode every benchmark greedily at K=1 .
Benchmark
Initial student
Δ -MOPD
Endpoint composite
Δ -MOPD − Endpoint (pp)
AMC 2023
63.12
79.84±1.99
79.53±2.79
+0.31
MATH-500
36.00
59.40±1.29
62.80±1.34
−3.40
AIME 2025
23.33
32.92±3.04
32.92±3.24
0.00
English Math macro
40.82
57.39
58.42
−1.03
GPQA-Diamond
34.72
40.28±2.11
39.02±2.22
+1.26
IF-Eval
23.43
24.58±1.31
24.68±1.07
−0.10
Appendix
Table 6: Mechanism run ( M=2 ) student evaluation. AMC 2023 and AIME 2025 use Avg@ 16 , GPQA-Diamond and IF-Eval use Avg@ 4 , and MATH-500 uses greedy decoding. Entries with error bars are mean ± sample standard deviation over training seeds, in percent.
Target
Cancellation
Aggregate norm
Conflict rate
Signal cosine
Entropy
KL to anchor
Target–student KL
Δ -MOPD
0.2871
3247.5
0.5540
−0.0163
0.2458
0.1243
0.0307
Endpoint composite
0.1363
10231.5
0.4386
0.0592
0.2643
0.2567
0.1522
Equal-norm reference
0.2929
—
0.5000
0.0000
—
—
—
Appendix
Table 7: Late-training diagnostics for the mechanism run, averaged over steps 300,320,340,360,380 .
Target
Step
Cancellation
Target–student KL
Δ -MOPD
20
0.2728
0.1070
Δ -MOPD
60
0.2914
0.0750
Δ -MOPD
100
0.2832
0.0462
Endpoint composite
20
0.1265
0.2472
Endpoint composite
60
0.1378
0.1222
Endpoint composite
100
0.1381
0.1632
Appendix
Table 8: Early mechanism-run diagnostics at steps 20,60,100 .
Teachers
Target
AMC 2023
MATH-500
AIME 2025
Math macro
GPQA-D
IF-Eval
Sci/IF macro
Five-suite macro
M=2
Endpoint composite
35.50±2.56
34.04±1.37
17.33±1.93
28.96
16.77±2.27
14.64±1.66
15.70
23.66
M=2
Δ -MOPD
38.00±1.68
28.64±1.42
20.67±3.11
29.10
20.30±1.90
14.45±1.33
17.38
24.41
M=3
Endpoint composite
38.00±1.80
32.84±0.99
17.33±3.08
29.39
17.27±1.53
15.75±1.14
16.51
24.24
M=3
Δ -MOPD
38.00±1.61
31.84±0.99
30.67±2.50
33.50
17.27±1.48
13.16±1.50
15.22
26.19
Appendix
Table 9: Scaling run at step 100 . Greedy pass@ 1 ; entries with error bars are mean ± sample standard deviation over training seeds, in percent.
Order
Method
AMC 2023
MATH-500
AIME 2025
GPQA-Diamond
IF-Eval
Math macro
Sci/IF macro
Science/IF → Math
Endpoint
25.50±2.43
41.24±0.84
17.33±2.19
15.76±1.93
20.74±0.81
28.02
18.25
Δ -MOPD
33.00±2.06
42.84±1.31
27.33±2.69
24.85±2.28
20.92±1.92
34.39
22.89
Math → Science/IF
Endpoint
40.50±1.79
51.44±0.94
30.67±2.21
29.90±1.81
20.55±1.35
40.87
25.23
Δ -MOPD
40.50±2.48
55.44±1.12
30.67±2.52
32.42±2.43
22.03±1.47
42.20
27.23
Appendix
Table 10: Phased routing at the step- 300 checkpoint. Greedy pass@ 1 ; group macros are unweighted averages within each group. Entries with error bars are mean ± sample standard deviation over training seeds, in percent.
Target
AMC 2023
MATH-500
AIME 2025
English Math macro
Δ -MOPD
67.81±1.53
48.80±1.40
27.08±1.82
47.90
Endpoint
64.84±1.71
44.20±0.93
26.04±1.73
45.03
Appendix
Table 11: Acquisition run at step 100 . AMC 2023 and AIME 2025 use Avg@ 16 ; MATH-500 uses greedy decoding. The macro is the unweighted average of these three benchmarks. Entries with error bars are mean ± sample standard deviation over training seeds, in percent.
English-Math macro
Δ -MOPD
Endpoint
46.18%
≤28.0 h
≤33.0 h
46.54%
≤28.0 h
≤42.9 h
47.90%
≤28.0 h
not observed by 42.9 h
Appendix
Table 12: Earliest evaluated checkpoint meeting each held-out English-Math threshold in the acquisition run, in allocated H100-hours. Checkpoints were evaluated every 50 steps, so every entry is an upper bound on first-passage compute.
Teachers
Target
AMC 2023
MATH-500
AIME 2025
Math macro
GPQA-D
IF-Eval
Sci/IF macro
Five-suite macro
M=3
Δ -MOPD
38.00±1.61
31.84±0.99
30.67±2.50
33.50
17.27±1.48
13.16±1.50
15.22
26.19
M=4
Projected Δ -MOPD
40.50±1.92
32.64±1.19
37.33±2.63
36.82
16.77±1.76
12.94±1.42
14.85
28.04
Appendix
Table 13: Projected four-teacher extension at step 100 , with the M=3 shift composite for reference. Greedy pass@ 1 , in percent.
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabeled training mixtures and leaves complementary signals from other teachers unused. We introduce MOPD-Router, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model. Its plug-in interface supports different metrics for selecting and weighting teacher-specific OPD signals. Within this interface, we propose ExpertAlign, which scores each teacher by whether its correction to the student at the current token expresses the specialization that teacher acquired during post-training, and compare it against two reference metrics built on teacher confidence (Entropy) and teacher-student discrepancy (Novelty). Experiments on unlabeled and domain-labeled training mixtures under strong-to-weak and same-size distillation scenarios show that ExpertAlign achieves the strongest overall performance in all four settings. On unlabeled data, it improves the overall score by 5.88 (+12.3%) points over Mean aggregation; on domain-labeled data, it outperforms standard MOPD by 3.95 (+7.8%) points without using available domain labels. These results demonstrate token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment. Code is available at: https://github.com/TURLEing/MOPD-Router.
Tianze Xu, Yanzhao Zheng, Zhentao Zhang +10
Shanghai Jiao Tong University · GAIR · Alibaba Group +2
Multi-teacher on-policy distillation (MOPD) combines independently developed domain teachers into a single student by distilling their predictions on student-generated samples. We study a setting where teachers share a reference model but undergo different post-training procedures, and find that MOPD can struggle to recover some teacher capabilities. Because distillation occurs on student-generated prefixes, the student initialization can strongly affect subsequent recovery. However, initial benchmark performance is not a reliable predictor of a good MOPD initialization. For example, merge initialization can start below SFT warm-up yet finish higher after MOPD. We further find that effective merging depends on both the relative teacher contributions and the overall merge scale, with some strong configurations lying outside the simplex of convex parameter averaging. Thus, selecting a good merge initialization requires evaluating not only its immediate performance but also the learning it enables under MOPD, making one-shot coefficient search difficult. We propose Iterative Merging for MOPD (IM-MOPD), which starts from a uniform merge and progressively adds task-vector increments for under-recovered domains during distillation. In a 5-domain setting, IM-MOPD achieves higher average normalized recovery than MOPD with either uniform merge initialization or SFT warm-up, showing that effective teacher contributions can be determined progressively during training.
Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, however, select one teacher per example and keep it fixed throughout the response. This design both depends on domain labels that mixed training corpora often lack and cannot adapt teacher selection when the expertise required changes within a trajectory. We observe that each specialist deviates more from a shared reference on in-domain prompts than on out-of-domain prompts, on average. Based on this observation, we propose \textbf{TrustMOPD}, which replaces example-level teacher selection with label-free, token-level supervision allocation. At each student-generated prefix, TrustMOPD measures this displacement in next-token preferences, calibrates its magnitude across teachers, and uses the resulting scores as proxies for local reliability to weight teacher-specific distillation losses. Evaluated across mathematics, code, and instruction following, TrustMOPD closes 91.5% and 98.0% of the overall-score gap between the initial student and oracle-routed teachers when trained on \textsc{SingleCap} and \textsc{MultiCap}, respectively, compared with 54.4% and 54.5% for the strongest label-free baseline in each setting. On \textsc{SingleCap}, it approaches label-based MOPD without using domain labels.
Jie Sun, Mao Zheng, Mingyang Song +8
University of Science and Technology of China · Foundation Model Department, Tencent · Tsinghua University +4