Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillation, prompts from different domains are assigned to the corresponding specialist. Both settings usually transfer each teacher's endpoint policy, which mixes what post-training changed with preferences inherited from the teacher's base. We introduce Δ-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed. We first expose the mechanism that impedes endpoint transfer: inherited base pull can exceed the post-training shift. Removing it reduces the teacher-term norm ratio and target--student KL. Across our experiments, the results suggest that shift targets are particularly useful when teacher signals are combined at a state. With three composed teachers, Δ-MOPD exceeds endpoint composition by 4.11 Math and 1.95 five-benchmark points; with two, it matches endpoint accuracy. Under phased routing, it achieves higher mean performance in both phase orders and reduces the observed order gap from 10.50 to 6.42 points. Under interleaved routing, where each update involves one teacher, the two targets perform comparably. The phased results provide supporting evidence that the benefit may extend to signals accumulated across training phases. Target construction is thus an independent design axis in MOPD, complementary to teacher selection.
Figures & tables
Figure 1: Endpoint versus shift targets. Each endpoint term zTi−zA contains a post-training shift zTi−zBi and an inherited base difference zBi−zA . Δ -MOPD keeps only the shift and re-anchors it at A , giving the effective teacher Ti ; the two composite targets differ by the summed base differences. Teacher 1 is drawn with B1=A , so T1=T1 . Positions are schematic.
Role
Model
Base
Origin
Specialization
Use
Anchor / student init
DS-R1-Distill-Qwen-1.5B
self
self
–
all runs
Math teacher
JustRL-DeepSeek-1.5B
DS-R1-Distill-Qwen-1.5B
same
Math RL
composition ( M=3 )
Science/IF teacher
Nemotron-Reasoning-Qwen-1.5B
DS-R1-Distill-Qwen-1.5B
same
multi-domain RL
composition, routing
Math teacher
Polaris-7B-Preview
DS-R1-Distill-Qwen-7B
cross
Math RL
all runs
Math teacher
Qwen3-DAPO-449
Qwen3-8B
cross
DAPO Math RL
extension (App. F )
Table 1: Public checkpoints. “Base” is the exact precursor used to compute each shift. Nemotron is pinned to revision v1 .
Figure 2: Composition diagnostics in the mechanism run. (a) For Polaris, the base-reference pull has a larger gradient norm than the shift. (b) The endpoint composite has a 5.2:1 teacher-term norm ratio; the shift composite has 1.44:1 . (c) The shift target is about five times closer to the student in target–student KL. Values are late-training averages (Table 7 ).
Teachers
Target
Math
Sci/IF
Five-suite
M=2
Endpoint composite
28.96
15.70
23.66
M=2
Δ -MOPD
29.10
17.38
24.41
M=3
Endpoint composite
29.39
16.51
24.24
M=3
Δ -MOPD
33.50
15.22
26.19
Δ3−2
Endpoint composite
+0.43
+0.81
+0.58
Δ3−2
Δ -MOPD
+4.40
−2.16
+1.78
Table 2: Common-domain composition in the scaling run at step 100 . Greedy pass@ 1 macros (%). All teachers score every BigMath rollout. Δ3−2 is the change from adding same-origin JustRL. Per-benchmark results are in Table 9 .
Schedule
Macro
Endpoint
Δ -MOPD
Δ
Interleaved
Math
30.61
31.52
+0.91
Sci/IF
18.26
18.45
+0.19
Five-suite
25.67
26.29
+0.62
Phased: Sci/IF → Math
Math
28.02
34.39
+6.37
Sci/IF
18.25
22.89
+4.64
Five-suite
24.11
29.79
+5.68
Table 3: Routed-domain distillation. Interleaved routing mixes domains in each batch ( 50/25/25 Math/Science/IF); phased routing trains one teacher–domain phase at a time. Greedy pass@ 1 macros (%); Δ is Δ -MOPD minus Endpoint in pp.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Acquisition ( M=1 )
Mechanism ( M=2 )
Scaling ( M=2,3 ; proj. M=4 )
Phased routing
Interleaved routing
Phase structure
single
single
single
2 phases, optimizer resumed, dataloader reset
single, domains interleaved from step 1
On-policy prompts
BigMath, 25 K Math
BigMath, 25 K Math
BigMath Math
BigMath; TextbookReasoning + Llama-Nemotron IF
Math + Science/IF mixture ( 50/25/25 )
Temperature / top- p
1.0 / 1.0
1.0 / 1.0
1.0 / 1.0
1.0 / 1.0
1.0 / 1.0, top- k off
Maximum response
10,000
10,000
10,000
10,000
10,000
Learning rate
10−6
10−6
10−6
10−6
10−6
Hardware
4 × H100 80GB
4 × H100 80GB
8 × H200
4 × H100 80GB per arm
8 × H200 (Endpoint) / 8 × H100 ( Δ -MOPD)
Appendix
Table 4: Training protocols. Values are shared by both arms of a run unless an arm-specific value is shown. Fields marked “–” were not recorded in the run’s frozen configuration summary.
Group
Benchmark
Test items
Avg@ K
Generation / scoring
English Math
AMC 2023
40
16
sampled, official answer checker
English Math
MATH-500
500
1
greedy, official answer checker
English Math
AIME 2025
30
16
sampled, official answer checker
Science
GPQA-Diamond
198
4
sampled, final option-letter scorer
Instruction following
IF-Eval
541
4
sampled, vendored rule checker
Appendix
Table 5: Held-out evaluation suite. Avg@ K applies to the acquisition and mechanism runs; other runs decode every benchmark greedily at K=1 .
Benchmark
Initial student
Δ -MOPD
Endpoint composite
Δ -MOPD − Endpoint (pp)
AMC 2023
63.12
79.84±1.99
79.53±2.79
+0.31
MATH-500
36.00
59.40±1.29
62.80±1.34
−3.40
AIME 2025
23.33
32.92±3.04
32.92±3.24
0.00
English Math macro
40.82
57.39
58.42
−1.03
GPQA-Diamond
34.72
40.28±2.11
39.02±2.22
+1.26
IF-Eval
23.43
24.58±1.31
24.68±1.07
−0.10
Appendix
Table 6: Mechanism run ( M=2 ) student evaluation. AMC 2023 and AIME 2025 use Avg@ 16 , GPQA-Diamond and IF-Eval use Avg@ 4 , and MATH-500 uses greedy decoding. Entries with error bars are mean ± sample standard deviation over training seeds, in percent.
Target
Cancellation
Aggregate norm
Conflict rate
Signal cosine
Entropy
KL to anchor
Target–student KL
Δ -MOPD
0.2871
3247.5
0.5540
−0.0163
0.2458
0.1243
0.0307
Endpoint composite
0.1363
10231.5
0.4386
0.0592
0.2643
0.2567
0.1522
Equal-norm reference
0.2929
—
0.5000
0.0000
—
—
—
Appendix
Table 7: Late-training diagnostics for the mechanism run, averaged over steps 300,320,340,360,380 .
Target
Step
Cancellation
Target–student KL
Δ -MOPD
20
0.2728
0.1070
Δ -MOPD
60
0.2914
0.0750
Δ -MOPD
100
0.2832
0.0462
Endpoint composite
20
0.1265
0.2472
Endpoint composite
60
0.1378
0.1222
Endpoint composite
100
0.1381
0.1632
Appendix
Table 8: Early mechanism-run diagnostics at steps 20,60,100 .
Teachers
Target
AMC 2023
MATH-500
AIME 2025
Math macro
GPQA-D
IF-Eval
Sci/IF macro
Five-suite macro
M=2
Endpoint composite
35.50±2.56
34.04±1.37
17.33±1.93
28.96
16.77±2.27
14.64±1.66
15.70
23.66
M=2
Δ -MOPD
38.00±1.68
28.64±1.42
20.67±3.11
29.10
20.30±1.90
14.45±1.33
17.38
24.41
M=3
Endpoint composite
38.00±1.80
32.84±0.99
17.33±3.08
29.39
17.27±1.53
15.75±1.14
16.51
24.24
M=3
Δ -MOPD
38.00±1.61
31.84±0.99
30.67±2.50
33.50
17.27±1.48
13.16±1.50
15.22
26.19
Appendix
Table 9: Scaling run at step 100 . Greedy pass@ 1 ; entries with error bars are mean ± sample standard deviation over training seeds, in percent.
Order
Method
AMC 2023
MATH-500
AIME 2025
GPQA-Diamond
IF-Eval
Math macro
Sci/IF macro
Science/IF → Math
Endpoint
25.50±2.43
41.24±0.84
17.33±2.19
15.76±1.93
20.74±0.81
28.02
18.25
Δ -MOPD
33.00±2.06
42.84±1.31
27.33±2.69
24.85±2.28
20.92±1.92
34.39
22.89
Math → Science/IF
Endpoint
40.50±1.79
51.44±0.94
30.67±2.21
29.90±1.81
20.55±1.35
40.87
25.23
Δ -MOPD
40.50±2.48
55.44±1.12
30.67±2.52
32.42±2.43
22.03±1.47
42.20
27.23
Appendix
Table 10: Phased routing at the step- 300 checkpoint. Greedy pass@ 1 ; group macros are unweighted averages within each group. Entries with error bars are mean ± sample standard deviation over training seeds, in percent.
Target
AMC 2023
MATH-500
AIME 2025
English Math macro
Δ -MOPD
67.81±1.53
48.80±1.40
27.08±1.82
47.90
Endpoint
64.84±1.71
44.20±0.93
26.04±1.73
45.03
Appendix
Table 11: Acquisition run at step 100 . AMC 2023 and AIME 2025 use Avg@ 16 ; MATH-500 uses greedy decoding. The macro is the unweighted average of these three benchmarks. Entries with error bars are mean ± sample standard deviation over training seeds, in percent.
English-Math macro
Δ -MOPD
Endpoint
46.18%
≤28.0 h
≤33.0 h
46.54%
≤28.0 h
≤42.9 h
47.90%
≤28.0 h
not observed by 42.9 h
Appendix
Table 12: Earliest evaluated checkpoint meeting each held-out English-Math threshold in the acquisition run, in allocated H100-hours. Checkpoints were evaluated every 50 steps, so every entry is an upper bound on first-passage compute.
Teachers
Target
AMC 2023
MATH-500
AIME 2025
Math macro
GPQA-D
IF-Eval
Sci/IF macro
Five-suite macro
M=3
Δ -MOPD
38.00±1.61
31.84±0.99
30.67±2.50
33.50
17.27±1.48
13.16±1.50
15.22
26.19
M=4
Projected Δ -MOPD
40.50±1.92
32.64±1.19
37.33±2.63
36.82
16.77±1.76
12.94±1.42
14.85
28.04
Appendix
Table 13: Projected four-teacher extension at step 100 , with the M=3 shift composite for reference. Greedy pass@ 1 , in percent.