Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one domain suppresses capabilities acquired from another. Inspired by the distinctive update geometry of OPD, we find that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. We therefore propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions. We further develop a lightweight conflict probe to characterize task interactions and guide task ordering, together with a cycling strategy that balances subspace estimation and timely task revisitation. Experiments on representative Code, Reason, and Math tasks show that PMOPD improves every evaluated capability over MOPD, raising the average score across the three tasks by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. These consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.
Figures & tables
Figure 1: Visualization of task-preferred update directions and the optimization trajectories of MOPD and PMOPD. The axes θ1 and θ2 define a rank-2 PCA plane derived from actual checkpoint displacements. (a-c) show the normalized Math, Code, and Reason OPD losses. The arrows indicate their preferred next-update directions from a shared reference checkpoint. (d) shows the equally weighted multi-task relative OPD loss, while (e-f) overlay the actual MOPD and four-cycle PMOPD trajectories on this joint landscape. Lighter colors indicate lower relative loss and darker colors indicate higher relative loss. The task-specific arrows differ substantially at the same checkpoint, and MOPD does not consistently descend on the joint landscape. By correcting updates against protected task directions, PMOPD follows a more stable trajectory toward a shared low-loss region.
Math
Code
Reason
Math
1.000
0.151
0.142
Code
0.151
1.000
0.133
Reason
0.142
0.133
1.000
Table 1: Similarities among the top-16 OPD update subspaces induced by different training domains. The low cross-domain similarities support selective subspace protection.
Figure 2: Overview of PMOPD in cycle N . The upper row shows sequential OPD across the Code, Reason, and Math stages, and the lower row shows subspace-memory construction and projected protection. Code is trained with an empty memory. Its block displacement constructs the first protected basis, which is expanded after Reason and used to constrain the Math stage. The final student initializes cycle N+1 , where the memory is rebuilt from new task blocks.
Task
RL
OPD
Test
Math
2700
600
450
Reason
4870
600
1221
Code
2400
600
300
Table 2: Data configuration for the main experiments.
Qwen2.5-7B
Llama-3.1-8B
Model / method
Math
Reason
Code
Avg.
Math
Reason
Code
Avg.
Student Model
51.78
50.20
49.67
50.55
11.33
58.56
27.33
32.41
Math Teacher
66.89
56.84
49.67
57.80
18.22
57.08
27.33
34.21
Reason Teacher
58.89
70.35
50.00
59.75
8.89
68.55
24.33
33.92
Code Teacher
64.22
60.69
58.00
60.97
11.33
56.43
42.00
36.59
Parameter Merge
61.56
60.44
53.67
58.56
12.22
63.64
30.00
35.29
Table 3: Results across the Qwen2.5-7B and Llama-3.1-8B families. Each group reports task-specific scores and their unweighted average. Bold values mark the best within each model family and column.
Variant
Math
Reason
Code
Avg.
No projection
67.33
68.55
55.67
63.85
+ Gradient projection
67.11
73.87
57.67
66.22
+ Optimizer-update projection
67.78
73.46
59.67
66.97
Table 4: Projection ablation on Qwen2.5-7B under the same order and cycle schedule. The variants isolate the contributions of gradient projection and optimizer-update projection.
Order
Math
Reason
Code
Avg.
C → M → R
65.78
74.04
55.67
65.16
M → R → C
69.11
70.68
55.33
65.04
C → R → M
70.22
73.87
53.33
65.81
M → C → R
66.22
69.70
54.67
63.53
R → C → M
65.11
72.32
56.67
64.70
R → M → C
64.22
72.97
51.33
62.84
Table 5: One-cycle performance of all six task orders. C, R, and M denote Code, Reason, and Math, respectively.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Fold
Code
Reason
Math
Ascending order
1
27.55%
35.37%
36.98%
Code < Reason < Math
2
27.96%
35.98%
37.07%
Code < Reason < Math
3
27.29%
34.62%
35.98%
Code < Reason < Math
4
28.10%
35.53%
37.14%
Code < Reason < Math
5
29.68%
37.29%
37.86%
Code < Reason < Math
6
28.25%
36.00%
37.09%
Code < Reason < Math
Appendix
Table 6: Task-level average undirected conflict estimated from each 60-example fold. The final column gives the ascending task order within each fold.
Direction
Removed gradient
Math → Reason
36.86%
Reason → Math
34.22%
Code → Math
26.56%
Code → Reason
25.06%
Math → Code
19.87%
Reason → Code
17.85%
Appendix
Table 7: Directional conflict measured by the removed-gradient ratio.
Cycles
Code
Reason
Math
Mean
2
0.486
0.603
0.371
0.487
4
0.535
0.622
0.417
0.525
5
0.472
0.610
0.382
0.488
8
0.433
0.584
0.338
0.452
10
0.411
0.546
0.315
0.424
Appendix
Table 8: Same-task subspace similarity between cycle revisits.
Variant
Math
Reason
Code
Avg.
WB random
66.67
71.91
55.00
64.52
WB ordered
68.78
72.71
53.00
64.83
BB random
65.33
70.68
55.67
63.89
BB ordered
65.56
69.86
56.67
64.03
Appendix
Table 9: Task-mixing variants of MOPD on Qwen2.5-7B.
Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97% of FP32 master weights differ from initialization, but only 7--11% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.
Siqi Zhu, Suozhi Huang, Kaixuan Zhang +4
University of Illinois Urbana-Champaign · Princeton University · Westlake University
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sources and thereby elevate the performance frontier of distillation, an intuitive direction is to infuse privileged information to either teacher or student itself. However, this additional input induces a potential failure mode we dub privilege illusion: a pattern that conflates the transferable capability gap that students are meant to close, and the information asymmetry gap that can only be mimicked but never replicated. This issue is further amplified by the inherent non-uniformity of token-level supervision, where only a small subset of tokens carries pivotal capability-bearing signals. To this end, we propose DOPD, an advantage-aware dual distillation paradigm that dynamically routes token-level supervision between privileged teacher and privileged student policies based on their advantage gap and relative probabilities. Each token receives supervision of different strength, objective, and strategy from either teacher or student itself, which transfers credible capability while simultaneously receiving auxiliary signals, to alleviate privilege illusion. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that DOPD consistently outperforms Vanilla OPD and other counterparts. Further results on stability, robustness, continual learning, and out-of-distribution tasks validate its superiority.
Multi-teacher on-policy distillation (MOPD) combines independently developed domain teachers into a single student by distilling their predictions on student-generated samples. We study a setting where teachers share a reference model but undergo different post-training procedures, and find that MOPD can struggle to recover some teacher capabilities. Because distillation occurs on student-generated prefixes, the student initialization can strongly affect subsequent recovery. However, initial benchmark performance is not a reliable predictor of a good MOPD initialization. For example, merge initialization can start below SFT warm-up yet finish higher after MOPD. We further find that effective merging depends on both the relative teacher contributions and the overall merge scale, with some strong configurations lying outside the simplex of convex parameter averaging. Thus, selecting a good merge initialization requires evaluating not only its immediate performance but also the learning it enables under MOPD, making one-shot coefficient search difficult. We propose Iterative Merging for MOPD (IM-MOPD), which starts from a uniform merge and progressively adds task-vector increments for under-recovered domains during distillation. In a 5-domain setting, IM-MOPD achieves higher average normalized recovery than MOPD with either uniform merge initialization or SFT warm-up, showing that effective teacher contributions can be determined progressively during training.