Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one domain suppresses capabilities acquired from another. Inspired by the distinctive update geometry of OPD, we find that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. We therefore propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions. We further develop a lightweight conflict probe to characterize task interactions and guide task ordering, together with a cycling strategy that balances subspace estimation and timely task revisitation. Experiments on representative Code, Reason, and Math tasks show that PMOPD improves every evaluated capability over MOPD, raising the average score across the three tasks by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. These consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.
Figures & tables
Figure 1: Visualization of task-preferred update directions and the optimization trajectories of MOPD and PMOPD. The axes θ1 and θ2 define a rank-2 PCA plane derived from actual checkpoint displacements. (a-c) show the normalized Math, Code, and Reason OPD losses. The arrows indicate their preferred next-update directions from a shared reference checkpoint. (d) shows the equally weighted multi-task relative OPD loss, while (e-f) overlay the actual MOPD and four-cycle PMOPD trajectories on this joint landscape. Lighter colors indicate lower relative loss and darker colors indicate higher relative loss. The task-specific arrows differ substantially at the same checkpoint, and MOPD does not consistently descend on the joint landscape. By correcting updates against protected task directions, PMOPD follows a more stable trajectory toward a shared low-loss region.
Math
Code
Reason
Math
1.000
0.151
0.142
Code
0.151
1.000
0.133
Reason
0.142
0.133
1.000
Table 1: Similarities among the top-16 OPD update subspaces induced by different training domains. The low cross-domain similarities support selective subspace protection.
Figure 2: Overview of PMOPD in cycle N . The upper row shows sequential OPD across the Code, Reason, and Math stages, and the lower row shows subspace-memory construction and projected protection. Code is trained with an empty memory. Its block displacement constructs the first protected basis, which is expanded after Reason and used to constrain the Math stage. The final student initializes cycle N+1 , where the memory is rebuilt from new task blocks.
Task
RL
OPD
Test
Math
2700
600
450
Reason
4870
600
1221
Code
2400
600
300
Table 2: Data configuration for the main experiments.
Qwen2.5-7B
Llama-3.1-8B
Model / method
Math
Reason
Code
Avg.
Math
Reason
Code
Avg.
Student Model
51.78
50.20
49.67
50.55
11.33
58.56
27.33
32.41
Math Teacher
66.89
56.84
49.67
57.80
18.22
57.08
27.33
34.21
Reason Teacher
58.89
70.35
50.00
59.75
8.89
68.55
24.33
33.92
Code Teacher
64.22
60.69
58.00
60.97
11.33
56.43
42.00
36.59
Parameter Merge
61.56
60.44
53.67
58.56
12.22
63.64
30.00
35.29
Table 3: Results across the Qwen2.5-7B and Llama-3.1-8B families. Each group reports task-specific scores and their unweighted average. Bold values mark the best within each model family and column.
Variant
Math
Reason
Code
Avg.
No projection
67.33
68.55
55.67
63.85
+ Gradient projection
67.11
73.87
57.67
66.22
+ Optimizer-update projection
67.78
73.46
59.67
66.97
Table 4: Projection ablation on Qwen2.5-7B under the same order and cycle schedule. The variants isolate the contributions of gradient projection and optimizer-update projection.
Order
Math
Reason
Code
Avg.
C → M → R
65.78
74.04
55.67
65.16
M → R → C
69.11
70.68
55.33
65.04
C → R → M
70.22
73.87
53.33
65.81
M → C → R
66.22
69.70
54.67
63.53
R → C → M
65.11
72.32
56.67
64.70
R → M → C
64.22
72.97
51.33
62.84
Table 5: One-cycle performance of all six task orders. C, R, and M denote Code, Reason, and Math, respectively.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Fold
Code
Reason
Math
Ascending order
1
27.55%
35.37%
36.98%
Code < Reason < Math
2
27.96%
35.98%
37.07%
Code < Reason < Math
3
27.29%
34.62%
35.98%
Code < Reason < Math
4
28.10%
35.53%
37.14%
Code < Reason < Math
5
29.68%
37.29%
37.86%
Code < Reason < Math
6
28.25%
36.00%
37.09%
Code < Reason < Math
Appendix
Table 6: Task-level average undirected conflict estimated from each 60-example fold. The final column gives the ascending task order within each fold.
Direction
Removed gradient
Math → Reason
36.86%
Reason → Math
34.22%
Code → Math
26.56%
Code → Reason
25.06%
Math → Code
19.87%
Reason → Code
17.85%
Appendix
Table 7: Directional conflict measured by the removed-gradient ratio.
Cycles
Code
Reason
Math
Mean
2
0.486
0.603
0.371
0.487
4
0.535
0.622
0.417
0.525
5
0.472
0.610
0.382
0.488
8
0.433
0.584
0.338
0.452
10
0.411
0.546
0.315
0.424
Appendix
Table 8: Same-task subspace similarity between cycle revisits.
Variant
Math
Reason
Code
Avg.
WB random
66.67
71.91
55.00
64.52
WB ordered
68.78
72.71
53.00
64.83
BB random
65.33
70.68
55.67
63.89
BB ordered
65.56
69.86
56.67
64.03
Appendix
Table 9: Task-mixing variants of MOPD on Qwen2.5-7B.