MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
Organizations: Shanghai Jiao Tong University · GAIR · Alibaba Group · University of Science and Technology of China · Shanghai Innovation Institute
Abstract
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabeled training mixtures and leaves complementary signals from other teachers unused. We introduce MOPD-Router, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model. Its plug-in interface supports different metrics for selecting and weighting teacher-specific OPD signals. Within this interface, we propose ExpertAlign, which scores each teacher by whether its correction to the student at the current token expresses the specialization that teacher acquired during post-training, and compare it against two reference metrics built on teacher confidence (Entropy) and teacher-student discrepancy (Novelty). Experiments on unlabeled and domain-labeled training mixtures under strong-to-weak and same-size distillation scenarios show that ExpertAlign achieves the strongest overall performance in all four settings. On unlabeled data, it improves the overall score by 5.88 (+12.3%) points over Mean aggregation; on domain-labeled data, it outperforms standard MOPD by 3.95 (+7.8%) points without using available domain labels. These results demonstrate token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment. Code is available at: https://github.com/TURLEing/MOPD-Router.
Figures & tables
| Math | Code | IF | Overall | |||||||
| Method | AIME24 | AIME25 | HMMT Feb. | HMMT Nov. | Human Eval+ | MBPP+ | LCB-v6 | IFEval | IFBench | Avg. |
| Baseline Models | ||||||||||
| Qwen3-1.7B | 12.50 | 11.67 | 5.83 | 4.58 | 60.40 | 52.40 | 12.00 | 67.65 | 16.67 | 27.08 |
| Qwen3-4B | 23.33 | 10.00 | 10.00 | 7.08 | 82.30 | 67.20 | 18.29 | 81.88 | 24.14 | 36.02 |
| RL Teacher | 58.33 | 52.08 | 28.75 | 35.00 | 84.14 | 72.20 | 24.57 | 82.80 | 31.29 | 52.13 |
| Student: Qwen3-1.7B-Non-Thinking | ||||||||||
| Math | Code | IF | Overall | |||||||
| Method | AIME24 | AIME25 | HMMT Feb. | HMMT Nov. | Human Eval+ | MBPP+ | LCB-v6 | IFEval | IFBench | Avg. |
| Baseline Models | ||||||||||
| Qwen3-1.7B | 12.50 | 11.67 | 5.83 | 4.58 | 60.40 | 52.40 | 12.00 | 67.65 | 16.67 | 27.08 |
| Qwen3-4B | 23.33 | 10.00 | 10.00 | 7.08 | 82.30 | 67.20 | 18.29 | 81.88 | 24.14 | 36.02 |
| RL Teacher | 58.33 | 52.08 | 28.75 | 35.00 | 84.14 | 72.20 | 24.57 | 82.80 | 31.29 | 52.13 |
| Student: Qwen3-1.7B-Non-Thinking | ||||||||||
| Student: Qwen3-1.7B | Student: Qwen3-4B | |||||||
|---|---|---|---|---|---|---|---|---|
| Routing metric | Math | Code | IF | Overall | Math | Code | IF | Overall |
| Unlabeled Training Set | ||||||||
| Mean Aggregation | 20.31 | 46.83 | 44.33 | 34.49 | 37.29 | 59.68 | 51.34 | 47.88 |
| ExpertAlign- Uniform | 23.12 | 47.77 | 42.81 | 35.71 | 41.46 | 61.78 | 51.31 | 50.42 |
| ExpertAlign- Cosine | 26.04 | 50.33 | 46.02 | 38.58 | 46.35 | 62.93 | 54.81 | 53.76 |
| Labeled Training Set | ||||||||
| Student: Qwen3-1.7B | Student: Qwen3-4B | |||||||
|---|---|---|---|---|---|---|---|---|
| Method | Math | Code | IF | Overall | Math | Code | IF | Overall |
| MOPD-Router w/ ExpertAlign | 26.98 | 53.56 | 46.56 | 40.19 | 46.98 | 64.22 | 55.31 | 54.58 |
| w/o Token-Level Routing | 22.50 | 50.25 | 44.18 | 36.57 | 44.48 | 62.37 | 54.24 | 52.61 |
| w/o Cross-Domain Teachers | 26.04 | 51.84 | 40.20 | 37.79 | 45.83 | 60.50 | 50.80 | 51.83 |
| Standard MOPD | 25.00 | 50.02 | 43.80 | 37.52 | 42.92 | 60.46 | 51.31 | 50.63 |
| Method | Avg. resp. length | GPU-hours | E2E ms/token | Rel. E2E cost |
|---|---|---|---|---|
| Standard MOPD | 4,893 | 692.9 | 0.32184 | |
| MOPD-Router | ||||
| w/ Entropy | 3,278 | 524.2 | 0.36344 | |
| w/ Novelty | 3,496 | 539.6 | 0.35079 | |
| w/ ExpertAlign | 4,278 | 716.9 | 0.38086 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| (a) MOPD experiments | |
|---|---|
| Configuration | Value |
| Training epochs | 3 |
| Train batch size | 1,024 |
| Mini-batch size | 1,024 |
| Learning rate | , constant |
| Gradient clipping | 1.0 |
| Domain label | Count | Proportion |
|---|---|---|
| Instruction following | 30,545 | 50.91% |
| Mathematics | 27,476 | 45.79% |
| Code | 1,979 | 3.30% |
| AI-generated label | Math | Code | IF | Total |
|---|---|---|---|---|
| Mathematics | 99 | 1 | 0 | 100 |
| Code | 24 | 76 | 0 | 100 |
| Instruction following | 8 | 5 | 87 | 100 |
| Setting | Standard MOPD | Mean Aggregation | Open-MOPD | ExpertAlign |
|---|---|---|---|---|
| Unlabeled, 1.7B | ||||
| Unlabeled, 4B | ||||
| Labeled, 1.7B | ||||
| Labeled, 4B |
| Method | Math | Code | IF | Overall |
|---|---|---|---|---|
| Standard MOPD | ||||
| Open-MOPD | ||||
| ExpertAlign |
| Predictive entropy | Routing weight (%) | Top-1 token share (%) | ||||
|---|---|---|---|---|---|---|
| Teacher | 1.7B | 4B | 1.7B | 4B | 1.7B | 4B |
| Mathematics | 0.337 | 0.294 | 36.79 | 36.57 | 69.02 | 70.31 |
| Code | 0.404 | 0.366 | 33.53 | 33.36 | 8.97 | 8.69 |
| Instruction Following | 0.421 | 0.390 | 29.68 | 30.07 | 22.01 | 21.00 |
| Routing metric | Math | Code | IF | Overall |
|---|---|---|---|---|
| Mean Aggregation | 20.31 | 46.83 | 44.33 | 34.49 |
| Entropy | 21.35 | 47.95 | 41.27 | 34.65 |
| w/ Batch Calibration | 21.25 | 46.42 | 44.60 | 34.83 |
| Student: Qwen3-1.7B | Student: Qwen3-4B | |||||||
|---|---|---|---|---|---|---|---|---|
| Routing metric | Math | Code | IF | Overall | Math | Code | IF | Overall |
| Novelty | 22.92 | 49.47 | 45.13 | 36.70 | 40.73 | 61.70 | 52.34 | 50.30 |
| w/o Accessibility | 21.56 | 44.38 | 43.82 | 34.12 | 40.62 | 61.86 | 54.81 | 50.85 |
| Student | Teacher | Top- overlap | -mass | -mass | Overlap JS | |
|---|---|---|---|---|---|---|
| Qwen3-1.7B | Math | 0.728 | 0.990 | 0.992 | 0.991 | 0.0386 |
| Code | 0.736 | 0.991 | 0.990 | 0.990 | 0.0348 | |
| IF | 0.671 | 0.984 | 0.978 | 0.981 | 0.0606 | |
| Qwen3-4B | Math | 0.874 | 0.995 | 0.997 | 0.996 | 0.0158 |
| Code | 0.878 | 0.995 | 0.996 | 0.995 | 0.0131 | |
| IF | 0.762 | 0.991 | 0.989 | 0.990 | 0.0385 |