Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
Organizations: Nanyang Technological University · Yale University · University of Manchester
Abstract
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.
Figures & tables
| Math (avg@64) | Code (avg@6) | IF (avg@16) | |||||
| Method | AIME25 | AIME26 | LCB v5 | LCB v6 | IFEval | IFBench | Total |
| Initial student | 57.7 | 62.6 | 54.9 | 51.4 | 82.4 | 33.8 | 57.1 |
| RL experts | |||||||
| Math expert | 59.7 | 69.4 | 53.4 | 49.4 | 82.3 | 34.9 | 58.2 |
| Code expert | 58.5 | 65.4 | 58.5 | 54.3 | 82.3 | 34.6 | 58.9 |
| IF expert | 54.6 | 60.5 | 54.0 | 51.5 | 86.5 | 42.1 | 58.2 |
| Qwen3.5-4B | Qwen3.5-2B | |||||||
| Method | Math | Code | IF | Total | Math | Code | IF | Total |
| Initial student | 52.2 | 38.1 | 52.8 | 47.7 | 17.6 | 11.3 | 43.3 | 24.0 |
| RL experts | ||||||||
| Math expert | 54.2 | 39.7 | 54.0 | 49.3 | 22.1 | 13.4 | 44.7 | 26.7 |
| Code expert | 51.6 | 47.4 | 53.6 | 50.8 | 20.7 | 21.6 | 44.4 | 28.9 |
| IF expert | 50.7 | 38.1 | 60.6 | 49.8 | 15.6 | 12.5 | 52.8 | 27.0 |
| Variant | Teacher rule | Signal change | 9B | 4B | 2B |
|---|---|---|---|---|---|
| Uniform pool | Mixture | Unchanged | -0.41 | -0.69 | +0.30 |
| Dynamic router | Response | Unchanged | -0.25 | +0.28 | -0.07 |
| Label-routed | Domain | Unchanged | +0.00 | +0.00 | +0.00 |
| Annealed injection | Domain | Early imitation | -0.24 | +1.11 | +0.30 |
| Math 2 only | Domain | Math scale | — | +1.20 | +0.27 |
| IF 0.25 only | Domain | IF scale | — | +1.79 | +2.19 |
| Method | Total | Math tokens | Code tokens | IF tokens | At cap (%) |
| Qwen3.5-9B | |||||
| SeqKD-SFT | 62.6 | 5,547 | 5,896 | 422 | 1.5 |
| ParamMerge-TA | 60.5 | 5,083 | 5,479 | 430 | 1.7 |
| Label-routed | 58.4 | 8,637 | 6,846 | 450 | 9.8 |
| DN-MOPD | 59.6 | 7,032 | 6,427 | 441 | 5.2 |
| Qwen3.5-4B | |||||
| Method | 80 | 160 | Total |
|---|---|---|---|
| Qwen3.5-9B | |||
| Single teacher (IF) | 57.3 | 57.6 | +0.4 |
| Label-routed | 58.4 | 59.4 | +1.0 |
| DN-MOPD | 59.6 | 60.7 | +1.2 |
| Qwen3.5-4B | |||
| Single teacher (code) | 50.9 | 51.2 | +0.4 |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Size | Math expert | Code expert | IF expert |
|---|---|---|---|
| 9B | 250 | 200 | 400 |
| 4B | 320 | 300 | 400 |
| 2B | 400 | 400 | 400 |
| Math (avg@64) | Code (avg@6) | IF (avg@16) | |||||
| Method | AIME25 | AIME26 | LCB v5 | LCB v6 | IFEval | IFBench | Total |
| Initial student | 48.4 | 56.0 | 39.1 | 37.0 | 78.3 | 27.2 | 47.7 |
| RL experts | |||||||
| Math expert | 50.6 | 57.9 | 41.1 | 38.4 | 79.0 | 29.0 | 49.3 |
| Code expert | 48.5 | 54.6 | 49.5 | 45.2 | 79.2 | 28.0 | 50.8 |
| IF expert | 47.9 | 53.4 | 38.4 | 37.8 | 84.3 | 36.9 | 49.8 |
| Math (avg@64) | Code (avg@6) | IF (avg@16) | |||||
| Method | AIME25 | AIME26 | LCB v5 | LCB v6 | IFEval | IFBench | Total |
| Initial student | 17.7 | 17.4 | 9.0 | 13.5 | 62.8 | 23.9 | 24.0 |
| RL experts | |||||||
| Math expert | 20.9 | 23.3 | 10.1 | 16.7 | 64.1 | 25.3 | 26.7 |
| Code expert | 19.6 | 21.8 | 20.1 | 23.0 | 63.5 | 25.3 | 28.9 |
| IF expert | 16.0 | 15.2 | 10.0 | 15.0 | 74.6 | 30.9 | 27.0 |
| Math (avg@64) | Code (avg@6) | IF (avg@16) | |||||
| Method | AIME25 | AIME26 | LCB v5 | LCB v6 | IFEval | IFBench | Total |
| Initial student | 45.5 | 50.7 | 45.0 | 42.2 | 82.0 | 34.0 | 49.9 |
| RL experts | |||||||
| Math expert | 56.2 | 64.6 | 45.5 | 44.3 | 82.5 | 35.0 | 54.7 |
| Code expert | 48.3 | 55.7 | 52.1 | 46.4 | 82.2 | 34.3 | 53.2 |
| IF expert | 45.2 | 50.0 | 45.1 | 43.9 | 86.4 | 41.9 | 52.1 |
| Math (avg@64) | Code (avg@6) | IF (avg@16) | |||||
| Method | AIME25 | AIME26 | LCB v5 | LCB v6 | IFEval | IFBench | Total |
| Initial student | 37.8 | 43.0 | 30.6 | 32.1 | 78.5 | 27.6 | 41.6 |
| RL experts | |||||||
| Math expert | 49.4 | 57.1 | 34.0 | 33.2 | 79.4 | 28.9 | 47.0 |
| Code expert | 42.3 | 47.7 | 43.5 | 41.5 | 79.3 | 28.3 | 47.1 |
| IF expert | 37.4 | 41.8 | 31.3 | 32.9 | 84.4 | 36.7 | 44.1 |
| Math (avg@64) | Code (avg@6) | IF (avg@16) | |||||
| Method | AIME25 | AIME26 | LCB v5 | LCB v6 | IFEval | IFBench | Total |
| Initial student | 11.1 | 8.0 | 7.6 | 12.0 | 63.6 | 24.1 | 21.1 |
| RL experts | |||||||
| Math expert | 19.6 | 21.6 | 8.4 | 15.2 | 63.5 | 25.2 | 25.6 |
| Code expert | 17.8 | 17.3 | 15.7 | 19.9 | 63.9 | 24.8 | 26.6 |
| IF expert | 12.1 | 8.2 | 8.0 | 12.8 | 74.8 | 30.5 | 24.4 |
| Size | Math | Code | IF |
|---|---|---|---|
| 9B | 792/900 | 724/900 | 812/900 |
| 4B | 794/900 | 686/900 | 763/900 |
| 2B | 599/900 | 363/900 | 771/900 |
| Method | Updates | AIME25 | AIME26 | LCB v5 | LCB v6 | IFEval | IFBench | Total |
|---|---|---|---|---|---|---|---|---|
| Qwen3.5-9B | ||||||||
| Label | 160 | 57.0 | 63.9 | 58.9 | 51.0 | 84.9 | 40.6 | 59.4 |
| DN-MOPD | 160 | 58.9 | 70.1 | 56.7 | 52.5 | 85.1 | 41.1 | 60.7 |
| Single IF | 160 | 55.8 | 61.7 | 52.4 | 50.1 | 85.1 | 40.6 | 57.6 |
| Annealed injection | 80 | 55.8 | 63.6 | 54.3 | 52.0 | 84.6 | 38.6 | 58.2 |
| Qwen3.5-4B | ||||||||
| Method | Updates | AIME25 | AIME26 | LCB v5 | LCB v6 | IFEval | IFBench | Total |
|---|---|---|---|---|---|---|---|---|
| Qwen3.5-9B | ||||||||
| Label | 160 | 47.3 | 54.5 | 48.6 | 44.5 | 85.1 | 40.3 | 53.4 |
| DN-MOPD | 160 | 52.2 | 62.3 | 50.3 | 45.8 | 85.1 | 41.2 | 56.2 |
| Single IF | 160 | 45.9 | 51.3 | 43.7 | 43.2 | 85.3 | 41.0 | 51.8 |
| Annealed injection | 80 | 46.4 | 52.0 | 45.1 | 42.9 | 84.5 | 38.2 | 51.5 |
| Qwen3.5-4B | ||||||||
| Contrast (updates) | Total | 95% CI |
|---|---|---|
| Qwen3.5-9B | ||
| DN Label (80) | ||
| DN Label (160) | ||
| Annealed Label (80) | ||
| DN Annealed (80) | ||
| Qwen3.5-4B | ||
| Contrast (updates) | Total | 95% CI |
|---|---|---|
| Qwen3.5-9B | ||
| DN Label (80) | ||
| DN Label (160) | ||
| Annealed Label (80) | ||
| DN Annealed (80) | ||
| Qwen3.5-4B | ||
| Size | Strongest single | Label single | DN-MOPD single |
|---|---|---|---|
| 16K evaluation | |||
| 9B | Code | ||
| 4B | Code | ||
| 2B | Code | ||
| 8K evaluation | |||
| 9B | Math | ||
| Cap | DN-MOPD code (80) | DN-MOPD code (160) | Label code (160) | Code: 160 80 |
|---|---|---|---|---|
| 16K | ||||
| 8K |
| Size | Updates | Math | Code | IF |
|---|---|---|---|---|
| 16K evaluation | ||||
| 9B | 80 | |||
| 9B | 160 | |||
| 4B | 80 | |||
| 4B | 160 | |||
| 2B | 80 | |||
| Method | Math tokens | Code tokens | IF tokens | Cap hits (%) |
| Qwen3.5-9B (16K evaluation) | ||||
| Label | 8,637 | 6,846 | 450 | 9.8 |
| DN-MOPD | 7,032 | 6,427 | 441 | 5.2 |
| Qwen3.5-4B (16K evaluation) | ||||
| Label | 8,653 | 8,238 | 523 | 12.4 |
| DN-MOPD | 6,146 | 7,649 | 522 | 5.3 |
| Size | Seed 42 | Seed 43 | Seed 44 | Three-seed mean |
|---|---|---|---|---|
| 16K evaluation | ||||
| 9B | ||||
| 4B | ||||
| 2B | ||||
| 8K evaluation | ||||
| 9B | ||||
| Size | Control | Weights | Label | DN-MOPD | Math Label |
|---|---|---|---|---|---|
| 16K evaluation | |||||
| 9B | Update 0 | 1.77/0.89/0.25 | |||
| 9B | Global | 2/1/0.25 | |||
| 4B | Update 0 | 2.01/0.81/0.34 | |||
| 4B | Global | 2/1/0.25 | |||
| 4B | Math only | 2/1/1 | |||
| Size | Initial | Label | Math-only OPD | DN-MOPD | Label initial | DN-MOPD Label |
| 16K evaluation | ||||||
| 9B | 95.84 | 95.76 | 97.30 | 96.50 | ||
| 4B | 94.46 | 94.08 | 95.84 | 95.33 | ||
| 2B | 76.91 | 77.21 | 82.93 | 83.16 | ||
| 8K evaluation | ||||||
| 9B | 94.08 | 94.38 | 96.80 | 95.83 | ||
| Domain | Update 0 | Raw factor | Applied factor | Clipped | Tokens (%) | |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | ||||||
| Math | 0.57 | 0.65 | 1.54 | 1.54 | 0/115 | 48.9 |
| Code | 1.13 | 0.64 | 1.57 | 1.57 | 0/115 | 50.1 |
| IF | 4.41 | 7.37 | 0.14 | 0.25 | 115/115 | 0.9 |
| Qwen3.5-4B | ||||||
| Math | 0.50 | 0.49 | 2.03 | 2.03 | 0/134 | 46.9 |
| Domain | Tokens (%) | Pooled variance (%) | Would-be under Label | Applied by DN-MOPD |
|---|---|---|---|---|
| Qwen3.5-9B | ||||
| Math | 48.4 | 20.8 | 1.49 | 1.59 |
| Code | 50.6 | 28.3 | 0.99–1.00 | 1.53 |
| IF | 1.0 | 50.9 | 0.25 | 0.25 |
| Qwen3.5-4B | ||||
| Math | 47.2 | 10.9 | 1.91–1.93 | 2.15 |
| Domain | Split-half cos. | Share, equal (%) | Share, DN (%) | |
|---|---|---|---|---|
| Initial student | ||||
| Math | 3.6 | 1 | 16 | |
| Code | 10.6 | 5 | 20 | |
| IF | 51.7 | 94 | 64 | |
| Label, 80 updates | ||||
| Math | 2.6 | 0 | 0 | |
| Size | Math expert | Math-only OPD | Label | DN-MOPD |
|---|---|---|---|---|
| Evaluation cap: 16K | ||||
| 9B | ||||
| 4B | ||||
| 2B | ||||
| Evaluation cap: 8K | ||||
| 9B | ||||
| Metric | Label | DN-MOPD | (pp) | 95% CI |
|---|---|---|---|---|
| Math | 53.67 | 53.41 | ||
| Code | 38.36 | 38.27 | ||
| IF | 56.13 | 56.53 | ||
| Total | 49.39 | 49.41 |