Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97% of FP32 master weights differ from initialization, but only 7--11% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.
Figures & tables
Figure 1: From supervision choices to parameter updates and task gains. (a) Averaging changes gradients more than Adam updates. (b) Shared momentum aligns teacher updates. (c) BF16 rounding hides widespread FP32 changes. (d) Local gradient fidelity and one-step BF16 changes, followed by mathematics gains in separate 500-step training runs. DR/GT denote domain response/global token averaging.
Loss averaging
Math
Code
IF
Science
Mean
Initial student
68.80
16.40
18.33
32.20
33.93
Adam, PG loss
Domain responses
73.36 ±0.79
18.26 ±0.71
28.07 ±1.19
35.94 ±1.42
38.91 ±0.69
Domain tokens
73.28 ±0.69
19.24 ±0.83
27.07 ±1.01
34.93 ±1.35
38.63 ±0.65
Global tokens
76.32 ±0.73
17.91 ±0.66
27.13 ±1.24
33.81 ±1.54
38.79 ±0.70
Adam, top-16 intersection KL loss
Table 1: MOPD scores by loss averaging and distillation loss. Scores (%) are mean ± standard deviation across five training seeds; Mean equally weights the four task means. Domain responses and Domain tokens assign equal domain weights, averaging responses equally or by length, respectively. Global token averaging weights all response tokens equally.
Figure 2: Loss averaging changes gradients, while shared Adam state aligns updates. (a) Domain token shares under global token averaging. (b) Gradient and update cosines for six batches (three per checkpoint). (c–f) Per-task learning curves with the top-64 intersection KL loss. Error bars show standard deviations from three training seeds at steps 100 and 250, and five at step 500. GT/DT/DR denote global token, domain token, and domain response averaging. S/T mark the initial student and domain teacher.
Figure 3: Optimizer processing attenuates differences between averaging rules. (a) Gradient retention and clipping rates. (b) Adam step norms with saved or reset moments (step count retained). (c) Held-out reverse KL after minus before one step; SGD matches each rule’s saved-Adam step norm. Open points: batches; filled points/bars: mean/SD. (d–f) Raw gradient norms, FP32 step norms, and cumulative FP32 displacement under Adam with the top-64 intersection KL loss. Faint/solid traces show individual steps/25-step medians; dotted line: clipping threshold. GT/DT/DR denote global token, domain token, and domain response averaging.
Figure 4: Loss averaging influences the effect of doubling the maximum training response length from 4,096 (4K) to 8,192 (8K) tokens. (a) Domain token shares and total response tokens (millions) during 100 steps. (b) Relative 4K/8K changes in gradients ( g ), clipped gradients, and Adam steps ( U ); groups name source 4K run and training step, rows mark averaging rules. Each group averages three paired batches. (c) Task score differences (8K minus 4K, percentage points); adjacent points show equal-domain means at steps 50 (open) and 100 (filled). GT/DT/DR denote global token, domain token, and domain response averaging.
Figure 5: BF16 rounding hides small changes; shared Adam momentum aligns student updates. (a) FP32/BF16 parameter fractions: changed from initialization (NZ), or carrying 90% of squared change (E90). S/M: math-only/MOPD training; PG: sampled-token loss; Top-64: intersection KL loss. M-DR/M-DT/M-GT: Top-64 with response/domain-token/global-token averaging; other rows average responses. (b,c) Pairwise update cosines: own-domain responses (Routed) or shared prefixes (Common). Saved retains trained Adam state; −U0 subtracts the zero-gradient step; m=0 resets the first moment. Small points average teacher pairs within batches; large points average batches.
Figure 6: Top-64 closely matches full-vocabulary gradient directions in Qwen. (a) Rows: training loss@step. Columns compare one-step BF16 changed fractions for different losses at fixed parameters, responses, and Adam state. (b) PG-to-full gradient and update cosines across training snapshots; gray shows Top-64-to-full ranges. (c) Score differences from PG under response averaging. PG: sampled-token loss; Top-16/Top-64: intersection KL losses; Teacher-64: teacher-supported loss without renormalization (Appendix D ); Full: full-vocabulary KL loss. S/M: mathematics-only/MOPD training.
Figure 7: High mean coverage can hide gradient error in SmolLM. (a,b) Mean intersection coverage and fraction below 90%; domains, responses within domains, and positions within responses are equally weighted. Bands: response-cluster 95% intervals. (c,d) Gradient cosine and relative error against full-vocabulary KL after training; relative error is the gradient-difference norm divided by the full-gradient norm. Dashed lines: k=64 . Thin/bold lines: batches/means. Initial: initialization; DR500: 500 steps with top-64 intersection KL and response averaging.
Figure 8: Sampling improves gradient alignment; Adam history affects BF16 changes. (a) Mean gradient cosine to full KL. (b,c) Changed BF16 fractions with shared axes; inset repeats (b)’s categories. PG: sampled-token loss; Top-64: intersection KL; Teacher-64: teacher-supported loss; Full: full-vocabulary KL. Saved retains Adam state; reset- m clears its first moment; fresh Adam clears both moments and step count. In (b,c), open points show measurements; bars show means.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Upstream dataset
Rows
Mathematics
BytedTsinghua-SIA/DAPO-Math-17k and Skywork/Skywork-OR1-RL-Data
Table 2: Domain data sources and converted pool sizes before stage-specific filtering.
Domain
Responses per prompt
Batch Size
Updates
Mathematics
8
16
500
Code
16
16
500
Instruction following
16
16
500
Science
4
16
1000
Appendix
Table 3: Teacher GRPO setup.
Figure 9: Cumulative BF16 changes remain much sparser than FP32 changes. PG and top-64 trajectories under domain-response averaging in single-teacher and MOPD training, with additional top-64 MOPD runs using domain-token/global-token averaging. Panels (a,b) show nonzero coordinate fractions; panels (c,d) show the fraction of coordinates containing 90% of squared displacement from initialization. Lines connect steps 1, 50, 100, 250, and 500; dashed lines denote single-teacher training. Figure labels use S/M for single-teacher/MOPD training and PG/Top-64 for the PG/top-64 intersection KL losses.
Figure 10: Mean probe-loss reductions can coexist with domain-level increases. Full-vocabulary KL after minus before BF16 writeback, in units of 10−3 . Cells show three-batch means; dots show batches 1042–1044 from left to right. The shared symmetric-log color scale is linear within ±0.1 . SGD is matched to each rule’s saved-state Adam step norm.
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.
Xin Li, Hao Jiang, Xin Gao +6
Nanyang Technological University · Yale University · University of Manchester
Multi-teacher on-policy distillation (MOPD) combines independently developed domain teachers into a single student by distilling their predictions on student-generated samples. We study a setting where teachers share a reference model but undergo different post-training procedures, and find that MOPD can struggle to recover some teacher capabilities. Because distillation occurs on student-generated prefixes, the student initialization can strongly affect subsequent recovery. However, initial benchmark performance is not a reliable predictor of a good MOPD initialization. For example, merge initialization can start below SFT warm-up yet finish higher after MOPD. We further find that effective merging depends on both the relative teacher contributions and the overall merge scale, with some strong configurations lying outside the simplex of convex parameter averaging. Thus, selecting a good merge initialization requires evaluating not only its immediate performance but also the learning it enables under MOPD, making one-shot coefficient search difficult. We propose Iterative Merging for MOPD (IM-MOPD), which starts from a uniform merge and progressively adds task-vector increments for under-recovered domains during distillation. In a 5-domain setting, IM-MOPD achieves higher average normalized recovery than MOPD with either uniform merge initialization or SFT warm-up, showing that effective teacher contributions can be determined progressively during training.
Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one domain suppresses capabilities acquired from another. Inspired by the distinctive update geometry of OPD, we find that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. We therefore propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions. We further develop a lightweight conflict probe to characterize task interactions and guide task ordering, together with a cycling strategy that balances subspace estimation and timely task revisitation. Experiments on representative Code, Reason, and Math tasks show that PMOPD improves every evaluated capability over MOPD, raising the average score across the three tasks by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. These consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.