From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation
Organizations: University of Illinois Urbana-Champaign · Princeton University · Westlake University
Abstract
Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97% of FP32 master weights differ from initialization, but only 7--11% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.
Figures & tables
| Loss averaging | Math | Code | IF | Science | Mean |
|---|---|---|---|---|---|
| Initial student | 68.80 | 16.40 | 18.33 | 32.20 | 33.93 |
| Adam, PG loss | |||||
| Domain responses | 73.36 | 18.26 | 28.07 | 35.94 | 38.91 |
| Domain tokens | 73.28 | 19.24 | 27.07 | 34.93 | 38.63 |
| Global tokens | 76.32 | 17.91 | 27.13 | 33.81 | 38.79 |
| Adam, top-16 intersection KL loss | |||||
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Upstream dataset | Rows |
|---|---|---|
| Mathematics | BytedTsinghua-SIA/DAPO-Math-17k and Skywork/Skywork-OR1-RL-Data | 22,056 |
| Code | nvidia/Nemotron-RL-coding-competitive_coding | 19,169 |
| Instruction following | nvidia/Nemotron-RL-instruction_following | 16,575 |
| Science | nvidia/Nemotron-RL-knowledge-mcqa (STEM multiple-choice component) | 19,670 |
| Domain | Responses per prompt | Batch Size | Updates |
|---|---|---|---|
| Mathematics | 8 | 16 | 500 |
| Code | 16 | 16 | 500 |
| Instruction following | 16 | 16 | 500 |
| Science | 4 | 16 | 1000 |