Domain experts trained from a shared checkpoint can transfer their specialized capabilities to a single student through on-policy distillation (OPD). Existing research primarily focuses on improving this merging process, while the algorithms used to train the experts have received limited systematic comparison. We investigate which training algorithm produces teachers better suited to OPD through controlled single-teacher comparisons of supervised fine-tuning (SFT) and reinforcement learning (RL) across Agentic, Reasoning, and Perception. Teachers and students share the same Qwen3.5-9B initialization, and the two teacher types are compared at similar task performance. Our experiments show that RL teachers yield stronger students and higher recovery of teacher performance gains across all three domains. At their best checkpoints, RL-guided students outperform SFT-guided students by 4.27, 1.50, and 0.86 percentage points, respectively. In Agentic, the best SFT-guided student recovers only 44.44% of its teacher's performance gain over the base model, whereas the best RL-guided student recovers 115.00%, surpassing its teacher. Our further analysis shows that RL teachers undergo smaller parameter displacements from the shared initialization than SFT teachers. These findings support the hypothesis that RL teachers' smaller departures from the student's starting point facilitate learning through OPD, resulting in stronger students.
Figures & tables
Domain
Evaluation benchmarks
Task metric
Agentic
MobileWorld ( Kong et al., 2025 )
Task success rate
Reasoning
Geo3K test set ( Lu et al., 2021 ) , MathVista ( Lu et al., 2024 ) , MathVision ( Wang et al., 2024 ) , MathVerse ( Zhang et al., 2024 )
Answer accuracy
Perception
ScreenSpot-v2 ( Wu et al., 2024 ) , ScreenSpot-Pro ( Li et al., 2025 ) , UI-Vision ( Nayak et al., 2025 ) , MMBench-GUI L2 ( Wang et al., 2025 ) , OSWorld-G, OSWorld-G-Refine ( Xie et al., 2025 )
Grounding accuracy
Table 1: Evaluation benchmarks and metrics.
Figure 1: MobileWorld success rates across OPD checkpoints.
Figure 2: Geo3K accuracies across OPD checkpoints.
Figure 3: Mean accuracies across the six Perception benchmarks at OPD checkpoints. The horizontal axis is expanded for the first ten steps.
Figure 4: Teacher displacement and OPD performance. (a) Comparison across domains. (b) Agentic accuracy trajectories. (c) Gain recovery at each group’s best student checkpoint.
Figure 5: Teacher–student Top-20 overlap on student rollouts. Step 0 denotes the base student before its first update. Shading marks the early steps compared in the text.
Table 4: Teacher training settings. Batch: demonstrations (SFT) or prompts × sampled responses (RL) per update. Max tokens: total (SFT) or prompt/response (RL). Learning rates are configured values.
Parameter
Agentic
Reasoning
Perception
Optimizer
AdamW
AdamW
AdamW
Learning rate
10−6
1.875×10−6
1.875×10−6
Weight decay
0
0.01
0.01
Input batch
128
32
112
Responses per prompt ( n )
8
1
1
Warmup (updates)
0
3
3
Appendix
Table 5: Student OPD training settings.
SFT
RL
Checkpoint
R1
R2
R3
Mean
R1
R2
R3
Mean
Teacher
28
23
29
22.79
26
28
28
23.36
Qwen3.5-9B
22
20
20
17.66
22
20
20
17.66
Step 8
21
28
21
19.94
17
23
25
18.52
Step 16
19
19
20
16.52
25
18
23
18.80
Step 24
23
22
22
19.09
23
29
24
21.65
Appendix
Table 6: MobileWorld results across OPD checkpoints. R1–R3: successes per round; Mean: success rate (%).
Model
Step
Geo3K
MathVista
MathVision
MathVerse
Avg.
Qwen3.5-9B
–
79.70
84.00
57.24
70.18
72.78
SFT teacher
–
88.35
80.80
53.95
67.89
72.75
RL teacher
–
84.19
83.60
57.57
73.10
74.61
Student ← SFT
4
86.36
83.60
51.32
71.45
73.18
8
86.19
81.50
55.59
64.85
72.03
12
89.02
82.10
56.91
68.15
74.04
Appendix
Table 7: Reasoning accuracy (%) across OPD checkpoints.
Model
Step
ScreenSpot- v2
ScreenSpot- Pro
UI-Vision
MMBench- GUI L2
OSWorld- G
OSWorld- G-R
Avg.
Qwen3.5-9B
–
91.75
63.00
26.92
80.41
61.35
67.73
65.19
SFT teacher
–
92.92
65.15
33.67
83.87
62.94
70.92
68.24
RL teacher
–
93.16
65.28
30.52
82.39
63.65
71.28
67.71
Student ← SFT
2
91.12
62.87
26.14
79.53
61.52
68.44
64.94
4
91.35
63.19
27.52
81.17
61.52
68.44
65.53
6
92.61
63.38
32.94
82.66
63.30
67.91
67.13
Appendix
Table 8: Perception accuracy (%) across OPD checkpoints.
Multi-teacher on-policy distillation (MOPD) combines independently developed domain teachers into a single student by distilling their predictions on student-generated samples. We study a setting where teachers share a reference model but undergo different post-training procedures, and find that MOPD can struggle to recover some teacher capabilities. Because distillation occurs on student-generated prefixes, the student initialization can strongly affect subsequent recovery. However, initial benchmark performance is not a reliable predictor of a good MOPD initialization. For example, merge initialization can start below SFT warm-up yet finish higher after MOPD. We further find that effective merging depends on both the relative teacher contributions and the overall merge scale, with some strong configurations lying outside the simplex of convex parameter averaging. Thus, selecting a good merge initialization requires evaluating not only its immediate performance but also the learning it enables under MOPD, making one-shot coefficient search difficult. We propose Iterative Merging for MOPD (IM-MOPD), which starts from a uniform merge and progressively adds task-vector increments for under-recovered domains during distillation. In a 5-domain setting, IM-MOPD achieves higher average normalized recovery than MOPD with either uniform merge initialization or SFT warm-up, showing that effective teacher contributions can be determined progressively during training.
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.
Xin Li, Hao Jiang, Xin Gao +6
Nanyang Technological University · Yale University · University of Manchester
Reinforcement learning can substantially improve a reasoning teacher, but it is unclear which of those improvements survive when the teacher supervises a smaller on-policy student. We study this question in mathematical reasoning by comparing teacher lineages before and after GRPO, multiple student scales, direct GRPO, and several on-policy distillation objectives. The central finding is that transfer is structured rather than scalar: teacher strength alone does not make dense distillation competitive, while an RL-improved teacher creates useful but metric-dependent student gains. This motivates Transfer-Stratified On-Policy Distillation (TS-OPD), which screens training problems by the joint sampled success of the student and teacher, routes acquisition problems to gated forward KL, routes consolidation problems to gated reverse KL, and adds an entropy brake to protect sampled coverage. Across the main comparison, TS-OPD is the strongest student objective for macro average correctness with the GRPO-improved teacher, while pass@K remains more mixed. Ablations show that the gains come from routing and token gating rather than skipping problems. These results support a transfer-aware view of OPD: stronger teachers help when the supervision direction and token budget match the student's observed ability, not merely because the teacher endpoint is stronger.
Xiaoyu Chen, Bo Shao, Tiangang Zhu +5
Institute of Software, Chinese Academy of Sciences · Microsoft · Work done during an internship at Microsoft +1