Domain experts trained from a shared checkpoint can transfer their specialized capabilities to a single student through on-policy distillation (OPD). Existing research primarily focuses on improving this merging process, while the algorithms used to train the experts have received limited systematic comparison. We investigate which training algorithm produces teachers better suited to OPD through controlled single-teacher comparisons of supervised fine-tuning (SFT) and reinforcement learning (RL) across Agentic, Reasoning, and Perception. Teachers and students share the same Qwen3.5-9B initialization, and the two teacher types are compared at similar task performance. Our experiments show that RL teachers yield stronger students and higher recovery of teacher performance gains across all three domains. At their best checkpoints, RL-guided students outperform SFT-guided students by 4.27, 1.50, and 0.86 percentage points, respectively. In Agentic, the best SFT-guided student recovers only 44.44% of its teacher's performance gain over the base model, whereas the best RL-guided student recovers 115.00%, surpassing its teacher. Our further analysis shows that RL teachers undergo smaller parameter displacements from the shared initialization than SFT teachers. These findings support the hypothesis that RL teachers' smaller departures from the student's starting point facilitate learning through OPD, resulting in stronger students.
Figures & tables
Domain
Evaluation benchmarks
Task metric
Agentic
MobileWorld ( Kong et al., 2025 )
Task success rate
Reasoning
Geo3K test set ( Lu et al., 2021 ) , MathVista ( Lu et al., 2024 ) , MathVision ( Wang et al., 2024 ) , MathVerse ( Zhang et al., 2024 )
Answer accuracy
Perception
ScreenSpot-v2 ( Wu et al., 2024 ) , ScreenSpot-Pro ( Li et al., 2025 ) , UI-Vision ( Nayak et al., 2025 ) , MMBench-GUI L2 ( Wang et al., 2025 ) , OSWorld-G, OSWorld-G-Refine ( Xie et al., 2025 )
Grounding accuracy
Table 1: Evaluation benchmarks and metrics.
Figure 1: MobileWorld success rates across OPD checkpoints.
Figure 2: Geo3K accuracies across OPD checkpoints.
Figure 3: Mean accuracies across the six Perception benchmarks at OPD checkpoints. The horizontal axis is expanded for the first ten steps.
Figure 4: Teacher displacement and OPD performance. (a) Comparison across domains. (b) Agentic accuracy trajectories. (c) Gain recovery at each group’s best student checkpoint.
Figure 5: Teacher–student Top-20 overlap on student rollouts. Step 0 denotes the base student before its first update. Shading marks the early steps compared in the text.
Table 4: Teacher training settings. Batch: demonstrations (SFT) or prompts × sampled responses (RL) per update. Max tokens: total (SFT) or prompt/response (RL). Learning rates are configured values.
Parameter
Agentic
Reasoning
Perception
Optimizer
AdamW
AdamW
AdamW
Learning rate
10−6
1.875×10−6
1.875×10−6
Weight decay
0
0.01
0.01
Input batch
128
32
112
Responses per prompt ( n )
8
1
1
Warmup (updates)
0
3
3
Appendix
Table 5: Student OPD training settings.
SFT
RL
Checkpoint
R1
R2
R3
Mean
R1
R2
R3
Mean
Teacher
28
23
29
22.79
26
28
28
23.36
Qwen3.5-9B
22
20
20
17.66
22
20
20
17.66
Step 8
21
28
21
19.94
17
23
25
18.52
Step 16
19
19
20
16.52
25
18
23
18.80
Step 24
23
22
22
19.09
23
29
24
21.65
Appendix
Table 6: MobileWorld results across OPD checkpoints. R1–R3: successes per round; Mean: success rate (%).
Model
Step
Geo3K
MathVista
MathVision
MathVerse
Avg.
Qwen3.5-9B
–
79.70
84.00
57.24
70.18
72.78
SFT teacher
–
88.35
80.80
53.95
67.89
72.75
RL teacher
–
84.19
83.60
57.57
73.10
74.61
Student ← SFT
4
86.36
83.60
51.32
71.45
73.18
8
86.19
81.50
55.59
64.85
72.03
12
89.02
82.10
56.91
68.15
74.04
Appendix
Table 7: Reasoning accuracy (%) across OPD checkpoints.
Model
Step
ScreenSpot- v2
ScreenSpot- Pro
UI-Vision
MMBench- GUI L2
OSWorld- G
OSWorld- G-R
Avg.
Qwen3.5-9B
–
91.75
63.00
26.92
80.41
61.35
67.73
65.19
SFT teacher
–
92.92
65.15
33.67
83.87
62.94
70.92
68.24
RL teacher
–
93.16
65.28
30.52
82.39
63.65
71.28
67.71
Student ← SFT
2
91.12
62.87
26.14
79.53
61.52
68.44
64.94
4
91.35
63.19
27.52
81.17
61.52
68.44
65.53
6
92.61
63.38
32.94
82.66
63.30
67.91
67.13
Appendix
Table 8: Perception accuracy (%) across OPD checkpoints.