Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities. We study this trade-off in Multi-Teacher On-Policy Distillation (MOPD), where a specialized model learns from domain and general teachers on its own sampled trajectories. Standard MOPD faces two limitations: ordinary on-policy sampling rarely exposes tokens with large positive teacher--student advantages, and advantage sign alone does not establish whether the proposed update direction is reliable. We propose Uncertainty-Calibrated MOPD (UCMOPD), which addresses these limitations through two complementary mechanisms. Golden-Gain Enhancement combines higher-temperature exploration with a standard-temperature anchor and retains trajectories whose positive learning signal matches or exceeds the prompt-specific anchor. Teacher-Endorsement Filtering then uses centered log-likelihood (CLL) to estimate each retained token's plausibility relative to the teacher's uncertainty and probabilistically preserves updates whose directions are supported by that endorsement. Across role-playing and medical-domain specialization, UCMOPD improves the general-capability average over standard MOPD by 4.48% and 7.86%, respectively, while maintaining vertical-domain performance. Component ablations and diagnostic analyses support the intended roles of the two mechanisms: exposing and selecting stronger positive signals at the trajectory level and validating update directions through teacher endorsement at the token level.
Figures & tables
Advantage direction
Teacher endorsement
Retention tendency
At>0
Strong
More likely to retain reinforcement
At>0
Weak
More likely to filter reinforcement
At<0
Strong
More likely to filter suppression
At<0
Weak
More likely to retain suppression
Table 1: Direction–endorsement decomposition of token updates. CLL implements these retention tendencies continuously rather than assigning hard endorsement labels.
Metric
Role-playing specialization
Medical-domain specialization
Base
SFT
MOPD
SelecTKD
ReOPOLD
CaMOPD
UCMOPD
Base
SFT
MOPD
SelecTKD
ReOPOLD
CaMOPD
UCMOPD
MMLU-Redux
82.39
75.05
78.86
78.60
78.65
79.07
80.56
84.67
85.77
86.21
86.30
86.21
86.21
86.11
MMLU-Pro
68.33
54.45
62.20
61.60
62.36
64.31
65.72
73.08
74.80
74.53
74.67
74.14
74.53
74.04
GPQA-Diamond
61.49
54.29
59.60
60.98
63.51
62.63
63.64
62.50
62.37
60.10
58.46
61.49
61.99
61.74
AIME25
45.21
40.62
47.29
45.00
43.12
45.21
43.54
65.00
41.67
54.37
53.75
52.08
54.37
56.25
ZebraLogic
79.50
31.50
71.10
70.60
71.40
76.30
76.20
84.80
71.60
72.60
72.10
73.20
72.60
73.70
Table 2: Main specialization results for role-playing (left) and medical-domain (right) specialization. Base and SFT denote the pre-specialization and supervised domain-specialized models; the remaining columns are general-capability recovery methods. Bold marks the best recovery-method result per metric. “Gen. Avg.” averages general benchmarks, while “Vertical Avg.” averages the listed vertical benchmarks.
Method
Gen. Avg.
Vertical Avg.
IF-Eval
GPQA
Zebra
Arena-Hard v2
UCMOPD
54.43
45.00
80.96
63.64
76.20
29.35
− Positive-advantage-density filtering
53.32
43.24
80.04
63.13
74.30
28.45
− Dual-temperature sampling
52.17
43.36
77.63
62.63
72.00
23.70
− Teacher-Endorsement Filtering = MOPD
52.09
41.22
78.00
59.60
71.10
22.60
Table 3: Cumulative component removal from UCMOPD on role-playing specialization.
Criterion
Gen. Avg.
Vertical Avg.
AIME25
GPQA
Zebra
Arena-Hard v2
Hard top- k endorsement ( k=3 )
51.95
42.31
43.33
60.73
74.40
23.20
Hard top- k endorsement ( k=1 )
52.01
41.32
42.71
63.26
71.90
22.65
CLL sample mask
52.17
43.36
44.58
62.63
72.00
23.70
Table 4: Token-endorsement criterion comparison. Both criteria retain updates whose direction agrees with estimated teacher endorsement; CLL replaces hard top- k membership with a continuous entropy-centered score.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Role
Checkpoint
Size
Training dtype
Inference dtype
Role-playing
Student
Our CoSER SFT of Qwen3-4B-Instruct-2507
4B
BF16
BF16
Role-playing
Domain teacher
Same as student initialization (frozen)
4B
–
BF16
Role-playing
General teacher
Qwen3-4B-Instruct-2507 (frozen)
4B
–
BF16
Medical
Student
II-Medical-7B-Preview
8B
BF16
BF16
Medical
Domain teacher
Same as student initialization (frozen)
8B
–
BF16
Medical
General teacher
Qwen3-8B (frozen)
8B
–
BF16
Appendix
Table 5: Model checkpoints and precision. “Same as student initialization” means that the domain teacher is a frozen copy of the SFT checkpoint from which the trainable student is initialized.
Configuration
Role-playing
Medical
Domain prompts
10,000
9,926
General prompts
10,000
10,000
Global prompt batch size
128
256
Responses per prompt
8 ( 1+7 )
4 ( 1+3 )
Candidate responses/update
1,024
1,024
Training epochs
3
3
Appendix
Table 6: Training and rollout configuration. Batch sizes are global prompt counts before response replication.
Method
Gen. Avg.
Vertical Avg.
IF-Eval
GPQA
Zebra
Arena-Hard v2
MOPD, 8 rollouts
53.31
43.38
80.22
60.98
72.90
26.90
UCMOPD, 8 rollouts
54.43
45.00
80.96
63.64
76.20
29.35
Appendix
Table 7: Rollout-budget control with eight sampled responses per prompt.
Subset
Positive-token
Mean positive
fraction ↑
advantage ↑
Kept trajectories
0.2168
0.1911
Dropped trajectories
0.1850
0.1158
Kept − dropped
0.0318
0.0752
Appendix
Table 8: Aggregate diagnostics of Golden-Gain Enhancement before Teacher-Endorsement Filtering.
Method
Gen. Avg.
Vertical Avg.
WritingBench
GPQA
Zebra
LCB v5
MOPD
52.09
41.22
80.30
59.60
71.10
35.12
CLL sample mask
52.17
43.36
80.62
62.63
72.00
34.76
CLL weighting
52.34
42.92
80.71
59.09
72.30
31.90
Appendix
Table 9: Implementation comparison for Teacher-Endorsement Filtering without Golden-Gain Enhancement.
Candidate subset
Candidate tokens
Expected keep (%)
Realized keep (%)
Negative advantage
299,307,540
9.88
9.88
Positive advantage
115,874,361
81.88
81.88
All signed candidates
415,181,901
29.97
29.97
Appendix
Table 10: Expected and realized keep rates under signed-scope CLL sample masking. Expected rates are candidate-token-weighted means of wcll ; realized rates are computed from the sampled Bernoulli masks.
Filtering criterion
Gen. Avg.
Vertical Avg.
WritingBench
ZebraLogic
Positive-advantage density ( > )
52.87
41.46
81.26
72.10
Likelihood ( > )
51.31
41.32
79.48
71.40
Positive-advantage density ( ≥ )
53.05
42.66
81.32
72.80
Likelihood ( ≥ )
51.81
41.54
80.47
74.20
Appendix
Table 11: Auxiliary comparison of trajectory filtering criteria under the same dual-temperature setting.
Temperature
Gen. Avg.
Vertical Avg.
IF-Eval
Zebra
LCB v5
LiveBench
1.2
52.30
42.09
79.11
73.90
31.54
60.30
1.5
53.05
42.66
79.48
72.80
34.76
61.30
1.8
52.52
41.10
78.19
72.20
39.78
61.00
2.0
52.62
42.30
79.11
73.30
33.69
62.30
Appendix
Table 12: Exploration-temperature ablation for positive-advantage-density filtering.
Domain specialization can improve LLM behavior, but often weakens the general capabilities inherited from the original model. Recent Multi-Teacher On-Policy Distillation (MOPD) pipelines recover model capabilities by supervising student-generated trajectories with teacher feedback, but typically assume teacher-aligned prompt coverage, requiring prompts to match the teachers' training distributions. This assumption is difficult to satisfy when the general teacher is an open-source model whose post-training data are unknown. Instead of attempting to reconstruct this hidden distribution, we study general capability recovery with readily available proxy general prompts. We identify two failure modes of vanilla MOPD in this incomplete-coverage situation: recovery-preservation counteraction from mixing conflicting recovery and preservation gradients, and weak-signal flattening from uniformly averaging samples with unequal correction demand. We propose \textbf{Counteraction-Aware Multi-Teacher On-Policy Distillation} (\textbf{CaMOPD}), which addresses these issues with decoupled alternating training and gap-based sample selection. CaMOPD allocates dedicated updates to general recovery, periodically performs domain-preservation updates, and selects samples with larger averaged token-level teacher-student log-probability gaps to concentrate correction signals. Across role-play dialogue and medical reasoning QA scenarios, CaMOPD outperforms all baselines in general capability recovery while maintaining domain-specific behavior. Gradient coherence analyses further support the intended effect of CaMOPD in producing more coherent correction signals.
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.
Xin Li, Hao Jiang, Xin Gao +6
Nanyang Technological University · Yale University · University of Manchester
Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, however, select one teacher per example and keep it fixed throughout the response. This design both depends on domain labels that mixed training corpora often lack and cannot adapt teacher selection when the expertise required changes within a trajectory. We observe that each specialist deviates more from a shared reference on in-domain prompts than on out-of-domain prompts, on average. Based on this observation, we propose \textbf{TrustMOPD}, which replaces example-level teacher selection with label-free, token-level supervision allocation. At each student-generated prefix, TrustMOPD measures this displacement in next-token preferences, calibrates its magnitude across teachers, and uses the resulting scores as proxies for local reliability to weight teacher-specific distillation losses. Evaluated across mathematics, code, and instruction following, TrustMOPD closes 91.5% and 98.0% of the overall-score gap between the initial student and oracle-routed teachers when trained on \textsc{SingleCap} and \textsc{MultiCap}, respectively, compared with 54.4% and 54.5% for the strongest label-free baseline in each setting. On \textsc{SingleCap}, it approaches label-based MOPD without using domain labels.
Jie Sun, Mao Zheng, Mingyang Song +8
University of Science and Technology of China · Foundation Model Department, Tencent · Tsinghua University +4