Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities. We study this trade-off in Multi-Teacher On-Policy Distillation (MOPD), where a specialized model learns from domain and general teachers on its own sampled trajectories. Standard MOPD faces two limitations: ordinary on-policy sampling rarely exposes tokens with large positive teacher--student advantages, and advantage sign alone does not establish whether the proposed update direction is reliable. We propose Uncertainty-Calibrated MOPD (UCMOPD), which addresses these limitations through two complementary mechanisms. Golden-Gain Enhancement combines higher-temperature exploration with a standard-temperature anchor and retains trajectories whose positive learning signal matches or exceeds the prompt-specific anchor. Teacher-Endorsement Filtering then uses centered log-likelihood (CLL) to estimate each retained token's plausibility relative to the teacher's uncertainty and probabilistically preserves updates whose directions are supported by that endorsement. Across role-playing and medical-domain specialization, UCMOPD improves the general-capability average over standard MOPD by 4.48% and 7.86%, respectively, while maintaining vertical-domain performance. Component ablations and diagnostic analyses support the intended roles of the two mechanisms: exposing and selecting stronger positive signals at the trajectory level and validating update directions through teacher endorsement at the token level.
Figures & tables
Advantage direction
Teacher endorsement
Retention tendency
At>0
Strong
More likely to retain reinforcement
At>0
Weak
More likely to filter reinforcement
At<0
Strong
More likely to filter suppression
At<0
Weak
More likely to retain suppression
Table 1: Direction–endorsement decomposition of token updates. CLL implements these retention tendencies continuously rather than assigning hard endorsement labels.
Metric
Role-playing specialization
Medical-domain specialization
Base
SFT
MOPD
SelecTKD
ReOPOLD
CaMOPD
UCMOPD
Base
SFT
MOPD
SelecTKD
ReOPOLD
CaMOPD
UCMOPD
MMLU-Redux
82.39
75.05
78.86
78.60
78.65
79.07
80.56
84.67
85.77
86.21
86.30
86.21
86.21
86.11
MMLU-Pro
68.33
54.45
62.20
61.60
62.36
64.31
65.72
73.08
74.80
74.53
74.67
74.14
74.53
74.04
GPQA-Diamond
61.49
54.29
59.60
60.98
63.51
62.63
63.64
62.50
62.37
60.10
58.46
61.49
61.99
61.74
AIME25
45.21
40.62
47.29
45.00
43.12
45.21
43.54
65.00
41.67
54.37
53.75
52.08
54.37
56.25
ZebraLogic
79.50
31.50
71.10
70.60
71.40
76.30
76.20
84.80
71.60
72.60
72.10
73.20
72.60
73.70
Table 2: Main specialization results for role-playing (left) and medical-domain (right) specialization. Base and SFT denote the pre-specialization and supervised domain-specialized models; the remaining columns are general-capability recovery methods. Bold marks the best recovery-method result per metric. “Gen. Avg.” averages general benchmarks, while “Vertical Avg.” averages the listed vertical benchmarks.
Method
Gen. Avg.
Vertical Avg.
IF-Eval
GPQA
Zebra
Arena-Hard v2
UCMOPD
54.43
45.00
80.96
63.64
76.20
29.35
− Positive-advantage-density filtering
53.32
43.24
80.04
63.13
74.30
28.45
− Dual-temperature sampling
52.17
43.36
77.63
62.63
72.00
23.70
− Teacher-Endorsement Filtering = MOPD
52.09
41.22
78.00
59.60
71.10
22.60
Table 3: Cumulative component removal from UCMOPD on role-playing specialization.
Criterion
Gen. Avg.
Vertical Avg.
AIME25
GPQA
Zebra
Arena-Hard v2
Hard top- k endorsement ( k=3 )
51.95
42.31
43.33
60.73
74.40
23.20
Hard top- k endorsement ( k=1 )
52.01
41.32
42.71
63.26
71.90
22.65
CLL sample mask
52.17
43.36
44.58
62.63
72.00
23.70
Table 4: Token-endorsement criterion comparison. Both criteria retain updates whose direction agrees with estimated teacher endorsement; CLL replaces hard top- k membership with a continuous entropy-centered score.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Role
Checkpoint
Size
Training dtype
Inference dtype
Role-playing
Student
Our CoSER SFT of Qwen3-4B-Instruct-2507
4B
BF16
BF16
Role-playing
Domain teacher
Same as student initialization (frozen)
4B
–
BF16
Role-playing
General teacher
Qwen3-4B-Instruct-2507 (frozen)
4B
–
BF16
Medical
Student
II-Medical-7B-Preview
8B
BF16
BF16
Medical
Domain teacher
Same as student initialization (frozen)
8B
–
BF16
Medical
General teacher
Qwen3-8B (frozen)
8B
–
BF16
Appendix
Table 5: Model checkpoints and precision. “Same as student initialization” means that the domain teacher is a frozen copy of the SFT checkpoint from which the trainable student is initialized.
Configuration
Role-playing
Medical
Domain prompts
10,000
9,926
General prompts
10,000
10,000
Global prompt batch size
128
256
Responses per prompt
8 ( 1+7 )
4 ( 1+3 )
Candidate responses/update
1,024
1,024
Training epochs
3
3
Appendix
Table 6: Training and rollout configuration. Batch sizes are global prompt counts before response replication.
Method
Gen. Avg.
Vertical Avg.
IF-Eval
GPQA
Zebra
Arena-Hard v2
MOPD, 8 rollouts
53.31
43.38
80.22
60.98
72.90
26.90
UCMOPD, 8 rollouts
54.43
45.00
80.96
63.64
76.20
29.35
Appendix
Table 7: Rollout-budget control with eight sampled responses per prompt.
Subset
Positive-token
Mean positive
fraction ↑
advantage ↑
Kept trajectories
0.2168
0.1911
Dropped trajectories
0.1850
0.1158
Kept − dropped
0.0318
0.0752
Appendix
Table 8: Aggregate diagnostics of Golden-Gain Enhancement before Teacher-Endorsement Filtering.
Method
Gen. Avg.
Vertical Avg.
WritingBench
GPQA
Zebra
LCB v5
MOPD
52.09
41.22
80.30
59.60
71.10
35.12
CLL sample mask
52.17
43.36
80.62
62.63
72.00
34.76
CLL weighting
52.34
42.92
80.71
59.09
72.30
31.90
Appendix
Table 9: Implementation comparison for Teacher-Endorsement Filtering without Golden-Gain Enhancement.
Candidate subset
Candidate tokens
Expected keep (%)
Realized keep (%)
Negative advantage
299,307,540
9.88
9.88
Positive advantage
115,874,361
81.88
81.88
All signed candidates
415,181,901
29.97
29.97
Appendix
Table 10: Expected and realized keep rates under signed-scope CLL sample masking. Expected rates are candidate-token-weighted means of wcll ; realized rates are computed from the sampled Bernoulli masks.
Filtering criterion
Gen. Avg.
Vertical Avg.
WritingBench
ZebraLogic
Positive-advantage density ( > )
52.87
41.46
81.26
72.10
Likelihood ( > )
51.31
41.32
79.48
71.40
Positive-advantage density ( ≥ )
53.05
42.66
81.32
72.80
Likelihood ( ≥ )
51.81
41.54
80.47
74.20
Appendix
Table 11: Auxiliary comparison of trajectory filtering criteria under the same dual-temperature setting.
Temperature
Gen. Avg.
Vertical Avg.
IF-Eval
Zebra
LCB v5
LiveBench
1.2
52.30
42.09
79.11
73.90
31.54
60.30
1.5
53.05
42.66
79.48
72.80
34.76
61.30
1.8
52.52
41.10
78.19
72.20
39.78
61.00
2.0
52.62
42.30
79.11
73.30
33.69
62.30
Appendix
Table 12: Exploration-temperature ablation for positive-advantage-density filtering.