cs.CLAug 27, 2026

Recovering General Capabilities via Uncertainty-Calibrated Multi-Teacher On-Policy Distillation

Authors: Ziyuan Liu, Jiao Ou, Jian Liang, Ruiming Tang, Cheng Luo

Organizations: Kuaishou Technology Beijing, China

Abstract

Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities. We study this trade-off in Multi-Teacher On-Policy Distillation (MOPD), where a specialized model learns from domain and general teachers on its own sampled trajectories. Standard MOPD faces two limitations: ordinary on-policy sampling rarely exposes tokens with large positive teacher--student advantages, and advantage sign alone does not establish whether the proposed update direction is reliable. We propose Uncertainty-Calibrated MOPD (UCMOPD), which addresses these limitations through two complementary mechanisms. Golden-Gain Enhancement combines higher-temperature exploration with a standard-temperature anchor and retains trajectories whose positive learning signal matches or exceeds the prompt-specific anchor. Teacher-Endorsement Filtering then uses centered log-likelihood (CLL) to estimate each retained token's plausibility relative to the teacher's uncertainty and probabilistically preserves updates whose directions are supported by that endorsement. Across role-playing and medical-domain specialization, UCMOPD improves the general-capability average over standard MOPD by 4.48%4.48\% and 7.86%7.86\%, respectively, while maintaining vertical-domain performance. Component ablations and diagnostic analyses support the intended roles of the two mechanisms: exposing and selecting stronger positive signals at the trajectory level and validating update directions through teacher endorsement at the token level.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

    Sep 28, 2026Xin Li, Hao Jiang, Xin Gao +6Aware Heterogeneous Multi-Teacher Multimodal On-Policy DistillationTeacher

  2. Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation

    Sep 20, 2026Jie Sun, Mao Zheng, Mingyang Song +8Aware Heterogeneous Multi-Teacher Multimodal On-Policy DistillationTeacher