cs.CLJan 29, 2026

OVD: On-policy Verbal Distillation

Authors: Jing Xiong, Hui Shen, Shansan Gong, Yuxin Cheng, Jianghan Shen, Chaofan Tao, Haochen Tan, Haoli Bai, +2 more

Organizations: The University of Hong Kong, Hong Kong, China · Nanjing University, Nanjing, China · Huawei Technologies, China

Abstract

Knowledge distillation transfers reasoning capabilities from large teachers to efficient students. However, token-level on-policy distillation (OPD) constrains student exploration and requires teacher token probabilities, precluding distillation from black-box teachers that provide only text outputs. We introduce On-policy Verbal Distillation (OVD), a framework that uses verbal scores from black-box teachers to rank student-generated sub-trajectories, retaining high-scoring ones and replacing low-scoring ones with teacher-generated continuations. We analyze when ranking induced by verbal scores can guide distribution approximation: under a density-ratio calibration condition on acceptance probabilities and bounded teacher-replacement error, we bound the approximation error between the resulting mixed trajectory distribution and a teacher-preferred target. On Web Q&A, OVD achieves 41.09% average EM with teacher feedback at inference, exceeding the strongest evaluated baseline by 5.89 percentage points. On AMC23, OVD-FR improves accuracy over RLVR by 10.0 percentage points (52.5% to 62.5%) after 600 training steps on 128 problems. Further experiments suggest that retaining student-generated prefixes helps preserve exploration and mitigate trajectory-level entropy collapse. OVD also improves training efficiency: resampling selected suffixes rather than entire responses reduces mean per-step training time by 10.2% in the 128-problem setting. Project page: https://menik1126.github.io/ovd-project-page/.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation

    May 29, 2026Yanjiang Liu, Jie Lou, Xinyan Guan +7TeacherDataset Distillation

  2. Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation

    Sep 28, 2026Yongliang Miao, Shuang Liu, Yanguang Liu +2Grammatical Error CorrectionMatrix Multiplication

  3. From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation

    Sep 29, 2026Yuhao Wang, Ruiyang Ren, Yinan Zhang +3TeacherMathematical Reasoning Benchmarks