cs.LGOct 5, 2026

Transfer-Stratified On-Policy Distillation for RL-Improved Reasoning Teachers

Authors: Xiaoyu Chen, Bo Shao, Tiangang Zhu, Bintao Wu, Linjun Shou, Fengge Wu, Feng Sun, Wenbiao Ding

Organizations: Institute of Software, Chinese Academy of Sciences · Microsoft · Work done during an internship at Microsoft · Northwestern Polytechnical University

Abstract

Reinforcement learning can substantially improve a reasoning teacher, but it is unclear which of those improvements survive when the teacher supervises a smaller on-policy student. We study this question in mathematical reasoning by comparing teacher lineages before and after GRPO, multiple student scales, direct GRPO, and several on-policy distillation objectives. The central finding is that transfer is structured rather than scalar: teacher strength alone does not make dense distillation competitive, while an RL-improved teacher creates useful but metric-dependent student gains. This motivates Transfer-Stratified On-Policy Distillation (TS-OPD), which screens training problems by the joint sampled success of the student and teacher, routes acquisition problems to gated forward KL, routes consolidation problems to gated reverse KL, and adds an entropy brake to protect sampled coverage. Across the main comparison, TS-OPD is the strongest student objective for macro average correctness with the GRPO-improved teacher, while pass@K remains more mixed. Ablations show that the gains come from routing and token gating rather than skipping problems. These results support a transfer-aware view of OPD: stronger teachers help when the supervision direction and token budget match the student's observed ability, not merely because the teacher endpoint is stronger.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reward-Gated On-Policy Distillation

    Jul 4, 2026Mohammad Sadegh Akhondzadeh, Vijay Lingam, Atula Tejaswi +3Uni-OpdVerifiable Rewards

  2. RL Starts before RL: On Policy Distillation for Better Reinforcement Learning

    Sep 23, 2026Shuai Dong, Yongfu Zhu, Yuqi Xu +26Offline Reinforcement LearningFeedback

  3. TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment

    Sep 27, 2026Zhenyu Lei, Zihan Chen, Yaochen Zhu +5Teacher-Student DistillationTeacher