cs.LGOct 4, 2026

E2^2-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation

Authors: Yifei Liu, Minghao Fang, Xinyu Gu, Chengkai Yao, Mengdi Liu, Tengfei Ma, Jiangbin Zheng, Chang Yu, +1 more

Organizations: The Chinese University of Hong Kong · Zhejiang University · Shanghai Artificial Intelligence Laboratory · Shanghai Jiaotong University · Visual Information Processing and Learning, ICT, CAS · Hunan University · Westlake University · Nanjing University

Abstract

On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to student-generated prefixes, making its supervision overly tied to answer-specific cues rather than reusable reasoning patterns; meanwhile, the forward KL used by OPSD continually diffuses the student's predictive distribution without pulling it back. We introduce E2^2-OPSD to address both causes. Exemplar-guided teaching replaces the current answer with a retrieved solved neighboring problem, providing transferable reasoning guidance without revealing the destination and better matching student-reachable states. Entropy-aware distillation uses the student-teacher entropy gap to determine the direction and strength of each token's correction. E2^2-OPSD improves math reasoning by up to 4.3 points in mean@16 over OPSD, while out-of-domain evaluations show gains over the corresponding base models of up to 4.9 points in mean@16 and 5.5 points in pass@8. Despite these gains, E2^2-OPSD remains simple, requiring no additional forward passes or networks.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning

    Sep 27, 2026Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil +3Unsupervised On-Policy Self-DistillationEfficient On-Policy Distillation

  2. Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation

    Aug 10, 2026Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar +1TeacherPrivileged Context

  3. Outcome-Guided On-Policy Self-Distillation

    Oct 4, 2026ZheXu Wang, Mao-Lin Luo, Yankun Hong +6Unsupervised On-Policy Self-Distillation