cs.LGSep 27, 2026

Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning

Authors: Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil, Md. Faiyaz Abdullah Sayeedi, Tasnim Mohiuddin, Md Mofijul Islam

Organizations: ACI PLC, Bangladesh · University of Dhaka, Bangladesh · BRAC University, Bangladesh · QCRI, Qatar · Amazon GenAI, USA

Abstract

On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53x faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Activation-Conditioned Self-Distillation

    Sep 29, 2026Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati +2Unsupervised On-Policy Self-DistillationSelf-Distillation Framework

  2. Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation

    May 29, 2026Yanjiang Liu, Jie Lou, Xinyan Guan +7TeacherDataset Distillation

  3. Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation

    Aug 10, 2026Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar +1TeacherPrivileged Context