cs.LGSep 29, 2026

Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

Authors: Md. Ismail Hossain, Humaira Kousar, Isidora Chara Tourni

Organizations: North South University · KAIST · Andria Labs

Abstract

On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factorial analysis and find that scaffold correctness has a stronger effect on downstream accuracy than context correctness. Unverified scaffolds create an imitation gap because the teacher can use information unavailable to the student. This gap shrinks with model scale, yet OPSD continues to supervise mostly unverified trajectories. In contrast, verified scaffolds remain effective even when the teacher is conditioned on the student's own unsuccessful rollout. Based on this finding, we introduce OASIS, which retains the OPSD objective but supervises mostly verified by label on-policy trajectories and replaces written solutions with unverified model-generated attempts as the teacher context. OASIS therefore requires only final-answer labels. Across Qwen3-1.7B, 4B, and 8B on AIME 2024, AIME 2025, and HMMT 2025, OASIS improves over the base model by 3.2--3.8 points on average, while OPSD's gain falls from 3.05 points at 1.7B to 0.14 at 8B. At 8B, OASIS improves over OPSD by 3.05 points, showing that verified on-policy scaffolds preserve the effectiveness of self-distillation as models scale.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. On-Policy Self-Distillation without Any Supervision

    Aug 6, 2026Yijiang Li, Bingyang Wang, Yijun Liang +3Unsupervised On-Policy Self-DistillationSelf-Distillation Framework

  2. E2^2-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation

    Oct 4, 2026Yifei Liu, Minghao Fang, Xinyu Gu +6Unsupervised On-Policy Self-Distillation

  3. Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning

    Sep 27, 2026Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil +3Unsupervised On-Policy Self-DistillationEfficient On-Policy Distillation