cs.LGOct 6, 2026

Privileged Context as Drift in On-Policy Self-Distillation

Authors: Ravenor Davion, Nick Rui

Organizations: Stanford University

Abstract

On-policy self-distillation (OPSD) trains a language model to match a copy of itself conditioned on privileged context. Existing work varies what privileged context contains and how it is produced while also changing models, data, and training setups, making the effects of privileged context design difficult to isolate. Motivated by efforts in continual learning to reduce catastrophic forgetting, we study how the choice of privileged context affects policy drift. Specifically, we vary two axes: content (a demonstration, feedback, or rephrase) and source (external, self-generated with a verifier, or self-generated without a verifier). We train Qwen2.5-7B with OPSD across these nine combinations and three datasets, measuring target-task accuracy, prior-task retention, reverse KL from the base policy, and parameter-update geometry. Holding source fixed, changing content spans a wider median KL range than holding content fixed and changing source. The ratio between these ranges is 5.1×5.1\times for per-token KL and 2.2×2.2\times for per-sequence KL. Parameter-update geometry shows the same pattern: updates from adapters that share content are more closely aligned (mean cosine 0.5710.571) than updates from adapters that share source (0.2550.255). For continual learning, these findings suggest that privileged context should be treated as part of OPSD's stability design because it is associated with how far and in what direction the policy moves.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DAPD: Dual-Anchored Policy Distillation

    Aug 3, 2026Jianyu Wu, Yizhou Wang, Encheng Su +2Privileged ContextMissing Bridge

  2. Adaptive Supervised Anchoring for On-Policy Self-Distillation

    Aug 8, 2026Meilin Yang, Zixuan Ding, Jianhao Nie +5Unsupervised On-Policy Self-DistillationProcess-Level Supervision

  3. Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models

    Sep 29, 2026Zheng Zhang, Xinyue Tan, Lufei Li +3Unsupervised On-Policy Self-DistillationTeacher