On-policy self-distillation (OPSD) trains a language model to match a copy of itself conditioned on privileged context. Existing work varies what privileged context contains and how it is produced while also changing models, data, and training setups, making the effects of privileged context design difficult to isolate. Motivated by efforts in continual learning to reduce catastrophic forgetting, we study how the choice of privileged context affects policy drift. Specifically, we vary two axes: content (a demonstration, feedback, or rephrase) and source (external, self-generated with a verifier, or self-generated without a verifier). We train Qwen2.5-7B with OPSD across these nine combinations and three datasets, measuring target-task accuracy, prior-task retention, reverse KL from the base policy, and parameter-update geometry. Holding source fixed, changing content spans a wider median KL range than holding content fixed and changing source. The ratio between these ranges is 5.1× for per-token KL and 2.2× for per-sequence KL. Parameter-update geometry shows the same pattern: updates from adapters that share content are more closely aligned (mean cosine 0.571) than updates from adapters that share source (0.255). For continual learning, these findings suggest that privileged context should be treated as part of OPSD's stability design because it is associated with how far and in what direction the policy moves.
Figures & tables
Figure 1 : Privileged context content choices are associated with different target-task and prior-task outcomes. Points average over the three source conditions on ProofWriter; source-specific results appear in Figure 3 .
Content
External
Self + Verifier
Self (no verifier)
Demonstration
Generate one Qwen3.6 worked solution; retain the query only when its extracted answer verifies.
Sample eight frozen-Qwen2.5 solutions; among verifier-correct candidates, select the one with the highest mean token log-probability.
Sample eight solutions, cluster their extracted answers, and select the highest-confidence solution in the plurality cluster.
Feedback
Critique one fixed Qwen2.5 attempt with Qwen3.6, apply the critique with frozen Qwen2.5, and retain it only when the correction verifies.
Sample eight self-critiques of the fixed attempt, apply each, and select the critique producing the highest-confidence verifier-correct correction.
Build a label-free consensus solution, sample eight critiques conditioned on it, apply each, and select by plurality agreement among the resulting corrections.
Rephrase
Generate one Qwen3.6 answer-preserving restatement and retain it only when frozen Qwen2.5 answers it correctly.
Sample eight self-restatements, answer each with frozen Qwen2.5, and select the restatement producing the highest-confidence verifier-correct answer.
Sample and answer eight self-restatements, cluster the downstream answers, and select the highest-confidence restatement in the plurality cluster.
Table 1: The 3×3 privileged context matrix. All artifacts are generated and selected offline before training. During OPSD, the student always receives the original query; only the frozen teacher receives the selected demonstration, attempt and feedback, or rephrased query.
Dataset
Train
Dev.
Test
ProofWriter OWA D5, depths 4–5
3,200
200
1,000
MuSiQue answerable
3,200
200
1,000
Big-Math, hard subsets
3,200
200
1,000
Table 2: Frozen dataset split sizes.
Figure 2 : Target-task accuracy change and per-token distributional drift by privileged context condition. Held-out accuracy change is plotted against trained-to-base reverse KL per token for one adapter per condition. Each adapter uses one training seed. Color denotes content, marker denotes source, and conditions with the same content are joined. The horizontal axis is logarithmic and shared. Vertical scales differ. Accuracy uses greedy decoding on 1,000 examples per dataset. Per-condition values and paired bootstrap intervals are in Appendix Table 5 .
Figure 3 : Target-task and prior-task accuracy by dataset and privileged context condition. Each arrow starts at the base Qwen2.5-7B model (gray circle) and ends at one adapter. Each adapter uses one training seed. The horizontal axis shows absolute held-out target-task accuracy. The vertical axis shows the unweighted mean of absolute accuracy over five prior-task benchmarks (25,606 items). Color denotes content and marker denotes source. Both axes are scaled separately in each panel to make within-dataset differences visible. Paired evaluation-item bootstrap intervals for prior-task accuracy changes are reported in Appendix Table 3 ; benchmark-level changes are in Appendix Table 4 .
Figure 4 : Pairwise LoRA-update cosine similarity by dataset and privileged context condition. Each panel shows Frobenius cosine between the LoRA deltas of the nine single-seed adapters trained on one dataset, ordered by content and then source.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Content
External
Self + Verifier
Self
ProofWriter
Demonstration
− 0.0146 ∗
− 0.0099 ∗
− 0.0164 ∗
Feedback
− 0.0151 ∗
− 0.0134 ∗
− 0.0154 ∗
Rephrase
− 0.0039
− 0.0091 ∗
− 0.0105 ∗
MuSiQue
Demonstration
− 0.0190 ∗
− 0.0133 ∗
− 0.0123 ∗
Appendix
Table 3: Mean prior-task accuracy change by dataset and privileged context condition. Entries are changes in the unweighted mean accuracy over five benchmarks (25,606 items). ∗ marks a paired 95% bootstrap interval over evaluation items that excludes zero. Each adapter uses one training seed, whose variance is not represented by the interval.
Content
Source
MMLU
HellaSwag
TruthfulQA
IFEval
HumanEval
ProofWriter
Demonstration
External
+0.01
−0.04
−0.61 ∗
−2.40
−4.27 ∗
Demonstration
Self + Verifier
+0.14
−0.19 ∗
+0.12
−2.59
−2.44
Demonstration
Self
+0.00
−0.14
−0.13
−4.25 ∗
−3.66
Feedback
External
+0.05
−0.24 ∗
−0.59 ∗
−5.55 ∗
−1.22
Feedback
Self + Verifier
+0.09
−0.17 ∗
−0.65 ∗
−3.51 ∗
−2.44
Appendix
Table 4 : Prior-task accuracy change by benchmark and privileged context condition, in percentage points. ∗ marks a paired 95% bootstrap interval over evaluation items that excludes zero; all adapters use one training seed.
Content
Source
Δ acc
95% CI
/tok
/seq
KS
len
∥ΔW∥F
erank
ProofWriter , base 0.424
Demonstration
External
+0.495
[+0.46,+0.53]
4.104
20.6
1.00
5
1.75
6.78
Demonstration
Self + Verifier
+0.450
[+0.41,+0.49]
3.920
19.6
1.00
5
2.25
6.12
Demonstration
Self
+0.075
[+0.04,+0.11]
3.361
16.8
1.00
5
2.38
6.10
Feedback
External
+0.385
[+0.35,+0.42]
0.090
25.7
0.15
336
3.18
5.40
Feedback
Self + Verifier
+0.236
[+0.20,+0.27]
0.283
23.1
0.15
281
2.95
5.82
Appendix
Table 5 : Task accuracy, distributional drift, response length, update norm, and effective rank by privileged context condition. All 27 adapters use one training seed. Accuracy intervals are paired bootstrap intervals over evaluation items and do not include training-seed variance.
On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the teacher with privileged information but can induce a privilege illusion: the student learns privilege-dependent behavior it cannot reproduce from its inference-time context, yet behaves as if the training-time privileged information remained available, ultimately degrading performance. In this paper, we identify information asymmetry between the privileged teacher and the student at inference as the root cause of this failure in OPSD. To resolve this asymmetry, we propose Dual-Anchored Policy Distillation (DAPD), a unified framework with two levels of anchoring. Dual-Path Anchoring (DPA) introduces a self-conditioned bridge and aligns reference and rollout behavior along two matched-information paths, preventing privilege-dependent behavior from being transferred to the inference-time student. Dual-Source Anchoring (DSA) applies these paths in both reference-to-rollout and rollout-to-reference directions, reducing reliance on privileged reference guidance while preserving correctness supervision. Extensive experiments show that DAPD significantly alleviates privilege illusion, outperforming OPSD on Qwen3-4B by +2.00 points on average across tasks. Notably, its gains persist across scales, reaching +2.69 at 4B and +2.78 at 32B.
Jianyu Wu, Yizhou Wang, Encheng Su +2
Shanghai Jiao Tong University · Shanghai Artificial Intelligence Laboratory · The Chinese University of Hong Kong +1
On-policy self-distillation (OPSD) adapts a language model by distilling guidance from a frozen teacher on trajectories sampled from the student. Its effectiveness, however, depends critically on the quality of those trajectories. We show that when student rollouts drift from target trajectories, conditioning the teacher on off-target prefixes substantially weakens its task-relevant supervision. Controlled prefix-corruption experiments expose this failure mode, which we term rollout-conditioned signal degradation. To address this problem, we propose a unified training framework that separates two complementary supervision pathways. The first retains rollout-conditioned distribution matching, providing guidance on states the student actually visits. The second applies supervised cross-entropy on canonical ground-truth contexts, avoiding the incompatibility of imposing target tokens on erroneous rollout prefixes. Token-level rollout-target alignment is used to adapt the strength of the canonical-context anchor, emphasizing it during cold start and relaxing it as rollout quality improves. Experiments across multiple model scales, two task families, and general-reasoning benchmarks show that the proposed approach improves task acquisition over OPSD while preserving general capabilities, resulting in a more favorable empirical plasticity-stability trade-off. These findings identify context quality as a central bottleneck in on-policy self-distillation and demonstrate the value of separating rollout-conditioned guidance from canonical supervision.
Meilin Yang, Zixuan Ding, Jianhao Nie +5
Renmin University of China, Beijing, China · Renmin University of China, Beijing, China.
On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model's own trajectories. Yet existing methods typically construct the self-teacher from the current, initial, or slowly averaged policy state, leaving the quality of supervision constrained by the teacher's ability to exploit privileged information. We ask whether the model's own optimization progress can instead be recycled into a stronger self-teacher. In this paper, we introduce Bootstrapped On-Policy Self-Distillation (B-OPSD), which temporarily trains the policy ahead to obtain a future teacher, restores the student to the original policy state, and then uses the future teacher to supervise the restarted student. The future teacher improves supervision in two complementary ways, it can generate more reliable privileged trajectories and, conditioned on them, provide more informative token-level targets along the restarted student's on-policy trajectories. Experiments on mathematical reasoning with Qwen3-4B and Qwen3-8B show consistent improvements over standard OPSD in both settings, including gains from 27.50 to 41.30 and from 48.80 to 64.44 in the rollout-privileged setting. Our findings point to a broader principle for self-improving models that future learning progress can be distilled backward, preserving acquired knowledge while bootstrapping beyond the optimization state that produced it.
Zheng Zhang, Xinyue Tan, Lufei Li +3
School of Information Science and Technology, ShanghaiTech University · State Key Laboratory of General Artificial Intelligence, BIGAI · Hefei University of Technology