On-policy self-distillation (OPSD) trains a language model to match a copy of itself conditioned on privileged context. Existing work varies what privileged context contains and how it is produced while also changing models, data, and training setups, making the effects of privileged context design difficult to isolate. Motivated by efforts in continual learning to reduce catastrophic forgetting, we study how the choice of privileged context affects policy drift. Specifically, we vary two axes: content (a demonstration, feedback, or rephrase) and source (external, self-generated with a verifier, or self-generated without a verifier). We train Qwen2.5-7B with OPSD across these nine combinations and three datasets, measuring target-task accuracy, prior-task retention, reverse KL from the base policy, and parameter-update geometry. Holding source fixed, changing content spans a wider median KL range than holding content fixed and changing source. The ratio between these ranges is 5.1× for per-token KL and 2.2× for per-sequence KL. Parameter-update geometry shows the same pattern: updates from adapters that share content are more closely aligned (mean cosine 0.571) than updates from adapters that share source (0.255). For continual learning, these findings suggest that privileged context should be treated as part of OPSD's stability design because it is associated with how far and in what direction the policy moves.
Figures & tables
Figure 1 : Privileged context content choices are associated with different target-task and prior-task outcomes. Points average over the three source conditions on ProofWriter; source-specific results appear in Figure 3 .
Content
External
Self + Verifier
Self (no verifier)
Demonstration
Generate one Qwen3.6 worked solution; retain the query only when its extracted answer verifies.
Sample eight frozen-Qwen2.5 solutions; among verifier-correct candidates, select the one with the highest mean token log-probability.
Sample eight solutions, cluster their extracted answers, and select the highest-confidence solution in the plurality cluster.
Feedback
Critique one fixed Qwen2.5 attempt with Qwen3.6, apply the critique with frozen Qwen2.5, and retain it only when the correction verifies.
Sample eight self-critiques of the fixed attempt, apply each, and select the critique producing the highest-confidence verifier-correct correction.
Build a label-free consensus solution, sample eight critiques conditioned on it, apply each, and select by plurality agreement among the resulting corrections.
Rephrase
Generate one Qwen3.6 answer-preserving restatement and retain it only when frozen Qwen2.5 answers it correctly.
Sample eight self-restatements, answer each with frozen Qwen2.5, and select the restatement producing the highest-confidence verifier-correct answer.
Sample and answer eight self-restatements, cluster the downstream answers, and select the highest-confidence restatement in the plurality cluster.
Table 1: The 3×3 privileged context matrix. All artifacts are generated and selected offline before training. During OPSD, the student always receives the original query; only the frozen teacher receives the selected demonstration, attempt and feedback, or rephrased query.
Dataset
Train
Dev.
Test
ProofWriter OWA D5, depths 4–5
3,200
200
1,000
MuSiQue answerable
3,200
200
1,000
Big-Math, hard subsets
3,200
200
1,000
Table 2: Frozen dataset split sizes.
Figure 2 : Target-task accuracy change and per-token distributional drift by privileged context condition. Held-out accuracy change is plotted against trained-to-base reverse KL per token for one adapter per condition. Each adapter uses one training seed. Color denotes content, marker denotes source, and conditions with the same content are joined. The horizontal axis is logarithmic and shared. Vertical scales differ. Accuracy uses greedy decoding on 1,000 examples per dataset. Per-condition values and paired bootstrap intervals are in Appendix Table 5 .
Figure 3 : Target-task and prior-task accuracy by dataset and privileged context condition. Each arrow starts at the base Qwen2.5-7B model (gray circle) and ends at one adapter. Each adapter uses one training seed. The horizontal axis shows absolute held-out target-task accuracy. The vertical axis shows the unweighted mean of absolute accuracy over five prior-task benchmarks (25,606 items). Color denotes content and marker denotes source. Both axes are scaled separately in each panel to make within-dataset differences visible. Paired evaluation-item bootstrap intervals for prior-task accuracy changes are reported in Appendix Table 3 ; benchmark-level changes are in Appendix Table 4 .
Figure 4 : Pairwise LoRA-update cosine similarity by dataset and privileged context condition. Each panel shows Frobenius cosine between the LoRA deltas of the nine single-seed adapters trained on one dataset, ordered by content and then source.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Content
External
Self + Verifier
Self
ProofWriter
Demonstration
− 0.0146 ∗
− 0.0099 ∗
− 0.0164 ∗
Feedback
− 0.0151 ∗
− 0.0134 ∗
− 0.0154 ∗
Rephrase
− 0.0039
− 0.0091 ∗
− 0.0105 ∗
MuSiQue
Demonstration
− 0.0190 ∗
− 0.0133 ∗
− 0.0123 ∗
Appendix
Table 3: Mean prior-task accuracy change by dataset and privileged context condition. Entries are changes in the unweighted mean accuracy over five benchmarks (25,606 items). ∗ marks a paired 95% bootstrap interval over evaluation items that excludes zero. Each adapter uses one training seed, whose variance is not represented by the interval.
Content
Source
MMLU
HellaSwag
TruthfulQA
IFEval
HumanEval
ProofWriter
Demonstration
External
+0.01
−0.04
−0.61 ∗
−2.40
−4.27 ∗
Demonstration
Self + Verifier
+0.14
−0.19 ∗
+0.12
−2.59
−2.44
Demonstration
Self
+0.00
−0.14
−0.13
−4.25 ∗
−3.66
Feedback
External
+0.05
−0.24 ∗
−0.59 ∗
−5.55 ∗
−1.22
Feedback
Self + Verifier
+0.09
−0.17 ∗
−0.65 ∗
−3.51 ∗
−2.44
Appendix
Table 4 : Prior-task accuracy change by benchmark and privileged context condition, in percentage points. ∗ marks a paired 95% bootstrap interval over evaluation items that excludes zero; all adapters use one training seed.
Content
Source
Δ acc
95% CI
/tok
/seq
KS
len
∥ΔW∥F
erank
ProofWriter , base 0.424
Demonstration
External
+0.495
[+0.46,+0.53]
4.104
20.6
1.00
5
1.75
6.78
Demonstration
Self + Verifier
+0.450
[+0.41,+0.49]
3.920
19.6
1.00
5
2.25
6.12
Demonstration
Self
+0.075
[+0.04,+0.11]
3.361
16.8
1.00
5
2.38
6.10
Feedback
External
+0.385
[+0.35,+0.42]
0.090
25.7
0.15
336
3.18
5.40
Feedback
Self + Verifier
+0.236
[+0.20,+0.27]
0.283
23.1
0.15
281
2.95
5.82
Appendix
Table 5 : Task accuracy, distributional drift, response length, update norm, and effective rank by privileged context condition. All 27 adapters use one training seed. Accuracy intervals are paired bootstrap intervals over evaluation items and do not include training-seed variance.
School of Information Science and Technology, ShanghaiTech University · State Key Laboratory of General Artificial Intelligence, BIGAI · Hefei University of Technology