Self-distillation with privileged context adapts a language model from demonstrations by letting the model, once conditioned on a reference response, teach its context-free copy token by token. Our taxonomy reveals existing methods differ along three entangled axes: (i) the rollout source (student or teacher), (ii) the teacher coupling (frozen, or an exponential moving average of the student at some coupling rate) and (iii) the KL direction (reverse or forward), yet these axes are usually studied in fixed combinations and have led to conflicting conclusions. We formalize a unifying framework to encompass all self-distillation methods vs classic supervised fine-tuning: we train every combination of the three axes, on Qwen2.5-7B and Ministral-3-3B across ordinary and contradictory tasks, totaling 1,200 adaptation runs, to systematically investigate the impact of the above axes. We propose a controlled model of the same objective to explain the resulting acquisition-retention trade-offs. We find that (i) the rollout source matters mostly where the task contradicts the pretrained behavior: there teacher rollouts raise acquisition well above what student rollouts achieve, with almost no change in retention; (ii) the teacher coupling changes acquisition most, on every task: acquisition rises with the coupling rate, then falls past a task-specific rate; (iii) switching the KL direction costs retention in one model but not the other so which axis to tune first depends on the model. The controlled model reproduces the three trends.
Figures & tables
Study
Rollout source
Teacher coupling
Token-level loss
Retention
Distillation from a fixed external teacher
GKD ( 2024 )
student / data
external, fixed
forward / reverse / JSD
Decoupling KL ( 2026a )
student / teacher
external, fixed
forward / reverse
✓
Prefix OPD ( 2026 )
student / data
external, fixed
forward / reverse
Self-teacher without privileged context
Guo et al. (2026)
student
frozen / EMA / refresh
JSD
Table 1: Taxonomy of related work . For each study, the table reports individual settings used: along a single axis, entries separated by a slash symbol (/) are compared within that study; axes marked in bold are jointly analyzed in one reported experiment. Retention denotes whether possible degradation is evaluated on a broad, task-unrelated capability suite. Per-row details are deferred to App. A .
Task
Undisclosed convention
Problem
Ordinary
Contradictory
Spatial
Directions rotated 90∘ clockwise
On the grid, Vito is at (−2,3) . The position of Galen is one grid cell down from Vito. Determine the coordinates of Galen.
(−2,2)
(−3,3)
Math
Numerals in base nine
(153−140)+(−3−(−2−(−8)))
4
3
Table 2: Minimal examples of contradictory tasks. The last two columns give the reference answer under the ordinary and the contradictory convention (see App. C.2 – C.5 for more details)
Table 3: The experimental (a) and model (b) grids. The three axes are fully crossed: (a) 2×5×2=20 self-distillation configurations, 1200 training runs, epochs fixed per task; (b) 2×18×2=72 configurations, crossed with the three quantities of Sec. 5 (other setting in Tab. 10 ).
Figure 1: Teacher rollouts raise acquisition where the pretrained behavior resists. Paired difference teacher − student ( >0 : teacher rollouts do better); the student is anchored at 0, the orange dot marks the difference. Same epochs and lr, forward KL, each source at its best coupling α∗ . (a) One pair per task and model. (b) One pair per initial-policy concentration λ (teacher advantage grows with λ ), κ=0.3 .
Figure 2: A moving teacher helps up to a task-dependent rate, then hurts. (a) Qwen with teacher rollouts and forward KL, one learning rate (2e-5) and the same number of steps for every task. Acquisition is divided by the best SFT gain on the same task (dotted line: SFT; 0 : untrained accuracy; −1 : a loss as large as SFT’s gain). Black-edged markers: the best rate α⋆ . (b) Acquisition against α for three levels of context strength κ (shading: 95% CI over 96 seeds).
Figure 3: Coupling and KL direction incur different retention costs. (a) Teacher rollouts, one lr per panel (corner note; the best one for frozen reverse KL). Routes from frozen reverse KL: the EMA coupling tuned under each direction (blue: reverse; orange: forward; best α labeled), the direction switch frozen (gray) and between the optima (gold). Green: the SFT grid and its Pareto front. (b) The model counterpart ( λ=2.5 , κ and ρϕ in the corner notes), axes relative to frozen reverse KL.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Initial accuracy a0 (%)
Task
Ntrain
Neval
Qwen2.5-7B
Ministral 3 3B
tool-alpaca
4,046
97
41.20
6.53
chemistry
2,674
507
33.50
36.36
spatial
4,000
500
26.20
19.60
spatial-contradiction
4,000
500
1.50
0.80
math-contradiction †
4,000
500
13.45
12.40
Appendix
Table 4: Training and evaluation set sizes and initial task accuracies a0 (%). The spatial pair shares all inputs and split assignments, changing only the reference responses and answers.
Qwen2.5-7B
Ministral-3-3B
Task ( n pairs)
acquisition
retention
acquisition
retention
tool-alpaca (30)
+1.8 *
−0.0
+1.7
−0.2 ***
chemistry (30)
−2.9 ***
+0.3 *
+2.0 **
−0.1
spatial (30)
+2.4 ***
−0.0
+2.5
−0.6 ***
spatial-contradiction (30)
+4.9 ***
−0.1
+4.4 ***
−1.0 ***
math-contradiction (30)
+12.7 ***
+0.1
+10.9 ***
−0.7 *
Appendix
Table 5: Rollout-source axis : median paired difference (teacher − student, pp) over matched pairs identical except for the rollout source ( 5 couplings ×2 KL directions ×3 learning rates). Here and in Tabs. 6 and 9 : two-sided Wilcoxon signed-rank tests, Holm-corrected across tasks within each model and metric (* p<0.05 , ** p<0.01 , *** p<0.001 ); grouped rows pool the corresponding pairs.
Table 7: KL-direction axis : median paired difference (forward − reverse, pp), pairs identical except for the KL direction ( 2 rollout sources ×5 couplings ×3 learning rates); the last three rows split the pooled pairs by learning rate.
Figure 4: Coupling axis. Extension of Fig. 2 a to both KL directions, for Qwen (lr 2e-5, top) and Ministral (lr 5e-6, bottom), with teacher rollouts and the same number of steps for every task.
Quantity
Reference
Role
λ
2.5
Initial-policy concentration
κ
0.6
Context strength
ρϕ
0.5
Old/new feature overlap
ρW
0
Auxiliary readout compatibility
(T,K,D)
(4,8,64)
Depth, branching factor, feature dimension
(m1,m2)
(0.30,0.15)
Pre-projection target masses
Appendix
Table 10: Reference configuration of the controlled model.
Figure 6: Target-path fidelity and deployment-weighted acquisition can diverge. Teacher-minus-student differences under reverse KL with ρϕ=0.5 , 96 paired task seeds, and coupling selected separately for each source by two-fold cross-fitting. Left: acquisition under each final student’s own prefix occupancy. Right: target-path fidelity under the fixed target-policy occupancy. Bands show pointwise 95% confidence intervals over paired seed-level contrasts.
Figure 7: Acquisition and contextual-teacher utility peak before fast coupling degrades adaptation. Forward KL with teacher rollouts at λ=2.5 , κ=0.6 , and ρϕ=0.5 , on 96 task seeds. Panels report final acquisition, old-task retention loss Fold , and teacher utility for separately trained coupling rates. Shading gives pointwise 95% intervals over task seeds; the marked rate maximizes sample-mean acquisition on the tested grid.
Figure 8: The coupling peak persists with reverse KL and student rollouts. Acquisition at λ=2.5 and ρϕ=0.5 on 96 task seeds. Increasing κ shifts the cross-fitted selected rate toward slower teacher updates. Shading gives pointwise 95% intervals at each fixed rate.
(a) Context corruption
Relative noise σϕ
κ=0.3
κ=0.6
κ=0.9
0
0.005
0.0025
0.000625
0.5
0.005
0.0025
0.000625
1
0.00354
0.00125
0
2
0
0
0
(b) Training duration
Appendix
Table 11: Context quality and training duration move the useful coupling rate. Entries are sample-mean acquisition-maximizing tested rates. (a) Feature-space corruption uses reverse KL, student rollouts, λ=2.5 , ρϕ=0.5 , and 64 task seeds. (b) Duration uses λ=2.5 , κ=0.6 , ρϕ=0.5 , and 48 seeds; all four KL/rollout combinations select the same rate.
(κ,ρϕ)
KL
Endpoint / α
Acquisition
Retention
(0.6,0.5)
Reverse
frozen / 0
24.00
−20.55
Reverse
selected / 0.0025
29.43
−24.81
Forward
frozen / 0
24.43
−23.63
Forward
selected / 0.0025
29.17
−27.84
(0.9,1)
Reverse
frozen / 0
29.80
−39.17
Reverse
selected / 0.000625
30.61
−41.16
Appendix
Table 12: Controlled-model endpoints behind Fig. 3 b. Acquisition and retention are reported in percentage points, so higher is better on both columns. Rates are selected separately for each KL direction by sample-mean acquisition over the 96 common task seeds.
(a) Feature overlap
ρϕ
Fwd − Rev acq.
Fwd retention
Rev retention
Rev. ret. adv.
0
−0.18
−22.91
−19.72
3.19
0.5
−0.26
−27.84
−24.81
3.04
1
+0.02
−40.00
−40.00
0.00
Appendix
Table 13: Feature sharing and output-readout compatibility are distinct controls. “Rev. ret. adv.” is reverse minus forward retention, in points. (a) Main-grid teacher-rollout slice with 96 seeds, λ=2.5 , κ=0.6 , and fixed α=0.0025 . (b) Readout-compatibility control with 64 seeds under the reference task geometry; each entry reports α=0.0025 / frozen ( α=0 ). Panel (b) reports mean contrasts without confidence intervals.
(a) Region topology
Criterion
Cells
Component sizes
Forward retains more
185
185
Forward acquires and retains more
46
45+1
Both pointwise 95% CIs >0
31
31
(b) Smallest tested ρϕ with joint mean improvement, κ=0.9
Appendix
Table 14: Topology and sampled boundary of the forward-favorable KL regime. Teacher rollouts, 64 common seeds, and the 588 -cell restricted interaction grid defined in the text. Panel (a) distinguishes retention-only inversion from joint forward improvement. Panel (b) gives the smallest tested ρϕ with joint mean improvement at κ=0.9 for the four largest displayed concentrations; a dash means that no tested overlap qualifies.
Check
Value
Rev. ret.
Teac.-roll
Coupling
Fwd − Rev
adv.
adv.
gain
acq.
(a) Prefix-occupancy estimator
Estimator
exact
2.75
–
4.80
–
MC-64
2.12
–
4.58
–
MC-256
3.35
–
4.70
–
MC-1024
3.30
–
4.85
–
Appendix
Table 15: Scoped numerical and structural checks. Each sweep uses 48 task seeds, with λ=2.5 , κ=0.6 , ρϕ=0.5 , learning rate 10−3 , and 200 updates. Structural sweeps vary one coordinate of the reference (T,K,D)=(4,8,64) at a time and use separate seed cohorts, so their reference rows need not coincide. The four contrasts are defined in the text; all values are points. Dashes denote contrasts not reported for the estimator sweep.
On-policy self-distillation (OPSD) adapts a language model by distilling guidance from a frozen teacher on trajectories sampled from the student. Its effectiveness, however, depends critically on the quality of those trajectories. We show that when student rollouts drift from target trajectories, conditioning the teacher on off-target prefixes substantially weakens its task-relevant supervision. Controlled prefix-corruption experiments expose this failure mode, which we term rollout-conditioned signal degradation. To address this problem, we propose a unified training framework that separates two complementary supervision pathways. The first retains rollout-conditioned distribution matching, providing guidance on states the student actually visits. The second applies supervised cross-entropy on canonical ground-truth contexts, avoiding the incompatibility of imposing target tokens on erroneous rollout prefixes. Token-level rollout-target alignment is used to adapt the strength of the canonical-context anchor, emphasizing it during cold start and relaxing it as rollout quality improves. Experiments across multiple model scales, two task families, and general-reasoning benchmarks show that the proposed approach improves task acquisition over OPSD while preserving general capabilities, resulting in a more favorable empirical plasticity-stability trade-off. These findings identify context quality as a central bottleneck in on-policy self-distillation and demonstrate the value of separating rollout-conditioned guidance from canonical supervision.
Meilin Yang, Zixuan Ding, Jianhao Nie +5
Renmin University of China, Beijing, China · Renmin University of China, Beijing, China.
Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response. Self-distillation trains the model to retain this improvement when the context is not present. The method works by matching the model's output distribution under two settings: a student that sees only the question, and a self-teacher that also sees the context. What the model learns therefore depends on what context the self-teacher receives, yet the design of this context remains largely unexplored. We study context design for self-distillation by training a solver on feedback from a frozen critic. We compare three conditions: (i) a binary reward (GRPO), (ii) the reference solution, and (iii) a step-by-step critique aligned to the solver's reasoning trace. Step-aligned critique yields the largest gains, outperforming GRPO by 16.11 points and reference-solution-conditioned self-distillation by 5.27 points (Avg@12). Per-token advantage analysis reveals why: step-aligned feedback targets only the tokens where reasoning fails, leaving correct behavior intact. Conditioning on the reference solution, by contrast, pressures the model to change its behavior at every token (even correct steps) because an alternative derivation inevitably differs in phrasing and approach. This suggests that structural alignment between feedback and the solver's reasoning is a key driver of self-distillation effectiveness.
Self-distillation is a promising recipe for self-improvement in language models. In this setting, a model can serve as its own teacher when given privileged information, such as a solution to a math problem. This seems especially appealing for thinking models, which can use test-time reasoning to absorb the privileged information. Surprisingly, we show that privileged self-distillation degrades thinking models on long reasoning traces: across five Qwen3 and OLMo thinking models evaluated on AIME24, AIME25, and HMMT25, privileged-context distillation causes a relative drop of up to 17% in avg@16 accuracy. The degradation scales with the amount of privileged context withheld from the student and is most pronounced at long rollout budgets, where thinking models otherwise obtain their largest gains. This failure mode is not specific to self-distillation: on-policy distillation (OPD) improves thinking models, but privileged OPD reverses these gains. Our diagnostics link this failure mode to how privileged teacher context reshapes learning at high-entropy forking positions, where multiple continuations remain plausible and may lead to different reasoning paths. Privileged context lowers fork rates in thinking-model rollouts but not in instruction-model rollouts. This leads to an interesting dichotomy, where privileged context can help instruction-tuned models but hurts stronger thinking models. The effect is visible when the student begins a self-correction branch, where privileged OPD penalizes sampled reconsideration tokens that vanilla OPD supports. Thinking models trained with a privileged teacher produce fewer verification, backtracking, and hedging markers, even after length normalization. These findings indicate that self-distillation for strong thinking models requires attention to token-level signal, especially around correction and reasoning steps.
Simran Kaur, Narutatsu Ri, Yinghui He +2
Princeton Language and Intelligence, Princeton University