On-Policy Context Distillation (OPCD) has recently emerged as a powerful technique for transferring context to student models and for self-improvement. In OPCD, the teacher is conditioned on privileged information, and the goal is to minimize the Kullback-Leibler (KL) divergence between the privileged teacher and the student, evaluated on student-generated tokens. Many existing studies show that using instance-specific gold answers or gold demonstrations as the default privilege can hurt training performance, especially out-of-distribution (OOD). In this work, we instead design general instructions that target common student mistakes observed on the training samples, and show that such simple instructions can outperform gold as the OPCD privilege. In autoformalization tasks, using a matched formatting instruction as the privilege could outperform gold in OOD accuracy by a large margin. In 7 out of 8 experiments using ProverQA, ProofWriter, and ProntoQA as datasets, and Qwen3-Thinking and Olmo3-Thinking families as models, matched instruction privileges outperform gold in OOD by 4 to 17 points, while remaining on par with gold in-domain. Each instruction is only a few sentences (and thus contains much less information compared to all instance-specific gold) and is applied uniformly to every training sample. These results indicate that a general instruction, which applies equally to source and target domain examples, can be substantially more transferable than instance-specific gold in OPCD while maintaining in-domain performance.
Figures & tables
Figure 1: Pass 4 accuracy change of using gold FOL and one matched instruction for each dataset as privileges, compared to not using any privilege baseline (OPD). PR, PW, PT refer to ProverQA, ProofWriter and ProntoQA datasets. Model A→ Model B refers to teacher-student pairs in OPCD, and dataset X→ dataset Y refers to OOD performance trained on dataset X and evaluated on dataset Y . The error bars denote ±1 standard error.
Figure 2: Illustration of training students on autoformalization with OPCD. Standard OPD uses no privilege in the teacher, and standard OPCD uses gold FOL. This paper studies using different types of formatting instructions as privileges.
ProverQA
ProofWriter
ProntoQA
avg@8
pass 4
avg@8
pass 4
avg@8
pass 4
Base (no distillation)
53.3
22.5
56.7
28.9
67.0
42.5
OPSD (temp=1.0)
53.4
22.4
54.7
25.9
67.2
43.3
Teacher: Qwen3-4B
CD
58.7
28.0
73.5
40.8
90.1
73.2
CD with gold
68.4 (+9.7)
42.7 (+14.7)
81.0 (+7.5)
51.4 (+10.6)
90.4 (+0.3)
71.5 (-1.7)
Table 1: Autoformalization results on ProverQA, ProofWriter and ProntoQA. Numbers in parentheses indicate the change of accuracy when gold FOL is added to the teacher prompt as privilege.
ProverQA
ProofWriter
ProntoQA
average@8
pass 4
average@8
pass 4
average@8
pass 4
Qwen3-1.7B, no distillation
Base
78.5
64.7
79.0
69.6
82.8
68.0
Teacher: Qwen3-4B
CD
84.2
70.9
78.0
63.2
85.9
67.4
CD with verdict
85.4 (+1.2)
72.1 (+1.2)
78.2 (+0.2)
61.2 (-2.0)
85.3 (-0.6)
65.3 (-2.1)
Table 2: Direct question-answering results on three benchmarks. Numbers in parentheses indicate the change of accuracy when gold verdict is added to the teacher prompt as privilege.
Figure 3: Description of all formatting errors and their corresponding instructions. Note that a single rollout could exhibit multiple formatting errors.
Dataset
Unparseable
XOR →∨
∀ -drop
∀ -scope
Coref-split
Semantic
ProverQA
12.5
30.9
14.3
7.1
2.8
2.8
ProofWriter
8.3
0.0
28.6
10.9
0.8
1.0
ProntoQA
4.5
0.0
11.9
0.4
8.2
10.7
Table 3: Error decomposition for the Qwen3-1.7B student, across all three datasets. The numbers are the percentage of rollouts in the training set that contains a specific type of error. Red denotes the most common problems for each dataset and blue denotes the semantic problem for ProntoQA.
Dataset
Gold
Unparseable
XOR
∀ -drop
∀ -scope
Coref
ProverQA
+36.9 / +53.8
-3.4 / -6.2
+0.7 / +0.3
-7.1 / -8.4
-1.1 / -1.6
-2.4 / -2.0
ProofWriter
+35.0 / +51.3
-3.8 / -5.6
-1.5 / -3.3
+8.1 / +0.0
-4.3 / -6.7
-4.1 / -7.5
ProntoQA
+24.5 / +30.2
-3.7 / -0.8
-0.2 / +0.9
-0.3 / -2.7
-2.5 / +1.1
-1.8 / -0.1
Table 4: Direct-prompting with in-prompt privileges for Qwen3-1.7B (no training): avg@8 and pass 4 changes are shown.
Dataset
Teacher
+Gold
+Unpars
+XOR
+ ∀ -drop
+ ∀ -scope
+Coref
Avg
Best
ProverQA
4B
+8.8
+9.2
+5.7
+0.6
+5.3
+13.0
+6.8
46.2
8B
+3.2
+2.9
+3.9
-3.1
+0.3
+3.3
+1.5
47.7
ProofWriter
4B
+5.0
+5.7
+1.2
+17.1
+0.9
+0.5
+5.1
70.3
8B
-2.7
+2.0
+2.0
+3.0
-0.7
-4.8
+0.3
79.3
ProntoQA
4B
+6.6
+1.1
-0.5
-1.1
-3.0
-15.9
-3.9
89.1
8B
+9.6
+0.2
-1.5
-1.8
-3.6
-4.1
-2.2
89.0
Table 5: In-distribution pass 4 improvement over no-privilege (base) OPCD. Teachers are Qwen3-4B and Qwen3-8B and student is Qwen3-1.7B.
Data
Priv.
Acc ↑
Unp. ↓
XOR ↓
∀ drop ↓
∀ scp ↓
Coref ↓
Sem ↓
ProverQA
None
+9.0
-4.5
-4.9
-5.8
+4.0
+1.0
-0.5
+Gold
+12.7
-7.0
-20.6
-11.4
+2.5
+12.6
+3.7
+Coref
+14.1
-8.1
-5.5
-5.1
+3.7
+1.4
-1.0
ProofWriter
None
+23.1
-6.5
0.0
-12.4
-3.7
-0.6
-0.8
+Gold
+26.6
-2.4
0.0
-25.7
-6.0
-0.5
+3.2
+ ∀ drop
+32.1
-6.6
0.0
-21.9
-2.9
-0.6
-0.8
Table 6: Change (percentage points) in the error decomposition after OPCD with the Qwen3-4B teacher, relative to the no-privilege Qwen3-1.7B student (Table 3 ).
ProverQA
ProofWriter
ProntoQA
Trained on
Privilege
avg@8
pass 4
avg@8
pass 4
avg@8
pass 4
ProverQA
+Gold
+2.3
+0.7
-3.2
-8.7
-12.6
-12.5
+Coref
+4.2
+4.7
+2.8
+7.9
-0.9
-1.3
ProofWriter
+Gold
-4.6
-10.0
+6.6
+17.3
-0.6
-1.0
+ ∀ -drop
+3.4
+6.4
+7.2
+18.0
+3.7
+5.9
Table 7: In-distribution and OOD pass 4 improvement over no privilege OPCD. Teacher is Olmo-32B-Thinking and student is Olmo-7B-Thinking.
Transfer
Teacher
+Gold
+Unpars
+XOR
+ ∀ -drop
+ ∀ -scope
+Coref
Avg
Best
PR → PW
4B
+1.5
+7.5
+6.0
+14.5
+7.4
+5.4
+8.2
53.2
8B
+2.8
+2.0
+1.6
+7.2
−1.0
−4.0
+1.2
60.9
PR → PT
4B
−5.5
+0.9
+1.6
+5.2
+0.6
+0.1
+1.7
57.0
8B
−7.6
+2.9
−0.6
−2.9
−2.0
−4.5
−1.4
60.6
PW → PR
4B
−14.2
+0.7
−0.5
+1.5
+0.5
+4.4
+1.3
37.3
8B
−13.8
+2.3
−2.5
+2.6
−0.6
+3.3
+1.0
40.4
Table 8: Out-of-distribution transfer (OPCD): pass 4 improvement over no-privilege (base) distillation. Student is Qwen3-1.7B and teachers are Qwen3-4B and Qwen3-8B. PR, PW, PT represent ProverQA, ProofWriter and ProntoQA respectively.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Optimization
Steps / checkpointing
400 steps, save every 50 (8 checkpoints)
Effective batch size
8 (per-device 1 × accum 1 × 8 GPUs)
Learning rate
5×10−6
Optimizer / parallelism
DeepSpeed ZeRO, CPU optimizer offload ( cpu_adam ), 8-way
Training data
all 1200 rows, no correctness filtering (incorrect traces kept)
Seed
1 (single training seed; 8 evaluation samples)
Appendix
Table 9: Hyperparameters.
Data
Tch.
Priv.
Acc ↑
Unp. ↓
XOR ↓
∀ drop ↓
∀ scp ↓
Coref ↓
Sem ↓
ProverQA
4B
None
+9.0
-4.5
-4.9
-5.8
+4.0
+1.0
-0.5
+Gold
+12.7
-7.0
-20.6
-11.4
+2.5
+12.6
+3.7
+Coref
+14.1
-8.1
-5.5
-5.1
+3.7
+1.4
-1.0
8B
None
+12.9
-8.7
-6.4
-7.2
+2.5
+2.0
+0.2
+Gold
+16.2
-9.7
-21.5
-11.2
-1.6
+12.0
+4.7
+Coref
+14.1
-7.9
-8.1
-6.2
+0.6
+2.1
-0.6
Appendix
Table 10: Change in the error decomposition after OPCD, relative to the no-privilege Qwen3-1.7B student on evaluation rollouts.
On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises the student's trajectory. However, this interpretation conflates two effects. The reference solution not only reveals the answer to the current instance but also changes the context under which the teacher provides token-level supervision. We investigate the role of target-specific privilege with OP2SD (On-Policy Self-Distillation from Other Problems), which replaces the paired reference with a problem and solution from a different example, while preserving the student rollout, teacher, and distillation objective. Across three models and three mathematics benchmarks, OP2SD improves over the base model, remains competitive with OPSD. The success of OP2SD implies that OPSD gains do not necessarily come from access to the reference solution, and that the teacher's context-induced behavior is an important factor.
Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar +1
Mohamed bin Zayed University of Artificial Intelligence · Nagoya University · RIKEN AIP
On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher is reliable for each prompt. Because reverse KL is mode-seeking, a confidently wrong teacher can induce a strong yet misleading update. Distributional proxies, such as entropy or teacher-student likelihood agreement, measure uncertainty or agreement but do not directly verify outcome correctness. We introduce Teacher-Gated On-Policy Distillation (TGOPD), built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted. TGOPD estimates reliability from a small set of verifier-scored teacher probes and routes each prompt exclusively to dense OPD when the reliability check passes or to verifier-grounded GRPO otherwise. Across 4B and 35B students in mathematics, code, and instruction following, TGOPD outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training. By using otherwise-idle teacher capacity for reliability estimation, TGOPD also reduces teacher-side compute waste in asynchronous OPD, increasing teacher-node GPU utilization from 9.8% to 78.9% in the measured 4B single-domain run.
On-Policy Distillation (OPD) has gained wide attraction as an LLM post-training paradigm due to its effectiveness in improving capabilities without introducing model distribution drift, and consequently, regression in general tasks. On-Policy Self-Distillation (OPSD) is an efficient use-case of OPD, which is appealing as it requires only a single model as a student and teacher, and it also has the benefit of providing privileged context that is a absent at inference time (e.g. a persona, a private fact, or a worked solution) to the teacher during the training process. The challenge in this approach is that the privileged information can change model behavior more than intended: it can modify reasoning, degrade general capabilities, and affect performance indicators like response length, style, or local token preferences. Consequently, OPSD may train the student on side effects rather than a desired, transferable behavior. In this paper, we study this problem in a rare-token/identity setting and propose EviDence GuidEd On-Policy Distillation (EDGE-OPD), a modification of OPSD with two distinct characteristics: a) it uses guided rollouts to inject privileged-context behavior to the student at sampling time, so that the rare target behavior is actually present in the on-policy data, and b) it applies an evidence mask: the student is updated only at token positions where the privileged context supports the sampled token, rather than on every token in the rollout. We empirically show that OPSD (and its variant RLSD, with and without a verifier) completely fail to learn a target identity, while the integration of guided rollouts allows them to succeed. Additionally, mask-region ablations show that the persona signal is localized to the positive-evidence tail, allows us to draw valuable insights about efficient knowledge transfer and preservation of general purpose capabilities.
Aristotelis Lazaridis, Dylan Bates, Aman Sharma +3