On-Policy Context Distillation (OPCD) has recently emerged as a powerful technique for transferring context to student models and for self-improvement. In OPCD, the teacher is conditioned on privileged information, and the goal is to minimize the Kullback-Leibler (KL) divergence between the privileged teacher and the student, evaluated on student-generated tokens. Many existing studies show that using instance-specific gold answers or gold demonstrations as the default privilege can hurt training performance, especially out-of-distribution (OOD). In this work, we instead design general instructions that target common student mistakes observed on the training samples, and show that such simple instructions can outperform gold as the OPCD privilege. In autoformalization tasks, using a matched formatting instruction as the privilege could outperform gold in OOD accuracy by a large margin. In 7 out of 8 experiments using ProverQA, ProofWriter, and ProntoQA as datasets, and Qwen3-Thinking and Olmo3-Thinking families as models, matched instruction privileges outperform gold in OOD by 4 to 17 points, while remaining on par with gold in-domain. Each instruction is only a few sentences (and thus contains much less information compared to all instance-specific gold) and is applied uniformly to every training sample. These results indicate that a general instruction, which applies equally to source and target domain examples, can be substantially more transferable than instance-specific gold in OPCD while maintaining in-domain performance.
Figures & tables
Figure 1: Pass 4 accuracy change of using gold FOL and one matched instruction for each dataset as privileges, compared to not using any privilege baseline (OPD). PR, PW, PT refer to ProverQA, ProofWriter and ProntoQA datasets. Model A→ Model B refers to teacher-student pairs in OPCD, and dataset X→ dataset Y refers to OOD performance trained on dataset X and evaluated on dataset Y . The error bars denote ±1 standard error.
Figure 2: Illustration of training students on autoformalization with OPCD. Standard OPD uses no privilege in the teacher, and standard OPCD uses gold FOL. This paper studies using different types of formatting instructions as privileges.
ProverQA
ProofWriter
ProntoQA
avg@8
pass 4
avg@8
pass 4
avg@8
pass 4
Base (no distillation)
53.3
22.5
56.7
28.9
67.0
42.5
OPSD (temp=1.0)
53.4
22.4
54.7
25.9
67.2
43.3
Teacher: Qwen3-4B
CD
58.7
28.0
73.5
40.8
90.1
73.2
CD with gold
68.4 (+9.7)
42.7 (+14.7)
81.0 (+7.5)
51.4 (+10.6)
90.4 (+0.3)
71.5 (-1.7)
Table 1: Autoformalization results on ProverQA, ProofWriter and ProntoQA. Numbers in parentheses indicate the change of accuracy when gold FOL is added to the teacher prompt as privilege.
ProverQA
ProofWriter
ProntoQA
average@8
pass 4
average@8
pass 4
average@8
pass 4
Qwen3-1.7B, no distillation
Base
78.5
64.7
79.0
69.6
82.8
68.0
Teacher: Qwen3-4B
CD
84.2
70.9
78.0
63.2
85.9
67.4
CD with verdict
85.4 (+1.2)
72.1 (+1.2)
78.2 (+0.2)
61.2 (-2.0)
85.3 (-0.6)
65.3 (-2.1)
Table 2: Direct question-answering results on three benchmarks. Numbers in parentheses indicate the change of accuracy when gold verdict is added to the teacher prompt as privilege.
Figure 3: Description of all formatting errors and their corresponding instructions. Note that a single rollout could exhibit multiple formatting errors.
Dataset
Unparseable
XOR →∨
∀ -drop
∀ -scope
Coref-split
Semantic
ProverQA
12.5
30.9
14.3
7.1
2.8
2.8
ProofWriter
8.3
0.0
28.6
10.9
0.8
1.0
ProntoQA
4.5
0.0
11.9
0.4
8.2
10.7
Table 3: Error decomposition for the Qwen3-1.7B student, across all three datasets. The numbers are the percentage of rollouts in the training set that contains a specific type of error. Red denotes the most common problems for each dataset and blue denotes the semantic problem for ProntoQA.
Dataset
Gold
Unparseable
XOR
∀ -drop
∀ -scope
Coref
ProverQA
+36.9 / +53.8
-3.4 / -6.2
+0.7 / +0.3
-7.1 / -8.4
-1.1 / -1.6
-2.4 / -2.0
ProofWriter
+35.0 / +51.3
-3.8 / -5.6
-1.5 / -3.3
+8.1 / +0.0
-4.3 / -6.7
-4.1 / -7.5
ProntoQA
+24.5 / +30.2
-3.7 / -0.8
-0.2 / +0.9
-0.3 / -2.7
-2.5 / +1.1
-1.8 / -0.1
Table 4: Direct-prompting with in-prompt privileges for Qwen3-1.7B (no training): avg@8 and pass 4 changes are shown.
Dataset
Teacher
+Gold
+Unpars
+XOR
+ ∀ -drop
+ ∀ -scope
+Coref
Avg
Best
ProverQA
4B
+8.8
+9.2
+5.7
+0.6
+5.3
+13.0
+6.8
46.2
8B
+3.2
+2.9
+3.9
-3.1
+0.3
+3.3
+1.5
47.7
ProofWriter
4B
+5.0
+5.7
+1.2
+17.1
+0.9
+0.5
+5.1
70.3
8B
-2.7
+2.0
+2.0
+3.0
-0.7
-4.8
+0.3
79.3
ProntoQA
4B
+6.6
+1.1
-0.5
-1.1
-3.0
-15.9
-3.9
89.1
8B
+9.6
+0.2
-1.5
-1.8
-3.6
-4.1
-2.2
89.0
Table 5: In-distribution pass 4 improvement over no-privilege (base) OPCD. Teachers are Qwen3-4B and Qwen3-8B and student is Qwen3-1.7B.
Data
Priv.
Acc ↑
Unp. ↓
XOR ↓
∀ drop ↓
∀ scp ↓
Coref ↓
Sem ↓
ProverQA
None
+9.0
-4.5
-4.9
-5.8
+4.0
+1.0
-0.5
+Gold
+12.7
-7.0
-20.6
-11.4
+2.5
+12.6
+3.7
+Coref
+14.1
-8.1
-5.5
-5.1
+3.7
+1.4
-1.0
ProofWriter
None
+23.1
-6.5
0.0
-12.4
-3.7
-0.6
-0.8
+Gold
+26.6
-2.4
0.0
-25.7
-6.0
-0.5
+3.2
+ ∀ drop
+32.1
-6.6
0.0
-21.9
-2.9
-0.6
-0.8
Table 6: Change (percentage points) in the error decomposition after OPCD with the Qwen3-4B teacher, relative to the no-privilege Qwen3-1.7B student (Table 3 ).
ProverQA
ProofWriter
ProntoQA
Trained on
Privilege
avg@8
pass 4
avg@8
pass 4
avg@8
pass 4
ProverQA
+Gold
+2.3
+0.7
-3.2
-8.7
-12.6
-12.5
+Coref
+4.2
+4.7
+2.8
+7.9
-0.9
-1.3
ProofWriter
+Gold
-4.6
-10.0
+6.6
+17.3
-0.6
-1.0
+ ∀ -drop
+3.4
+6.4
+7.2
+18.0
+3.7
+5.9
Table 7: In-distribution and OOD pass 4 improvement over no privilege OPCD. Teacher is Olmo-32B-Thinking and student is Olmo-7B-Thinking.
Transfer
Teacher
+Gold
+Unpars
+XOR
+ ∀ -drop
+ ∀ -scope
+Coref
Avg
Best
PR → PW
4B
+1.5
+7.5
+6.0
+14.5
+7.4
+5.4
+8.2
53.2
8B
+2.8
+2.0
+1.6
+7.2
−1.0
−4.0
+1.2
60.9
PR → PT
4B
−5.5
+0.9
+1.6
+5.2
+0.6
+0.1
+1.7
57.0
8B
−7.6
+2.9
−0.6
−2.9
−2.0
−4.5
−1.4
60.6
PW → PR
4B
−14.2
+0.7
−0.5
+1.5
+0.5
+4.4
+1.3
37.3
8B
−13.8
+2.3
−2.5
+2.6
−0.6
+3.3
+1.0
40.4
Table 8: Out-of-distribution transfer (OPCD): pass 4 improvement over no-privilege (base) distillation. Student is Qwen3-1.7B and teachers are Qwen3-4B and Qwen3-8B. PR, PW, PT represent ProverQA, ProofWriter and ProntoQA respectively.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Optimization
Steps / checkpointing
400 steps, save every 50 (8 checkpoints)
Effective batch size
8 (per-device 1 × accum 1 × 8 GPUs)
Learning rate
5×10−6
Optimizer / parallelism
DeepSpeed ZeRO, CPU optimizer offload ( cpu_adam ), 8-way
Training data
all 1200 rows, no correctness filtering (incorrect traces kept)
Seed
1 (single training seed; 8 evaluation samples)
Appendix
Table 9: Hyperparameters.
Data
Tch.
Priv.
Acc ↑
Unp. ↓
XOR ↓
∀ drop ↓
∀ scp ↓
Coref ↓
Sem ↓
ProverQA
4B
None
+9.0
-4.5
-4.9
-5.8
+4.0
+1.0
-0.5
+Gold
+12.7
-7.0
-20.6
-11.4
+2.5
+12.6
+3.7
+Coref
+14.1
-8.1
-5.5
-5.1
+3.7
+1.4
-1.0
8B
None
+12.9
-8.7
-6.4
-7.2
+2.5
+2.0
+0.2
+Gold
+16.2
-9.7
-21.5
-11.2
-1.6
+12.0
+4.7
+Coref
+14.1
-7.9
-8.1
-6.2
+0.6
+2.1
-0.6
Appendix
Table 10: Change in the error decomposition after OPCD, relative to the no-privilege Qwen3-1.7B student on evaluation rollouts.