Knowledge distillation can transfer reasoning from stronger teachers to frozen students through reusable prompts, but avoiding weight updates does not eliminate supervision. Without ground-truth answers, teacher solutions are unverified, and agreement with the teacher can reward shared mistakes. We introduce Knowledge-to-Prompt (K2P) for label-free knowledge distillation to prompts. K2P synthesizes reusable instructions from teacher solutions, refines them using paired teacher and student responses, and guides search and selection with answer agreement. It retains candidates that adaptive search may undervalue and selects on reserved questions. Deployment uses only the frozen student and selected prompt. Our theory separates generation and selection gaps and gives conditions under which agreement-guided construction yields accuracy guarantees despite imperfect teacher references. Across reasoning tasks and students, K2P outperforms label-free alternatives overall and remains competitive with supervised prompt optimization. Ablations and archive diagnostics assess the contributions of teacher solutions and refinement, while revealing the limits of agreement-guided selection.
Figures & tables
Figure 1: Label-dependent and label-free prompt construction. (a) PLD uses gold labels for extraction and correctness feedback. (b) K2P synthesizes initial prompts from unlabeled source questions and teacher solutions (1–3), then refines a search parent using paired teacher and student responses (4). Only strict improvement in search agreement updates the parent. All admissible, fully evaluated candidates enter the archive; agreement on reserved questions selects from this fixed archive after construction (5). Deployment uses only the selected prompt and frozen student.
Method
Construction
Tracking
Colored Objects
TabMWP
QuaRTz
Qwen3.5-0.8B
Zero-shot
–
21.18
49.00
65.90
82.53
PLD
Supervised
49.02 ± 28.91
43.70 ± 4.37
34.90 ± 20.79
71.98 ± 5.38
APE
79.48 ± 7.37
69.87 ± 3.35
61.57 ± 1.60
82.48 ± 0.82
GEPA
90.72 ± 6.91
79.43 ± 1.07
71.83 ± 4.76
82.48 ± 1.43
PLD-T
Label-free
47.52 ± 9.85
40.43 ± 6.72
42.90 ± 20.74
71.39 ± 0.60
Table 1: Test accuracy (%): mean ± sample standard deviation (SD) over three construction seeds; zero-shot uses one evaluation. Bold: highest mean among label-free construction methods per setting.
Figure 2: Refinement versus independent synthesis (Q: Qwen; M: Ministral). (a) Selected and (b) archive-best test accuracy: mean ± sample SD over three seeds; Δ is K2P minus independent synthesis in percentage points (pp). (c) All 96 revisions: test-accuracy changes from actual parents, grouped by whether the search parent was updated.
Figure 3: Agreement and selection: 288 candidates, 24 archives, three seeds. (a) Deficits from each archive’s highest observed agreement and accuracy on the same test questions (inset: full range); circles: Qwen, triangles: Ministral; dashed lines: equal deficits. (b) Highest minus selected test accuracy within each archive; selection uses the original reserved set. Points: means; bars: sample SD across seeds.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Detailed K2P construction and deployment workflow. Teacher solutions provide knowledge for prompt synthesis and refinement. After search evaluation, all initial candidates enter the archive C , while the search winner p0 starts a single refinement trajectory. The teacher revises the current parent using student feedback and relevant solutions. Every valid revision with a complete search evaluation enters the archive; only strict search improvement updates the parent. Reserved agreement selects pS∗ from the complete archive using the same target student. Reserved questions are held out from refinement. Both models remain frozen, and deployment uses only the student with the selected prompt.
Role
Model repository
Revision
Teacher
Qwen/Qwen3.5-9B
c202236235762e1c871ad0ccb60c8ee5ba337b9a
First target student
Qwen/Qwen3.5-0.8B
2fc06364715b967f1860aea9cf38778875588b17
Second target student
mistralai/ Ministral-3-3B- Instruct-2512-BF16
b6d637bef2393152b3da2b2fde72eecdee30557e
Appendix
Table A1: Pinned model identities. Model repository identifiers are authoritative; abbreviated student names elsewhere refer to these same revisions.
Setting
Method
Teacher input
Teacher output
Student input
Student output
Tracking/Q
PLD
0.112
0.027
0.276
0.321
APE
0.045
0.031
1.913
1.355
GEPA
0.112
0.031
1.592
0.922
PLD-T
0.198
0.108
0.438
0.241
GEPA-T †
0.118
0.093
1.036
0.718
SPO
0.198
0.060
0.015
0.012
Appendix
Table A2: Construction token usage (millions), seed 0. Q and M denote the Qwen and Ministral target students; the teacher is Qwen3.5-9B throughout. K2P provides the reference construction.
Task
Student
PLD
APE
GEPA
PLD-T
GEPA-T
K2P
Tracking
Qwen
412
604
957
481
743
585
Tracking
Ministral
287
253
600
780
859
135
Colored Objects
Qwen
256
413
801
641
1269
751
Colored Objects
Ministral
834
339
2089
389
1225
415
TabMWP
Qwen
156
68
724
262
1145
183
TabMWP
Ministral
1234
71
826
1025
556
782
Appendix
Table A3: Deployed instruction lengths, seed 0. Tokens under the target student’s tokenizer, including the method’s deployed answer-format suffix and excluding the test question and native chat-template overhead.
Task
Student
PLD
APE
GEPA
PLD-T
GEPA-T
Tracking
Qwen
[41.37, 50.98]
[4.31, 11.96]
[-1.96, 5.10]
[49.80, 59.41]
[0.20, 7.84]
Tracking
Ministral
[-5.49, -1.76]
[0.00, 5.49]
[-5.88, -1.76]
[-4.71, -0.20]
[-0.59, 4.90]
Colored Objects
Qwen
[40.40, 47.60]
[15.60, 22.20]
[4.40, 10.20]
[49.30, 56.40]
[5.40, 11.30]
Colored Objects
Ministral
[-1.20, 1.60]
[7.70, 12.10]
[0.80, 4.20]
[15.60, 20.80]
[22.50, 28.40]
TabMWP
Qwen
[8.72, 15.68]
[2.09, 8.98]
[-15.30, -8.43]
[8.25, 15.48]
[-10.38, -3.22]
TabMWP
Ministral
[39.84, 46.65]
[6.08, 11.13]
[-0.60, 3.68]
[30.15, 36.67]
[1.81, 6.12]
Appendix
Table A4: Paired test uncertainty for seed 0. Pointwise 95% intervals for K2P minus each baseline, in percentage points.
Task
Student
Seed
PLD
APE
GEPA
PLD-T
GEPA-T
K2P
Tracking
Qwen3.5-0.8B
0
45.29
83.33
89.80
36.67
87.45
91.37
20260715
79.61
84.12
98.04
50.00
91.37
89.41
20260609
22.16
70.98
84.31
55.88
85.49
86.86
Tracking
Ministral-3-3B
0
98.63
92.35
98.82
97.45
92.94
95.10
20260715
95.49
90.00
97.65
92.16
97.25
96.67
20260609
99.41
84.12
96.08
94.90
96.08
94.31
Appendix
Table A5: Results for three specified construction seeds. Accuracy (%). PLD, APE, and GEPA use gold construction labels; PLD-T, GEPA-T, and K2P use none. Qwen and Ministral refer to the same pinned students as in Table 1 . SPO results appear in Table A6 .
Seed 0
Seed 20260715
Seed 20260609
Task
Student
SPO
K2P
SPO
K2P
SPO
K2P
Tracking
Qwen
59.41
91.37
68.04
89.41
70.78
86.86
Tracking
Ministral
97.65
95.10
93.53
96.67
77.65
94.31
Colored Objects
Qwen
65.30
85.80
68.70
78.90
77.20
80.70
Colored Objects
Ministral
94.10
96.90
92.60
94.10
92.60
96.30
TabMWP
Qwen
64.20
65.40
33.90
66.30
68.90
74.40
Appendix
Table A6: SPO and K2P test accuracy (%) for each construction seed.
(a) Three-seed construction results
Task
Student
Without teacher solutions
K2P
Difference (pp)
Tracking
Qwen3.5-0.8B
89.41 ± 2.16
89.22 ± 2.26
-0.20
Tracking
Ministral-3-3B
94.18 ± 4.21
95.36 ± 1.20
+1.18
Colored Objects
Qwen3.5-0.8B
72.17 ± 3.79
81.80 ± 3.58
+9.63
Colored Objects
Ministral-3-3B
96.07 ± 1.40
95.77 ± 1.47
-0.30
TabMWP
Qwen3.5-0.8B
67.57 ± 2.14
68.70 ± 4.96
+1.13
Appendix
Table B1: Contribution of worked teacher solutions. (a) Accuracy (%): mean ± sample SD over three construction seeds. Differences are K2P minus solution removal, computed before rounding. (b) Seed 0 paired 95% test intervals, conditional on its frozen prompts. TabMWP resamples table groups; QuaRTz resamples background groups.
Task
Student
Seed 0
Seed 20260715
Seed 20260609
Tracking
Qwen
89.41
91.57
87.25
Tracking
Ministral
99.02
92.16
91.37
Colored Objects
Qwen
75.60
68.10
72.80
Colored Objects
Ministral
97.40
96.20
94.60
TabMWP
Qwen
66.00
70.00
66.70
TabMWP
Ministral
82.90
65.40
78.20
Appendix
Table B2: Individual test accuracies (%) without teacher solutions. Each construction uses complete 80-question reserved selection. The corresponding K2P accuracies appear in Table A5 .
Setting
Initial bank
Independent synthesis
K2P
(a) Reserved-selected test accuracy
Tracking/Q
71.70 ± 3.81
71.70 ± 3.81
89.22 ± 2.26
Tracking/M
95.36 ± 1.20
95.42 ± 1.31
95.36 ± 1.20
Colored/Q
53.00 ± 4.65
59.17 ± 4.58
81.80 ± 3.58
Colored/M
96.17 ± 2.00
95.03 ± 1.14
95.77 ± 1.47
TabMWP/Q
64.27 ± 2.78
63.03 ± 4.90
68.70 ± 4.96
Appendix
Table B3: Accuracy (%): mean ± sample SD across three construction seeds. Q and M denote Qwen and Ministral. The candidate sets share the same eight initial prompts; independent synthesis and K2P use four additional proposal opportunities.
Setting
One feedback case
Three feedback cases
Five feedback cases
Tracking/Q
83.14 ± 9.13
89.22 ± 2.26
81.24 ± 9.06
Tracking/M
95.56 ± 0.97
95.36 ± 1.20
96.60 ± 1.47
Colored/Q
62.10 ± 7.07
81.80 ± 3.58
71.03 ± 2.28
Colored/M
96.17 ± 2.00
95.77 ± 1.47
97.67 ± 1.14
TabMWP/Q
70.27 ± 7.66
68.70 ± 4.96
69.03 ± 5.53
TabMWP/M
77.07 ± 10.17
90.83 ± 1.86
82.40 ± 7.66
Appendix
Table B4: Test accuracy (%): mean ± sample SD across three construction seeds. All conditions use the full 80-question reserved set. Q and M denote Qwen and Ministral. The last row averages the eight setting means.
Selection questions
Test accuracy (%)
Selection loss (pp)
Full-set choice (%)
20
84.81
3.11
43.65
40
86.02
1.90
60.07
60
86.47
1.45
75.14
80
86.84
1.08
100.00
Appendix
Table B5: Selection-size sensitivity on 24 frozen K2P archives. Selection loss is archive-best minus selected test accuracy; full-set choice is the fraction of selections matching the original 80-question choice. Candidate archives and their responses are fixed throughout.
(a) Search size; reserved selection fixed at 80
Setting
40
Original 80
120
Δ80→120 (pp)
TabMWP/Q
69.83±6.91
68.70±4.96
65.80±0.46
-2.90
TabMWP/M
82.13±15.88
90.83±1.86
79.23±16.08
-11.60
QuaRTz/Q
82.19±1.94
81.59±3.38
81.59±3.38
+0.00
QuaRTz/M
91.41±0.99
91.45±1.42
91.45±1.42
+0.00
(b) Selection size; original search-80 archive fixed
Appendix
Table B6: Search and selection size sensitivity on TabMWP and QuaRTz (Q: Qwen; M: Ministral). Test accuracy (%), mean ± sample SD over three construction seeds. In (b), each seed contributes its mean over 1,000 paired subsets at 40 and 80, and one full-pool choice at 120; original 80 reports the unchanged main-evaluation selection set. Δ80→120 compares paired search constructions.
(a) Search size; selected test accuracy (%)
Setting
Seed
40
Original 80
120
Δ80→120 (pp)
TabMWP/Q
0
65.40
65.40
65.40
+0.00
TabMWP/Q
20260715
66.30
66.30
66.30
+0.00
TabMWP/Q
20260609
77.80
74.40
65.70
-8.70
TabMWP/M
0
91.20
91.70
91.40
-0.30
TabMWP/M
20260715
63.80
88.70
85.30
-3.40
Appendix
Table B7: Per-seed results for Table B6 . In (b), ± denotes the SD of selected test accuracy across the 1,000 subsets within one frozen archive ; Table B6 instead reports SD across the three seed-level means. Full 120 and original 80 each use one fixed selection set.
Setting
Seed 0
Seed 20260715
Seed 20260609
Mean ± SD
Tracking/Q
2.16
2.55
0.00
1.57 ± 1.37
Tracking/M
3.33
0.00
1.96
1.76 ± 1.68
Colored/Q
0.00
0.00
0.00
0.00 ± 0.00
Colored/M
1.20
2.70
1.70
1.87 ± 0.76
TabMWP/Q
0.00
0.00
0.00
0.00 ± 0.00
TabMWP/M
0.00
0.80
1.50
0.77 ± 0.75
Appendix
Table B8: K2P selection loss in percentage points: archive-best minus reserved-selected test accuracy, computed separately within each run.
Setting
Search agreement
Test agreement
Difference (pp)
Tracking/Q
90.83 ± 6.17
86.54 ± 4.74
4.30 ± 1.53
Tracking/M
96.67 ± 0.72
93.33 ± 3.98
3.33 ± 4.56
Colored/Q
85.00 ± 3.31
80.23 ± 1.82
4.77 ± 3.64
Colored/M
98.75 ± 1.25
96.63 ± 1.77
2.12 ± 1.03
TabMWP/Q
71.67 ± 11.27
65.23 ± 7.57
6.43 ± 3.91
TabMWP/M
89.58 ± 4.02
90.10 ± 2.10
-0.52 ± 2.19
Appendix
Table B9: Final search-parent agreement (%), mean ± sample SD over three seeds. Differences are computed within each seed before aggregation.
Setting
Selected gain [95% CI], pp
Selection loss [95% band], pp
Spearman’s ρ
Tracking/Q
20.59 [15.88, 25.10]
2.16 [0.00, 8.63]
0.998
Tracking/M
0.00 [0.00, 0.00]
3.33 [0.00, 7.84]
0.988
Colored/Q
37.90 [34.40, 41.50]
0.00 [0.00, 4.90]
0.998
Colored/M
-1.20 [-2.40, -0.10]
1.20 [0.00, 3.50]
0.979
TabMWP/Q
0.00 [0.00, 0.00]
0.00 [0.00, 5.19]
0.993
TabMWP/M
28.60 [25.47, 31.79]
0.00 [0.00, 4.38]
1.000
Appendix
Table B10: Seed 0 conditional uncertainty and agreement–accuracy rank correlation. Selected gain compares full-archive selection with initial-only selection.
Setting
Seed
Parent → revision
Update
Search
Repairs
Regressions
Test
Tracking/Q
0
I6 → R1
Yes
5.00
103
38
12.75
R1 → R2
Yes
3.75
46
26
3.92
R2 → R3
Yes
6.25
32
43
-2.16
R3 → R4
No
-2.50
38
38
0.00
Tracking/Q
20260715
I7 → R1
No
-1.25
96
78
3.53
I7 → R2
Yes
15.00
121
38
16.27
Appendix
Table B11: All 96 revisions relative to their actual parents. Search and test columns report agreement and accuracy changes (pp); repairs and regressions count test questions. “Update” denotes a strict search agreement increase. I and R index initial candidates and revisions.
Figure 5: K2P reserved-selected and archive-best test accuracy across three construction seeds (Q: Qwen; M: Ministral). Points and bars show means and sample standard deviations. Their within-run difference is the selection loss in Figure 3 (b).
Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across training stages remains unclear. Through controlled experiments, we find that forward Kullback-Leibler (KL) distillation--the standard KD formulation--with post-trained teachers behaves fundamentally differently during mid-training, an intermediate phase of self-supervised learning on curated corpora. Surprisingly, while forward KD simultaneously improves reasoning and factual recall during pre-training relative to standard next-token prediction (NTP), it instead slows factual recall acquisition during mid-training despite continued reasoning gains. We trace this stage dependence to an asymmetry in teacher confidence across data domains and the student's evolving knowledge state: teachers are more confident on procedural than knowledge-intensive data, while students acquire low-entropy factual knowledge earlier in training. To mitigate this imbalance, we propose Switch Distillation, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy. Switch Distillation consistently outperforms existing distillation objectives across teacher sizes. Relative to standard NTP, it achieves 1.61-1.71x the reasoning performance and 1.13-1.19x the knowledge and commonsense performance while preserving 96.7-96.8% of factual recall. Crucially, these benefits persist after post-training: Switch Distillation closes the factual recall gap while maintaining 1.25-1.32x and 1.13-1.20x gains in reasoning and knowledge and commonsense, respectively.
Jacqueline He, Howard Yen, Shuyue Stella Li +9
Meta AI · University of Washington · Princeton University
Distilling the capabilities from a large reasoning model (LRM) to a smaller student model often involves training on substantial amounts of reasoning data. However, knowledge distillation (KD) over lengthy sequences with prompt (P), chain-of-thought (CoT), and answer (A) sections makes the process computationally expensive. In this work, we investigate how the allocation of supervision across different sections (P, CoT, A) affects student performance. Our analysis shows that selective KD over only the CoT tokens can be effective when the prompt and answer information is encompassed by it. Building on this insight, we establish a truncation protocol to quantify computation-quality tradeoffs as a function of sequence length. We observe that beyond a specific length, longer training sequences provide marginal returns for downstream performance but require substantially higher memory and FLOPs. To this end, training on only the first 50% of tokens of every training sequence can retain, on average, ≈91% of full-sequence performance on math benchmarks while reducing training time, memory usage, and FLOPs by about 50% each. Codes are available at https://github.com/weiruichen01/distilling-the-essence.
Wei-Rui Chen, Vignesh Kothapalli, Ata Fatahibaarzi +5
1The University of British Columbia · 2LinkedIn · 3Canada Research Chair in NLP and ML
Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out challenging examples through data selection or introduce weaker intermediate assistant models, inherently compromising supervision coverage or quality. We propose Teacher Alignment, which directly adapts the teacher toward the student's distribution without discarding data or degrading reasoning quality. However, naive alignment through standard knowledge distillation triggers catastrophic collapse of the teacher's reasoning capabilities. To address this, we reformulate teacher alignment as reinforcement learning and introduce TeacherGRPO, built on Group Relative Policy Optimization with two key innovations: (i) Curriculum Selective Alignment applies dual token- and distribution-level curricula to focus rewards on high-signal reasoning gaps while filtering noise from trivial tokens and uncertain tail distributions, and (ii) Importance-Adaptive Length Regularization selectively penalizes verbose redundancy while preserving pedagogically critical reasoning steps. The aligned teacher then distills knowledge to students via standard pipelines. Extensive experiments show TeacherGRPO significantly outperforms baselines across diverse reasoning benchmarks and distillation methods. Our code is available at https://github.com/LzyFischer/TeacherGRPO.
Zhenyu Lei, Zihan Chen, Yaochen Zhu +5
University of Virginia · Netflix · University of Washington +2