Knowledge distillation (KD) transfers knowledge from stronger Teacher models to weaker Student models, but most methods require training the Student parameters, thereby binding the distilled knowledge to a specific architecture and checkpoint. This implicit representation is difficult to interpret or reuse across models and limits KD for API-only or costly-to-train models. This paper studies knowledge transfer for large language models (LLMs). We introduce Universal Textual Teaching (UTT), a parameter-update-free framework that distills observed Teacher-Student knowledge gaps into a textual, interpretable, and reusable natural-language artifact called Primer. Specifically, UTT first identifies representative gap cases through paired evaluations, and iteratively updates the Primer via multi-role interactions: the Student attempts each task, the Prompter turns evaluation feedback into a teaching instruction, the Teacher provides a targeted demonstration, and the Synthesizer consolidates validated lessons. Empirically, on the challenging math (Omni-MATH-2) and code generation (KernelBench) tasks, extensive results confirm the effectiveness of the method: UTT remarkably raises the Student's accuracy from 9.4% to 48.6% and Fast1 accuracy from 9% to 35% on KernelBench, while increasing mathematical reasoning accuracy from 27.6% to 51.7%. UTT also performs better than representative prompt engineering and parameter-based KD methods. Of note, UTT is shown to be generalizable across different Teachers and Students: a Primer synthesized for one Teacher-Student pair can generalize to other Students that do not participate in the synthesis.
Figures & tables
Figure 1: Left: We introduce U niversal T extual T eaching (UTT), a parameter-update-free framework that distills an LLM’s knowledge into a textual Primer through multi-role LLM interaction. Notably, the resulting Primer can improve the performance of the source Student and other Students that are not involved in creating the Primer . Right: For the target Student Qwen3.6-27B, Primers synthesized from the Teacher - Student pairs (shown at the bottom of the figure) consistently improve accuracy. Even when Qwen3.6-27B does not participate in the Primer synthesis, the absolute accuracy gains reach 41.8% and 31.5% (see the rightmost column) on KernelBench and Omni-MATH-2, respectively. These results suggest that text can serve as a universal medium for knowledge transfer among different LLMs for complex tasks.
Figure 2: Overview of Universal Textual Teaching (UTT). (1) Knowledge-gap construction. UTT first evaluates the knowledge gap between the Teacher and Student and accordingly partitions D into subsets (Section 3.2.1 ). (2) Primer synthesis. It then synthesizes a textual Primer through iterative interactions among multiple LLMs (Section 3.2.2 ). (3) Evaluation. Finally, the resulting Primer is applied to the source Student and transferred to other Students , improving both the source Student and the transfer Students . All model parameters remain frozen throughout the process. In the diagram, rectangles denote models, with the same color indicating the same model serving different roles; rounded rectangles denote data subsets, and diamonds denote evaluation procedures.
Primer synthesis
Target Student
Teacher
Teacher score
Source Student
Flash
Qwen
KernelBench Accuracy (%)
No Primer
9.4
9.2
Pro
19.6
Flash
48.6 (+39.2) ①
41.2 (+32.0) ③
Pro
19.6
Qwen
43.8 (+34.4) ④
50.0 (+40.8) ②
Opus
36.0
Flash
48.2 (+38.8) ⑤
51.0 (+41.8) ⑥
Table 1: Primer effectiveness and cross-model transferability. For each task, each Teacher and source Student pair produces one Primer , transferring to the target Student . ①–⑥ denote different experimental groups. Parentheses show gains over the corresponding Student baseline, and bold values indicate the best result. All scores are computed over the full task pool D .
Figure 3: Split-wise sample accuracy for KernelBench and Omni-MATH-2. Panels (a)–(c) show KernelBench results and panels (d)–(f) show Omni-MATH-2 results. From left to right, the three columns correspond to Pro → Flash, Pro → Qwen, and Opus → Flash, respectively.
Method
Omni-MATH-2
KernelBench
Flash
Qwen
Flash
Qwen
Acc.
Acc.
Acc.
Fast 1
Acc.
Fast 1
Student
27.6
34.3
9.4
9
9.2
12
Prompting Techniques (PT)
One-shot
35.0
57.7
39.2
23
45.8
31
Few-shot
35.4
54.8
40.6
25
46.6
30
Table 2: Comparison with prompt engineering methods. “ Teacher summary” means asking the Teacher to summarize the training set into a general prompt.
Method
Omni-MATH-2
KernelBench
Acc.
Acc.
Fast 1
Student
34.3
9.2
12
SeqKD ( Kim and Rush, 2016 )
53.5
18.8
18
Fine-tune-CoT ( Ho et al., 2023 )
39.5
41.8
28
RSR ( Yang et al., 2026 )
51.5
22.6
19
LUFFY ( Yan et al., 2025 )
55.2
11.4
13
Table 3: Comparison with KD methods under the Pro → Qwen setting.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Teacher → Student
Ddisttr
Ddistte
Dret
Dfrt
Total
KernelBench
Pro → Flash
135
55
100
210
500
Pro → Qwen
140
60
100
200
Opus → Flash
215
95
70
120
Omni-MATH-2
Pro → Flash
490
210
100
200
1000
Pro → Qwen
335
145
360
160
Opus → Flash
645
275
35
45
Appendix
Table 4: Numbers of generated samples associated with the knowledge-gap portions across tasks and Teacher– Student configurations.
Primer synthesis
Target Student
Teacher
Teacher score
Source Student
Flash 4.1
Qwen
No Primer
54.0
41.5
Max
94.0
Qwen
59.0 (+5.0)
76.0 (+34.5)
Max
94.0
Flash 4.1
56.0 (+2.0)
72.0 (+30.5)
Opus
93.5
Qwen
57.5 (+3.5)
80.5 (+39.0)
Appendix
Table 5: Primer effectiveness and cross-model transferability on Vision Dominant plane geometry problems from MathVerse. Each model is evaluated using two samples per problem. Teacher scores denote Teacher accuracy. Parentheses show percentage-point gains over the corresponding target Student baseline, and bold values indicate the best result for each target Student . Max, Qwen, Flash 4.1, and Opus denote Qwen3.8-Max, Qwen3.6-27B, DeepSeek-V4.1-Flash, and Claude Opus 5, respectively.
Figure 4: Effect of output-token budget on Omni-MATH-2 accuracy. Curves compare the unprimed Flash Student , the Pro Teacher, APE, GEPA, and UTT. UTT uses the same Primer synthesized under the 32k-token budget at every evaluated budget, without resynthesis.
Method / Role
Compute
Time (h)
Model calls
Input tokens
Output tokens
Parameter-updating KD
SeqKD
4 × A6000
1.2
–
–
–
Fine-tune-CoT
4 × A6000
3.5
–
–
–
RSR
4 × A6000
7.5
–
–
–
LUFFY
8 × A6000
22.0
–
–
–
UTT
Appendix
Table 6: Stage-specific resource usage for the Omni-MATH-2 results in Table 3 under the Pro → Qwen setting. We treat the task data as given and focus on the method-specific optimization stage: KD reports parameter training only, whereas UTT reports the time, model calls, and tokens used for API-based Primer synthesis. Because the two method families consume different resource types, these measurements characterize their respective resource profiles rather than constituting a strictly equivalent end-to-end cost comparison. Teacher side aggregates the Teacher, Prompter, and Synthesizer.
Variant
Batch
Ddisttr
Ddistte
Dret
Dfrt
Student without Primer
–
33.7
33.3
60.0
0.0
Full method
16
65.3
56.0
72.5
5.0
Random selection
16
55.6
50.0
72.5
3.8
Without Prompter
16
72.4
48.8
70.0
5.0
Without validation gate
16
59.2
39.3
80.0
3.8
Batch size 8
8
53.6
41.7
67.5
2.5
Appendix
Table 7: Core-component and batch-size ablations under Pro → Flash on Omni-MATH-2. Bold values indicate the best result for each subset.
Configuration
Accuracy (%)
Fast 1 (%)
Baseline
With Primer
Gain
Baseline
With Primer
Gain
Unquantized
9.2
41.2
+32.0
12
37
+25
Int4 RTN (G128)
4.4
34.6
+30.2
5
28
+23
Int4 RTN (Channelwise)
0.4
21.6
+21.2
0
16
+16
Thinking disabled
1.6
16.6
+15.0
1
7
+6
Appendix
Table 8: Primer gains under different quantization and thinking settings of Qwen3.6-27B on KernelBench.
Knowledge distillation is an established technique for improving the capabilities of small, efficient student models by training them with the representations of larger, more capable teacher models. Much of the recent work in the distillation of large language models (LLMs) has focused on distilling abilities learned during post-training, such as instruction following, chain-of-thought reasoning, and tool usage. This has left a large research gap in general knowledge distillation for LLMs, which is essential for developing efficient and private systems suitable for deployment on edge devices. We take a first-principles approach, evaluating previous lessons from prior works and conducting new explorations to develop a distillation methodology suitable for modern LLMs. We present KDFP, a novel methodology for white-box general knowledge distillation in LLMs. We demonstrate that KDFP outperforms existing methods by 1.6% − 4.9% across 9 benchmarks while increasing training efficiency by up to 99.1% through ephemeral parameter reduction.
Ryan Swift, Konstantinos Psounis
Thomas Lord Department of Computer Science University of Southern California Los Angeles, CA 90007, USA
Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time. Knowledge distillation (KD) offers a practical solution by transferring knowledge from a large teacher model to a smaller student model via alignment of discrete probability distributions. However, existing KD methods for LLMs primarily rely on divergences that evaluate discrepancies through probability values at each vocabulary index, without explicitly leveraging token-level semantic information. We propose Wasserstein-based knowledge distillation (WASD) for LLMs, which incorporates token-level semantic information via the Wasserstein-based distance with a cost matrix derived from token embeddings. To ensure computational tractability, we adopt the Sinkhorn divergence and derive a gradient-equivalent objective that can be efficiently optimized without introducing additional networks. Experiments across multiple LLM families and scales show that WASD consistently improves distillation performance on diverse tasks, including instruction following, mathematical reasoning, and code generation. Our results highlight the importance of semantic information encoded in the token space for effective distribution alignment in LLM distillation. The implementation is publicly available at https://github.com/aailab-kaist/WASD .
Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students' performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.
Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang +2
Stanford University · Georgia Tech · Carnegie Mellon University