Knowledge distillation (KD) transfers knowledge from stronger Teacher models to weaker Student models, but most methods require training the Student parameters, thereby binding the distilled knowledge to a specific architecture and checkpoint. This implicit representation is difficult to interpret or reuse across models and limits KD for API-only or costly-to-train models. This paper studies knowledge transfer for large language models (LLMs). We introduce Universal Textual Teaching (UTT), a parameter-update-free framework that distills observed Teacher-Student knowledge gaps into a textual, interpretable, and reusable natural-language artifact called Primer. Specifically, UTT first identifies representative gap cases through paired evaluations, and iteratively updates the Primer via multi-role interactions: the Student attempts each task, the Prompter turns evaluation feedback into a teaching instruction, the Teacher provides a targeted demonstration, and the Synthesizer consolidates validated lessons. Empirically, on the challenging math (Omni-MATH-2) and code generation (KernelBench) tasks, extensive results confirm the effectiveness of the method: UTT remarkably raises the Student's accuracy from 9.4% to 48.6% and Fast1 accuracy from 9% to 35% on KernelBench, while increasing mathematical reasoning accuracy from 27.6% to 51.7%. UTT also performs better than representative prompt engineering and parameter-based KD methods. Of note, UTT is shown to be generalizable across different Teachers and Students: a Primer synthesized for one Teacher-Student pair can generalize to other Students that do not participate in the synthesis.
Figures & tables
Figure 1: Left: We introduce U niversal T extual T eaching (UTT), a parameter-update-free framework that distills an LLM’s knowledge into a textual Primer through multi-role LLM interaction. Notably, the resulting Primer can improve the performance of the source Student and other Students that are not involved in creating the Primer . Right: For the target Student Qwen3.6-27B, Primers synthesized from the Teacher - Student pairs (shown at the bottom of the figure) consistently improve accuracy. Even when Qwen3.6-27B does not participate in the Primer synthesis, the absolute accuracy gains reach 41.8% and 31.5% (see the rightmost column) on KernelBench and Omni-MATH-2, respectively. These results suggest that text can serve as a universal medium for knowledge transfer among different LLMs for complex tasks.
Figure 2: Overview of Universal Textual Teaching (UTT). (1) Knowledge-gap construction. UTT first evaluates the knowledge gap between the Teacher and Student and accordingly partitions D into subsets (Section 3.2.1 ). (2) Primer synthesis. It then synthesizes a textual Primer through iterative interactions among multiple LLMs (Section 3.2.2 ). (3) Evaluation. Finally, the resulting Primer is applied to the source Student and transferred to other Students , improving both the source Student and the transfer Students . All model parameters remain frozen throughout the process. In the diagram, rectangles denote models, with the same color indicating the same model serving different roles; rounded rectangles denote data subsets, and diamonds denote evaluation procedures.
Primer synthesis
Target Student
Teacher
Teacher score
Source Student
Flash
Qwen
KernelBench Accuracy (%)
No Primer
9.4
9.2
Pro
19.6
Flash
48.6 (+39.2) ①
41.2 (+32.0) ③
Pro
19.6
Qwen
43.8 (+34.4) ④
50.0 (+40.8) ②
Opus
36.0
Flash
48.2 (+38.8) ⑤
51.0 (+41.8) ⑥
Table 1: Primer effectiveness and cross-model transferability. For each task, each Teacher and source Student pair produces one Primer , transferring to the target Student . ①–⑥ denote different experimental groups. Parentheses show gains over the corresponding Student baseline, and bold values indicate the best result. All scores are computed over the full task pool D .
Figure 3: Split-wise sample accuracy for KernelBench and Omni-MATH-2. Panels (a)–(c) show KernelBench results and panels (d)–(f) show Omni-MATH-2 results. From left to right, the three columns correspond to Pro → Flash, Pro → Qwen, and Opus → Flash, respectively.
Method
Omni-MATH-2
KernelBench
Flash
Qwen
Flash
Qwen
Acc.
Acc.
Acc.
Fast 1
Acc.
Fast 1
Student
27.6
34.3
9.4
9
9.2
12
Prompting Techniques (PT)
One-shot
35.0
57.7
39.2
23
45.8
31
Few-shot
35.4
54.8
40.6
25
46.6
30
Table 2: Comparison with prompt engineering methods. “ Teacher summary” means asking the Teacher to summarize the training set into a general prompt.
Method
Omni-MATH-2
KernelBench
Acc.
Acc.
Fast 1
Student
34.3
9.2
12
SeqKD ( Kim and Rush, 2016 )
53.5
18.8
18
Fine-tune-CoT ( Ho et al., 2023 )
39.5
41.8
28
RSR ( Yang et al., 2026 )
51.5
22.6
19
LUFFY ( Yan et al., 2025 )
55.2
11.4
13
Table 3: Comparison with KD methods under the Pro → Qwen setting.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Teacher → Student
Ddisttr
Ddistte
Dret
Dfrt
Total
KernelBench
Pro → Flash
135
55
100
210
500
Pro → Qwen
140
60
100
200
Opus → Flash
215
95
70
120
Omni-MATH-2
Pro → Flash
490
210
100
200
1000
Pro → Qwen
335
145
360
160
Opus → Flash
645
275
35
45
Appendix
Table 4: Numbers of generated samples associated with the knowledge-gap portions across tasks and Teacher– Student configurations.
Primer synthesis
Target Student
Teacher
Teacher score
Source Student
Flash 4.1
Qwen
No Primer
54.0
41.5
Max
94.0
Qwen
59.0 (+5.0)
76.0 (+34.5)
Max
94.0
Flash 4.1
56.0 (+2.0)
72.0 (+30.5)
Opus
93.5
Qwen
57.5 (+3.5)
80.5 (+39.0)
Appendix
Table 5: Primer effectiveness and cross-model transferability on Vision Dominant plane geometry problems from MathVerse. Each model is evaluated using two samples per problem. Teacher scores denote Teacher accuracy. Parentheses show percentage-point gains over the corresponding target Student baseline, and bold values indicate the best result for each target Student . Max, Qwen, Flash 4.1, and Opus denote Qwen3.8-Max, Qwen3.6-27B, DeepSeek-V4.1-Flash, and Claude Opus 5, respectively.
Figure 4: Effect of output-token budget on Omni-MATH-2 accuracy. Curves compare the unprimed Flash Student , the Pro Teacher, APE, GEPA, and UTT. UTT uses the same Primer synthesized under the 32k-token budget at every evaluated budget, without resynthesis.
Method / Role
Compute
Time (h)
Model calls
Input tokens
Output tokens
Parameter-updating KD
SeqKD
4 × A6000
1.2
–
–
–
Fine-tune-CoT
4 × A6000
3.5
–
–
–
RSR
4 × A6000
7.5
–
–
–
LUFFY
8 × A6000
22.0
–
–
–
UTT
Appendix
Table 6: Stage-specific resource usage for the Omni-MATH-2 results in Table 3 under the Pro → Qwen setting. We treat the task data as given and focus on the method-specific optimization stage: KD reports parameter training only, whereas UTT reports the time, model calls, and tokens used for API-based Primer synthesis. Because the two method families consume different resource types, these measurements characterize their respective resource profiles rather than constituting a strictly equivalent end-to-end cost comparison. Teacher side aggregates the Teacher, Prompter, and Synthesizer.
Variant
Batch
Ddisttr
Ddistte
Dret
Dfrt
Student without Primer
–
33.7
33.3
60.0
0.0
Full method
16
65.3
56.0
72.5
5.0
Random selection
16
55.6
50.0
72.5
3.8
Without Prompter
16
72.4
48.8
70.0
5.0
Without validation gate
16
59.2
39.3
80.0
3.8
Batch size 8
8
53.6
41.7
67.5
2.5
Appendix
Table 7: Core-component and batch-size ablations under Pro → Flash on Omni-MATH-2. Bold values indicate the best result for each subset.
Configuration
Accuracy (%)
Fast 1 (%)
Baseline
With Primer
Gain
Baseline
With Primer
Gain
Unquantized
9.2
41.2
+32.0
12
37
+25
Int4 RTN (G128)
4.4
34.6
+30.2
5
28
+23
Int4 RTN (Channelwise)
0.4
21.6
+21.2
0
16
+16
Thinking disabled
1.6
16.6
+15.0
1
7
+6
Appendix
Table 8: Primer gains under different quantization and thinking settings of Qwen3.6-27B on KernelBench.