Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time. Knowledge distillation (KD) offers a practical solution by transferring knowledge from a large teacher model to a smaller student model via alignment of discrete probability distributions. However, existing KD methods for LLMs primarily rely on divergences that evaluate discrepancies through probability values at each vocabulary index, without explicitly leveraging token-level semantic information. We propose Wasserstein-based knowledge distillation (WASD) for LLMs, which incorporates token-level semantic information via the Wasserstein-based distance with a cost matrix derived from token embeddings. To ensure computational tractability, we adopt the Sinkhorn divergence and derive a gradient-equivalent objective that can be efficiently optimized without introducing additional networks. Experiments across multiple LLM families and scales show that WASD consistently improves distillation performance on diverse tasks, including instruction following, mathematical reasoning, and code generation. Our results highlight the importance of semantic information encoded in the token space for effective distribution alignment in LLM distillation. The implementation is publicly available at https://github.com/aailab-kaist/WASD .
Figures & tables
Figure 1 : Motivation for Wasserstein-based distillation (WASD). (a) Loss contours on the probability simplex for a toy next-token prediction task (“ What do you sit on in the living room? ”) with candidates {sofa, couch, apple} , where we visualize divergence between fixed teacher and varying student distributions. Unlike other divergences, Sinkhorn divergence used in WASD reflects semantic relationships between tokens. (b) Quality-diversity trade-off (ROUGE-L vs. Self-BLEU) on five instruction-following benchmarks for GPT-2 (1.5B → 0.1B), varying the decoding temperature.
Figure 2 : Loss contour of entropy-regularized Wasserstein distance.
Figure 3
Method
Dolly Eval
Self Inst
Vicuna
Super NI
UnNI
Avg. ( ↑ )
GPT-2 XL (Teacher)
26.93 ± 0.60
14.77 ± 0.38
16.43 ± 0.28
27.19 ± 0.47
31.43 ± 0.12
23.35
KL [ 25 ]
23.33 ± 0.36
10.43 ± 0.23
15.34 ± 0.27
17.42 ± 0.14
19.34 ± 0.07
17.17
RKL [ 24 ]
24.62 ± 0.27
11.31 ± 0.28
15.97 ± 0.28
21.42 ± 0.30
23.82 ± 0.14
19.43
Sym-KL
23.32 ± 0.37
10.81 ± 0.27
15.41 ± 0.56
20.12 ± 0.23
21.23 ± 0.09
18.18
Jeffrey
22.84 ± 0.53
10.84 ± 0.24
15.30 ± 0.22
20.57 ± 0.37
22.66 ± 0.13
18.44
TV [ 64 ]
24.08 ± 0.41
10.83 ± 0.25
14.68 ± 0.43
25.57 ± 0.22
27.51 ± 0.09
20.53
Table 2 : ROUGE-L scores on five general instruction-following benchmarks for GPT-2 XL (1.5B) → GPT-2 Base (0.1B) , varying divergences. All results are obtained using our implementation, where all methods share identical teacher and student initialization. Performance is averaged over five evaluation seeds. Bold and underline indicate the best and second-best performance, respectively.
Method
Dolly Eval
Self Inst
Vicuna
Super NI
UnNI
Avg. ( ↑ )
GPT-2 XL (Teacher)
26.93 ± 0.60
14.77 ± 0.38
16.43 ± 0.28
27.19 ± 0.47
31.43 ± 0.12
23.35
GPT-2 XL (1.5B) → GPT-2 Base (0.1B)
GKD [ 1 ]
23.86 ± 0.22
11.21 ± 0.42
14.63 ± 0.19
19.43 ± 0.17
22.28 ± 0.12
18.28
TAID [ 54 ]
25.44 ± 0.42
12.55 ± 0.11
16.65 ± 0.32
24.45 ± 0.20
26.70 ± 0.09
21.16
DistiLLM (SKL) [ 35 ]
25.39 ± 0.43
12.42 ± 0.25
16.21 ± 0.33
25.13 ± 0.22
26.94 ± 0.18
21.22
DistiLLM (SRKL) [ 35 ]
25.10 ± 0.30
12.01 ± 0.26
17.40 ± 0.28
25.39 ± 0.21
26.40 ± 0.10
21.26
Table 3 : ROUGE-L scores on five general instruction-following benchmarks for GPT-2 and OpenLLaMA2 . Bold and underline indicate the best and second-best performance, respectively.
Figure 3 : ROUGE-L during training on Dolly Eval.
Table 7Figure 8Table 9
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Loss contours of various divergences on the probability simplex for a toy next-token prediction task with different teacher distributions, extending the visualization in Fig. 1(a) .
Figure 6 : Loss contours of Sinkhorn divergence on the probability simplex under different entropy regularization hyperparameters ϵ , where the value of ϵ is indicated in parentheses in each subplot title.
Figure 12
Method
Avg. ROUGE-L ( ↑ )
GPT-2 XL (Teacher)
23.35
CSD [ 32 ]
21.46
WASD
Default settings
22.47
Cost change (cos → L2)
22.45
Increased truncation k (8 → 16)
22.45
Appendix
Table 9 : Ablation study of WASD on five general instruction-following benchmarks for GPT-2 XL (1.5B) → GPT-2 Base (0.1B) under the same experimental setup as Table 2 .
Translation
Summarization
Arithmetic
Method
COMET ( ↑ )
ROUGE-L ( ↑ )
Accuracy ( ↑ )
CSD [ 32 ]
73.70 ± 0.08
34.75 ± 0.30
23.38 ± 0.57
AMiD [ 53 ]
73.78 ± 0.14
34.65 ± 0.39
24.31 ± 0.96
WASD (ours)
74.40 ± 0.31
34.96 ± 0.18
24.62 ± 0.25
Appendix
Table 10 : Repeated task-specific distillation experiments under the same settings as Section 4.2 . Results are reported as mean ± standard deviation over three training seeds.
Method
Dolly Eval
Self Inst
Vicuna
Super NI
UnNI
Average
DSKDv2 (FKL)
23.58
12.11
14.26
26.64
25.76
20.47
+ WASD-S
24.68
12.55
16.09
25.09
26.34
20.95
+ WASD-S+T
24.80
12.74
16.33
26.32
26.39
21.31
Appendix
Table 11 : Cross-family distillation from Qwen1.5-1.8B to GPT-2-0.1B within the DSKDv2 framework [ 69 ] . We report ROUGE-L with greedy decoding after three training epochs.
Distinct-2
Semantic coherence ( ↑ )
Temperature T
AMiD
WASD
AMiD
WASD
0.7
0.262
0.270
0.879
0.890
1.0
0.309
0.312
0.850
0.866
1.3
0.357
0.359
0.823
0.844
1.5
0.390
0.388
0.807
0.827
Appendix
Table 12 : Diversity and semantic coherence on Dolly Eval at different decoding temperatures.
Model
Top next tokens and probabilities
Teacher
speak (0.19), express (0.19), share (0.19), get (0.19), voice (0.03)
WASD
get (0.36), express (0.25), convince (0.16), share (0.07), finish (0.03)
AMiD
get (0.63), express (0.15), convince (0.14), share (0.02), make (0.01)
Appendix
Table 13 : Top next-token probabilities at a shared DialogSum context with the prefix “Sarah is upset because she couldn’t”.
Knowledge distillation (KD) is widely used to compress and post-train large language models (LLMs), yet many existing frameworks execute teacher inference with the same training-oriented backend as student optimization, leading to suboptimal efficiency. In this paper, we propose KDFlow, a novel framework for LLM distillation that features a decoupled architecture and employs SGLang for teacher inference. KDFlow combines SGLang for teacher inference with PyTorch FSDP2 for student optimization, allowing each model to run on a backend tailored to its workload. To enable efficient full-vocabulary distillation in this decoupled architecture, KDFlow transfers the teacher's final hidden states via Ray's object store and recomputes teacher logits on each student worker using a frozen copy of the teacher's output head. Furthermore, our framework supports both off-policy and on-policy distillation and incorporates cross-tokenizer algorithms through highly extensible and user-friendly APIs. Experiments show that KDFlow achieves a 1.44× to 6.36× speedup over MS-SWIFT in off-policy distillation and a 1.43× to 1.75× speedup over verl in on-policy distillation. KDFlow further scales to 64 GPUs, achieving 3.68× and 2.52× strong-scaling speedups in two representative model configurations. The code and documentation are publicly available.
Songming Zhang, Xue Zhang, Tong Zhang +3
School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China · Key Laboratory of Big Data & Artificial Intelligence in Transportation, (Beijing Jiaotong University), Ministry of Education · Weixin AI
Knowledge distillation is an established technique for improving the capabilities of small, efficient student models by training them with the representations of larger, more capable teacher models. Much of the recent work in the distillation of large language models (LLMs) has focused on distilling abilities learned during post-training, such as instruction following, chain-of-thought reasoning, and tool usage. This has left a large research gap in general knowledge distillation for LLMs, which is essential for developing efficient and private systems suitable for deployment on edge devices. We take a first-principles approach, evaluating previous lessons from prior works and conducting new explorations to develop a distillation methodology suitable for modern LLMs. We present KDFP, a novel methodology for white-box general knowledge distillation in LLMs. We demonstrate that KDFP outperforms existing methods by 1.6% − 4.9% across 9 benchmarks while increasing training efficiency by up to 99.1% through ephemeral parameter reduction.
Ryan Swift, Konstantinos Psounis
Thomas Lord Department of Computer Science University of Southern California Los Angeles, CA 90007, USA
Large language models (LLMs) have achieved remarkable performance across diverse domains, yet their enormous computational and memory requirements hinder deployment in resource-constrained environments. Knowledge distillation offers a promising solution by transferring knowledge from a large teacher model to a smaller student model. However, existing distillation methods typically treat all tokens equally, ignoring the fact that different tokens contribute unequally to model decisions. This can lead to inefficient knowledge transfer and reduced learning effectiveness. To address this limitation, we propose an entropy-based adaptive distillation strategy that dynamically adjusts the training process at the token level. Our method leverages the teacher's output entropy to guide three aspects of distillation. Specifically, we introduce a token-level curriculum by dynamically shifting focus from low- to high-entropy tokens during training. We further adjust the distillation temperature based on token entropy to better capture teacher confidence patterns. Moreover, we employ a dual-branch architecture for efficient logits-only distillation on easy tokens and deeper feature-based distillation on difficult tokens. Extensive experiments validate the soundness and effectiveness of our method.
Hao Zhang, Zhibin Zhang, Guangxin Wu +3
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences