Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time. Knowledge distillation (KD) offers a practical solution by transferring knowledge from a large teacher model to a smaller student model via alignment of discrete probability distributions. However, existing KD methods for LLMs primarily rely on divergences that evaluate discrepancies through probability values at each vocabulary index, without explicitly leveraging token-level semantic information. We propose Wasserstein-based knowledge distillation (WASD) for LLMs, which incorporates token-level semantic information via the Wasserstein-based distance with a cost matrix derived from token embeddings. To ensure computational tractability, we adopt the Sinkhorn divergence and derive a gradient-equivalent objective that can be efficiently optimized without introducing additional networks. Experiments across multiple LLM families and scales show that WASD consistently improves distillation performance on diverse tasks, including instruction following, mathematical reasoning, and code generation. Our results highlight the importance of semantic information encoded in the token space for effective distribution alignment in LLM distillation. The implementation is publicly available at https://github.com/aailab-kaist/WASD .
Figures & tables
Figure 1 : Motivation for Wasserstein-based distillation (WASD). (a) Loss contours on the probability simplex for a toy next-token prediction task (“ What do you sit on in the living room? ”) with candidates {sofa, couch, apple} , where we visualize divergence between fixed teacher and varying student distributions. Unlike other divergences, Sinkhorn divergence used in WASD reflects semantic relationships between tokens. (b) Quality-diversity trade-off (ROUGE-L vs. Self-BLEU) on five instruction-following benchmarks for GPT-2 (1.5B → 0.1B), varying the decoding temperature.
Figure 2 : Loss contour of entropy-regularized Wasserstein distance.
Figure 3
Method
Dolly Eval
Self Inst
Vicuna
Super NI
UnNI
Avg. ( ↑ )
GPT-2 XL (Teacher)
26.93 ± 0.60
14.77 ± 0.38
16.43 ± 0.28
27.19 ± 0.47
31.43 ± 0.12
23.35
KL [ 25 ]
23.33 ± 0.36
10.43 ± 0.23
15.34 ± 0.27
17.42 ± 0.14
19.34 ± 0.07
17.17
RKL [ 24 ]
24.62 ± 0.27
11.31 ± 0.28
15.97 ± 0.28
21.42 ± 0.30
23.82 ± 0.14
19.43
Sym-KL
23.32 ± 0.37
10.81 ± 0.27
15.41 ± 0.56
20.12 ± 0.23
21.23 ± 0.09
18.18
Jeffrey
22.84 ± 0.53
10.84 ± 0.24
15.30 ± 0.22
20.57 ± 0.37
22.66 ± 0.13
18.44
TV [ 64 ]
24.08 ± 0.41
10.83 ± 0.25
14.68 ± 0.43
25.57 ± 0.22
27.51 ± 0.09
20.53
Table 2 : ROUGE-L scores on five general instruction-following benchmarks for GPT-2 XL (1.5B) → GPT-2 Base (0.1B) , varying divergences. All results are obtained using our implementation, where all methods share identical teacher and student initialization. Performance is averaged over five evaluation seeds. Bold and underline indicate the best and second-best performance, respectively.
Method
Dolly Eval
Self Inst
Vicuna
Super NI
UnNI
Avg. ( ↑ )
GPT-2 XL (Teacher)
26.93 ± 0.60
14.77 ± 0.38
16.43 ± 0.28
27.19 ± 0.47
31.43 ± 0.12
23.35
GPT-2 XL (1.5B) → GPT-2 Base (0.1B)
GKD [ 1 ]
23.86 ± 0.22
11.21 ± 0.42
14.63 ± 0.19
19.43 ± 0.17
22.28 ± 0.12
18.28
TAID [ 54 ]
25.44 ± 0.42
12.55 ± 0.11
16.65 ± 0.32
24.45 ± 0.20
26.70 ± 0.09
21.16
DistiLLM (SKL) [ 35 ]
25.39 ± 0.43
12.42 ± 0.25
16.21 ± 0.33
25.13 ± 0.22
26.94 ± 0.18
21.22
DistiLLM (SRKL) [ 35 ]
25.10 ± 0.30
12.01 ± 0.26
17.40 ± 0.28
25.39 ± 0.21
26.40 ± 0.10
21.26
Table 3 : ROUGE-L scores on five general instruction-following benchmarks for GPT-2 and OpenLLaMA2 . Bold and underline indicate the best and second-best performance, respectively.
Figure 3 : ROUGE-L during training on Dolly Eval.
Table 7Figure 8Table 9
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Loss contours of various divergences on the probability simplex for a toy next-token prediction task with different teacher distributions, extending the visualization in Fig. 1(a) .
Figure 6 : Loss contours of Sinkhorn divergence on the probability simplex under different entropy regularization hyperparameters ϵ , where the value of ϵ is indicated in parentheses in each subplot title.
Figure 12
Method
Avg. ROUGE-L ( ↑ )
GPT-2 XL (Teacher)
23.35
CSD [ 32 ]
21.46
WASD
Default settings
22.47
Cost change (cos → L2)
22.45
Increased truncation k (8 → 16)
22.45
Appendix
Table 9 : Ablation study of WASD on five general instruction-following benchmarks for GPT-2 XL (1.5B) → GPT-2 Base (0.1B) under the same experimental setup as Table 2 .
Translation
Summarization
Arithmetic
Method
COMET ( ↑ )
ROUGE-L ( ↑ )
Accuracy ( ↑ )
CSD [ 32 ]
73.70 ± 0.08
34.75 ± 0.30
23.38 ± 0.57
AMiD [ 53 ]
73.78 ± 0.14
34.65 ± 0.39
24.31 ± 0.96
WASD (ours)
74.40 ± 0.31
34.96 ± 0.18
24.62 ± 0.25
Appendix
Table 10 : Repeated task-specific distillation experiments under the same settings as Section 4.2 . Results are reported as mean ± standard deviation over three training seeds.
Method
Dolly Eval
Self Inst
Vicuna
Super NI
UnNI
Average
DSKDv2 (FKL)
23.58
12.11
14.26
26.64
25.76
20.47
+ WASD-S
24.68
12.55
16.09
25.09
26.34
20.95
+ WASD-S+T
24.80
12.74
16.33
26.32
26.39
21.31
Appendix
Table 11 : Cross-family distillation from Qwen1.5-1.8B to GPT-2-0.1B within the DSKDv2 framework [ 69 ] . We report ROUGE-L with greedy decoding after three training epochs.
Distinct-2
Semantic coherence ( ↑ )
Temperature T
AMiD
WASD
AMiD
WASD
0.7
0.262
0.270
0.879
0.890
1.0
0.309
0.312
0.850
0.866
1.3
0.357
0.359
0.823
0.844
1.5
0.390
0.388
0.807
0.827
Appendix
Table 12 : Diversity and semantic coherence on Dolly Eval at different decoding temperatures.
Model
Top next tokens and probabilities
Teacher
speak (0.19), express (0.19), share (0.19), get (0.19), voice (0.03)
WASD
get (0.36), express (0.25), convince (0.16), share (0.07), finish (0.03)
AMiD
get (0.63), express (0.15), convince (0.14), share (0.02), make (0.01)
Appendix
Table 13 : Top next-token probabilities at a shared DialogSum context with the prefix “Sarah is upset because she couldn’t”.
School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China · Key Laboratory of Big Data & Artificial Intelligence in Transportation, (Beijing Jiaotong University), Ministry of Education · Weixin AI
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences