Knowledge distillation is an established technique for improving the capabilities of small, efficient student models by training them with the representations of larger, more capable teacher models. Much of the recent work in the distillation of large language models (LLMs) has focused on distilling abilities learned during post-training, such as instruction following, chain-of-thought reasoning, and tool usage. This has left a large research gap in general knowledge distillation for LLMs, which is essential for developing efficient and private systems suitable for deployment on edge devices. We take a first-principles approach, evaluating previous lessons from prior works and conducting new explorations to develop a distillation methodology suitable for modern LLMs. We present KDFP, a novel methodology for white-box general knowledge distillation in LLMs. We demonstrate that KDFP outperforms existing methods by 1.6% − 4.9% across 9 benchmarks while increasing training efficiency by up to 99.1% through ephemeral parameter reduction.
Figures & tables
Figure 1: Diagram of KDFP. The attention, MLP, and residual outputs are distilled in each block, as well as embeddings and model outputs. A single projection matrix is learned during training to project student activations into the teacher’s embedding space.
Figure 2: Comparison of JS divergence with Forward and Reverse KL divergence matching a Gaussian to a trimodal Gaussian mixture model (GMM).
Feature
Traditional KD
LWD
TED
KDFP (Ours)
Internal Distill.
No
Yes
Yes
Yes
Projection
None
Up-Online
Up (Task Filters)
Up-Online
Projector Count
N/A
Single
Per-layer
Single
Objectives
CE, Div
CE, Div, MSE
CE, Div, MSE
CE, Div, MSE
Divergence
Forward KL
Forward KL
Forward KL
Param. Div. ( β=0.25 )
Table 1: Comparison of distillation methodologies
Method
ARC
BoolQ
DROP
MMLU
PIQA
SciQ
TQA
WG
Avg.
Baseline
0.8469
0.5771
0.0280
0.2361
0.7187
0.773
0.0760
0.5572
0.4236
FT
0.8170
0.6110
0.0425
0.2398
0.7144
0.785
0.0244
0.5706
0.4227
KD
0.8326
0.5936
0.0367
0.2319
0.7231
0.785
0.0388
0.5675
0.4232
LWD
0.8432
0.6235
0.0463
0.2316
0.7203
0.785
0.0636
0.5485
0.4291
TED
0.8063
0.6125
0.0247
0.2294
0.7040
0.765
0.0378
0.5588
0.4154
KDFP
0.8542
0.6291
0.0463
0.2395
0.7182
0.798
0.0762
0.5612
0.4358
Table 2: Main results. 100M tokens of data and 2.8B parameter teacher without finetuning.
Table 3: Analyses of KDFP. Table 3(a) shows finetuning the teacher prior to distillation. Table 3(b) shows scaling the teacher from 2.8B to 6.9B parameters. Table 3(c) shows scaling the training data to 500M tokens. Italics indicate an improvement over the Main Results (Table 2 ). Full results for these experiments are provided in Appendix B .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Method
ARC
BoolQ
DROP
MMLU
PIQA
SciQ
TQA
WG
Avg.
KD
0.8327
0.5936
0.0367
0.2319
0.7127
0.785
0.0388
0.5604
0.4232
LWD
0.8440
0.6180
0.0442
0.2322
0.7231
0.799
0.0622
0.5675
0.4303
TED
0.7818
0.6315
0.0320
0.2294
0.7057
0.764
0.0154
0.5667
0.4140
KDFP
0.8407
0.6205
0.0447
0.2391
0.7165
0.786
0.0643
0.5722
0.4316
Appendix
Table 4: Results from finetuning teacher prior to distillation. Baseline and FT results are the same as in the Main Results in Table 2 , as there is no distillation. Italics indicate improvement over distillation without a finetuned teacher.
Method
ARC
BoolQ
DROP
MMLU
PIQA
SciQ
TQA
WG
Avg.
KD
0.8327
0.5936
0.0367
0.2319
0.7231
0.785
0.0388
0.5675
0.4232
LWD
0.8453
0.6080
0.0373
0.2352
0.7165
0.790
0.0595
0.5635
0.4284
TED
0.8119
0.6058
0.0237
0.2313
0.7046
0.765
0.0485
0.5564
0.4165
KDFP
0.8483
0.6144
0.0449
0.2339
0.7133
0.791
0.0600
0.5667
0.4303
Appendix
Table 5: Full results for the analysis of Pythia 6.9B as the teacher in the experiments in Section 7.3 . Italics indicate an improvement over the 2.8B parameter teacher in the Main Results in Table 2 .
Method
ARC
BoolQ
DROP
MMLU
PIQA
SciQ
TQA
WG
Avg.
FT
0.8339
0.6275
0.0401
0.2407
0.7165
0.783
0.0322
0.5627
0.4263
KD
0.8136
0.6177
0.0321
0.2329
0.7155
0.774
0.0332
0.5675
0.4207
LWD
0.8313
0.6232
0.0391
0.2310
0.7155
0.790
0.0517
0.5572
0.4266
TED
0.8080
0.6168
0.0257
0.2293
0.7013
0.762
0.0352
0.5604
0.4154
KDFP
0.8326
0.6242
0.0416
0.2317
0.7209
0.776
0.0499
0.5604
0.4264
Appendix
Table 6: Full results for the data scaling analysis experiments in Section 7.4 . Italics indicate an improvement over the 100M token trainig budget in the Main Results in Table 2 .
Figure 3: Distillation of internal activations between a student and teacher transformer with an upward projection. The output of the residual stream after student Block i ( hS ) is passed through a projection matrix P . This maps the student’s smaller embedding dimension ( dS ) into the teacher’s latent space ( dT ) so that it can be directly compared against the target teacher activation ( hT ) from Block j .
Figure 4: The fully linear autoencoder architecture for an upward offline projector. The student activations ( x=hS ) serve as the input and are reconstructed as h^S via linear weights Wenc and Wdec . The intermediate latent representation ( z^=h^T ) is explicitly trained to match the teacher activations ( z=hT ). The total objective is a convex combination of the reconstruction loss ( Lrecon ) and the latent alignment loss ( Lalign ).
Figure 5: Shallow student architectures used to be a common choice in KD, as it is maximally compatible with the teacher. However, shallow students require full pretraining, making them infeasible in most modern settings. Selecting a pretrained model for student initialization is the current best practice for continual pretraining.
Experiment
Accuracy
pretrained-ft
0.4146
pretrained-no_train
0.4164
pretrained-kd-down_offline
0.4167
pretrained-kd-up_offline
0.4167
pretrained-kd-down_randomized
0.4167
pretrained-kd-up_randomized
0.4167
Appendix
Table 7: Stage 1: Results of the basic exploration
Accuracy
β
0-shot
5-shot
0.0
0.4150
0.4132
0.25
0.4151
0.4133
0.5
0.4148
0.4128
0.75
0.4156
0.4128
1.0
0.4147
0.4134
Appendix
Table 8: Accuracy on HellaSwag across different divergence parameter values for 0-shot and 5-shot evaluations. Best results are bolded, and the lowest are underlined.
Figure 6: Internal architecture of a transformer block. Each selected hook point acts as a checkpoint on model operations in the model’s embedding space. attn_out checks the mixing across the sequence, mlp_out checks the transformation, and resid_post checks the full block operation.
Objective Components
Projection
Accuracy
Online
0.4160
Balanced
Offline
0.4150
Randomized
0.4150
Online
0.4151
Only Outputs
Offline
0.4151
Randomized
0.4151
Appendix
Table 9: Stage 4a results, comparing different projections and objective components.
Objective Components
Accuracy
No Resid
0.4154
No Attn Out
0.4145
No Embed
0.4113
No MLP
0.4040
Appendix
Table 10: Stage 4B subtractive results, removing specific objective components from original Balanced objective.
Table 14: Final Hyperparameter Settings for Model Training. FT indicates the parameters used for finetuning, and KD indicates the parameters used for all distillation methods.
Hyperparameter
Value
Learning Rate
4.5×10−4
Weight Decay
1.3×10−2
Tied Weights
True
Appendix
Table 15: Final Hyperparameters for Offline Projectors
Hyperparameter
Min
Max
LR (filters)
1e-4
1e-2
LR (distillation)
1e-6
1e-4
Appendix
Table 16: The search space for the hyperparameters in TED. Just as in the rest of our tuning, we searched this space over 25 trials.
Hyperparameter
Value
LR (filters)
3.9e-4
LR (distillation)
3.66e-5
Appendix
Table 17: The final hyperparameter settings used for TED training.
Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD). This recovery step largely decides the final quality, yet it is expensive. We present a practitioner's study of how to make distillation training efficient, organised around two systems contributions. First, we show that offline KD (caching the teacher's top-K logits once and training the student against the cache) matches online distillation at near-identical training loss while removing the teacher from memory, running about 29% faster per iteration, and reaching up to 41% higher throughput on a single H200 GPU. Second, we introduce a \emph{fused, chunked KL loss} that never materialises the full vocabulary-sized logit tensor, making peak memory linear in the sequence length. This removes the memory spike that otherwise caps context length and lets us train at four times the context (32{,}768 tokens) on a single GPU. A separate output-head-only toy benchmark isolates the loss kernel and confirms its memory and iteration-rate scaling from 4K to 256K tokens. Together these make large-scale healing and hundreds of ablations affordable. We also report supporting ablations on loss design and sequence packing. We release our chunked-loss implementation: https://github.com/CompactifAI/Full-Chunked-KL-Loss.
Bakbergen Ryskulov, Iker García-Ferrero, David Montero +5
Knowledge distillation (KD) is widely used to compress and post-train large language models (LLMs), yet many existing frameworks execute teacher inference with the same training-oriented backend as student optimization, leading to suboptimal efficiency. In this paper, we propose KDFlow, a novel framework for LLM distillation that features a decoupled architecture and employs SGLang for teacher inference. KDFlow combines SGLang for teacher inference with PyTorch FSDP2 for student optimization, allowing each model to run on a backend tailored to its workload. To enable efficient full-vocabulary distillation in this decoupled architecture, KDFlow transfers the teacher's final hidden states via Ray's object store and recomputes teacher logits on each student worker using a frozen copy of the teacher's output head. Furthermore, our framework supports both off-policy and on-policy distillation and incorporates cross-tokenizer algorithms through highly extensible and user-friendly APIs. Experiments show that KDFlow achieves a 1.44× to 6.36× speedup over MS-SWIFT in off-policy distillation and a 1.43× to 1.75× speedup over verl in on-policy distillation. KDFlow further scales to 64 GPUs, achieving 3.68× and 2.52× strong-scaling speedups in two representative model configurations. The code and documentation are publicly available.
Songming Zhang, Xue Zhang, Tong Zhang +3
School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China · Key Laboratory of Big Data & Artificial Intelligence in Transportation, (Beijing Jiaotong University), Ministry of Education · Weixin AI
Large language models (LLMs) achieve strong performance across many tasks, but their high computational cost limits deployment in resource-constrained environments. Knowledge Distillation (KD) offers a practical solution by transferring knowledge from a teacher model of a larger size to a smaller student model. While prior work has mainly examined task-specific or small-scale settings, the post-training stage for building general instruction-following models has received limited attention. In this paper, we conduct a systematic study of KD in post-training using the large-scale Tulu 3 dataset. We find that KD outperforms supervised fine-tuning (SFT) in low-data regimes, but its advantage diminishes as more training data is added. Distilling from a stronger instruction-tuned teacher restores substantial gains even with abundant data, indicating that KD remains effective when the teacher provides knowledge that the student cannot easily acquire from the training data alone. We further study domain-specific, low-resource scenarios and propose a two-stage KD strategy that leverages synthetic teacher-labeled data followed by refinement on human annotations. This method consistently improves student performance, providing practical guidance for building compact models in data-scarce environments.
Xin Liu, Simin Ma, Shujian Liu +5
University of Michigan · 2Zoom Video Communications