Knowledge distillation is an established technique for improving the capabilities of small, efficient student models by training them with the representations of larger, more capable teacher models. Much of the recent work in the distillation of large language models (LLMs) has focused on distilling abilities learned during post-training, such as instruction following, chain-of-thought reasoning, and tool usage. This has left a large research gap in general knowledge distillation for LLMs, which is essential for developing efficient and private systems suitable for deployment on edge devices. We take a first-principles approach, evaluating previous lessons from prior works and conducting new explorations to develop a distillation methodology suitable for modern LLMs. We present KDFP, a novel methodology for white-box general knowledge distillation in LLMs. We demonstrate that KDFP outperforms existing methods by 1.6% − 4.9% across 9 benchmarks while increasing training efficiency by up to 99.1% through ephemeral parameter reduction.
Figures & tables
Figure 1: Diagram of KDFP. The attention, MLP, and residual outputs are distilled in each block, as well as embeddings and model outputs. A single projection matrix is learned during training to project student activations into the teacher’s embedding space.
Figure 2: Comparison of JS divergence with Forward and Reverse KL divergence matching a Gaussian to a trimodal Gaussian mixture model (GMM).
Feature
Traditional KD
LWD
TED
KDFP (Ours)
Internal Distill.
No
Yes
Yes
Yes
Projection
None
Up-Online
Up (Task Filters)
Up-Online
Projector Count
N/A
Single
Per-layer
Single
Objectives
CE, Div
CE, Div, MSE
CE, Div, MSE
CE, Div, MSE
Divergence
Forward KL
Forward KL
Forward KL
Param. Div. ( β=0.25 )
Table 1: Comparison of distillation methodologies
Method
ARC
BoolQ
DROP
MMLU
PIQA
SciQ
TQA
WG
Avg.
Baseline
0.8469
0.5771
0.0280
0.2361
0.7187
0.773
0.0760
0.5572
0.4236
FT
0.8170
0.6110
0.0425
0.2398
0.7144
0.785
0.0244
0.5706
0.4227
KD
0.8326
0.5936
0.0367
0.2319
0.7231
0.785
0.0388
0.5675
0.4232
LWD
0.8432
0.6235
0.0463
0.2316
0.7203
0.785
0.0636
0.5485
0.4291
TED
0.8063
0.6125
0.0247
0.2294
0.7040
0.765
0.0378
0.5588
0.4154
KDFP
0.8542
0.6291
0.0463
0.2395
0.7182
0.798
0.0762
0.5612
0.4358
Table 2: Main results. 100M tokens of data and 2.8B parameter teacher without finetuning.
Table 3: Analyses of KDFP. Table 3(a) shows finetuning the teacher prior to distillation. Table 3(b) shows scaling the teacher from 2.8B to 6.9B parameters. Table 3(c) shows scaling the training data to 500M tokens. Italics indicate an improvement over the Main Results (Table 2 ). Full results for these experiments are provided in Appendix B .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Method
ARC
BoolQ
DROP
MMLU
PIQA
SciQ
TQA
WG
Avg.
KD
0.8327
0.5936
0.0367
0.2319
0.7127
0.785
0.0388
0.5604
0.4232
LWD
0.8440
0.6180
0.0442
0.2322
0.7231
0.799
0.0622
0.5675
0.4303
TED
0.7818
0.6315
0.0320
0.2294
0.7057
0.764
0.0154
0.5667
0.4140
KDFP
0.8407
0.6205
0.0447
0.2391
0.7165
0.786
0.0643
0.5722
0.4316
Appendix
Table 4: Results from finetuning teacher prior to distillation. Baseline and FT results are the same as in the Main Results in Table 2 , as there is no distillation. Italics indicate improvement over distillation without a finetuned teacher.
Method
ARC
BoolQ
DROP
MMLU
PIQA
SciQ
TQA
WG
Avg.
KD
0.8327
0.5936
0.0367
0.2319
0.7231
0.785
0.0388
0.5675
0.4232
LWD
0.8453
0.6080
0.0373
0.2352
0.7165
0.790
0.0595
0.5635
0.4284
TED
0.8119
0.6058
0.0237
0.2313
0.7046
0.765
0.0485
0.5564
0.4165
KDFP
0.8483
0.6144
0.0449
0.2339
0.7133
0.791
0.0600
0.5667
0.4303
Appendix
Table 5: Full results for the analysis of Pythia 6.9B as the teacher in the experiments in Section 7.3 . Italics indicate an improvement over the 2.8B parameter teacher in the Main Results in Table 2 .
Method
ARC
BoolQ
DROP
MMLU
PIQA
SciQ
TQA
WG
Avg.
FT
0.8339
0.6275
0.0401
0.2407
0.7165
0.783
0.0322
0.5627
0.4263
KD
0.8136
0.6177
0.0321
0.2329
0.7155
0.774
0.0332
0.5675
0.4207
LWD
0.8313
0.6232
0.0391
0.2310
0.7155
0.790
0.0517
0.5572
0.4266
TED
0.8080
0.6168
0.0257
0.2293
0.7013
0.762
0.0352
0.5604
0.4154
KDFP
0.8326
0.6242
0.0416
0.2317
0.7209
0.776
0.0499
0.5604
0.4264
Appendix
Table 6: Full results for the data scaling analysis experiments in Section 7.4 . Italics indicate an improvement over the 100M token trainig budget in the Main Results in Table 2 .
Figure 3: Distillation of internal activations between a student and teacher transformer with an upward projection. The output of the residual stream after student Block i ( hS ) is passed through a projection matrix P . This maps the student’s smaller embedding dimension ( dS ) into the teacher’s latent space ( dT ) so that it can be directly compared against the target teacher activation ( hT ) from Block j .
Figure 4: The fully linear autoencoder architecture for an upward offline projector. The student activations ( x=hS ) serve as the input and are reconstructed as h^S via linear weights Wenc and Wdec . The intermediate latent representation ( z^=h^T ) is explicitly trained to match the teacher activations ( z=hT ). The total objective is a convex combination of the reconstruction loss ( Lrecon ) and the latent alignment loss ( Lalign ).
Figure 5: Shallow student architectures used to be a common choice in KD, as it is maximally compatible with the teacher. However, shallow students require full pretraining, making them infeasible in most modern settings. Selecting a pretrained model for student initialization is the current best practice for continual pretraining.
Experiment
Accuracy
pretrained-ft
0.4146
pretrained-no_train
0.4164
pretrained-kd-down_offline
0.4167
pretrained-kd-up_offline
0.4167
pretrained-kd-down_randomized
0.4167
pretrained-kd-up_randomized
0.4167
Appendix
Table 7: Stage 1: Results of the basic exploration
Accuracy
β
0-shot
5-shot
0.0
0.4150
0.4132
0.25
0.4151
0.4133
0.5
0.4148
0.4128
0.75
0.4156
0.4128
1.0
0.4147
0.4134
Appendix
Table 8: Accuracy on HellaSwag across different divergence parameter values for 0-shot and 5-shot evaluations. Best results are bolded, and the lowest are underlined.
Figure 6: Internal architecture of a transformer block. Each selected hook point acts as a checkpoint on model operations in the model’s embedding space. attn_out checks the mixing across the sequence, mlp_out checks the transformation, and resid_post checks the full block operation.
Objective Components
Projection
Accuracy
Online
0.4160
Balanced
Offline
0.4150
Randomized
0.4150
Online
0.4151
Only Outputs
Offline
0.4151
Randomized
0.4151
Appendix
Table 9: Stage 4a results, comparing different projections and objective components.
Objective Components
Accuracy
No Resid
0.4154
No Attn Out
0.4145
No Embed
0.4113
No MLP
0.4040
Appendix
Table 10: Stage 4B subtractive results, removing specific objective components from original Balanced objective.
Table 14: Final Hyperparameter Settings for Model Training. FT indicates the parameters used for finetuning, and KD indicates the parameters used for all distillation methods.
Hyperparameter
Value
Learning Rate
4.5×10−4
Weight Decay
1.3×10−2
Tied Weights
True
Appendix
Table 15: Final Hyperparameters for Offline Projectors
Hyperparameter
Min
Max
LR (filters)
1e-4
1e-2
LR (distillation)
1e-6
1e-4
Appendix
Table 16: The search space for the hyperparameters in TED. Just as in the rest of our tuning, we searched this space over 25 trials.
Hyperparameter
Value
LR (filters)
3.9e-4
LR (distillation)
3.66e-5
Appendix
Table 17: The final hyperparameter settings used for TED training.
School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China · Key Laboratory of Big Data & Artificial Intelligence in Transportation, (Beijing Jiaotong University), Ministry of Education · Weixin AI