cs.LGOct 7, 2026

KDFP: A first-principles approach to knowledge distillation in large language models

Authors: Ryan Swift, Konstantinos Psounis

Organizations: Thomas Lord Department of Computer Science University of Southern California Los Angeles, CA 90007, USA

Abstract

Knowledge distillation is an established technique for improving the capabilities of small, efficient student models by training them with the representations of larger, more capable teacher models. Much of the recent work in the distillation of large language models (LLMs) has focused on distilling abilities learned during post-training, such as instruction following, chain-of-thought reasoning, and tool usage. This has left a large research gap in general knowledge distillation for LLMs, which is essential for developing efficient and private systems suitable for deployment on edge devices. We take a first-principles approach, evaluating previous lessons from prior works and conducting new explorations to develop a distillation methodology suitable for modern LLMs. We present KDFP, a novel methodology for white-box general knowledge distillation in LLMs. We demonstrate that KDFP outperforms existing methods by 1.6% −- 4.9% across 9 benchmarks while increasing training efficiency by up to 99.1% through ephemeral parameter reduction.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models

    Mar 2, 2026Songming Zhang, Xue Zhang, Tong Zhang +3Cross-Tokenizer Knowledge DistillationLanguage Model Distillation

  2. Understanding Knowledge Distillation in Post-Training: When It Helps and When It Fails

    Jun 22, 2026Xin Liu, Simin Ma, Shujian Liu +5Data-Free Knowledge DistillationLanguage Model Distillation