A Statistical Perspective on Knowledge Distillation: Foundations, Classical Methods, and Large Language Model Extensions
Authors: Luyang Fang, Haoran Lu, Jiazhang Cai, Tao Wang, Huimin Cheng, Wenxuan Zhong, Ping Ma
Organizations: Department of Statistics, University of Georgia, Athens, Georgia 30602, USA · Department of Biostatistics, Boston University, Boston, Massachusetts 02118, USA
Knowledge Distillation (KD) has emerged as a vital paradigm for transferring the capabilities of high-capacity models to efficient ``student'' counterparts, addressing critical challenges in computational cost, deployment constraints, and privacy-sensitive settings. Although KD is widely used in practice, it is often viewed primarily as an engineering technique, with a unified statistical perspective remaining less developed. This review bridges that gap by presenting a unified Bayesian formulation of KD that formulates teacher predictions as prior information. This provides a principled interpretation of how teacher information is incorporated into student learning and establishes a rigorous connection to uncertainty quantification. We demonstrate how this foundational lens reconciles classical distillation with modern extensions in generative and foundation-model systems, showing that contemporary developments remain rooted in these same statistical principles. By synthesizing theory with emerging methodologies and diverse applications, this review provides a conceptual roadmap and identifies critical open problems for the future of the field.
Figures & tables
Fig 1: Overview of knowledge distillation, Bayesian knowledge distillation, and representative KD methods. (a) In classical KD, the student learns from both hard labels and teacher-derived soft targets. (b) In Bayesian KD, the teacher’s predictive distribution is incorporated as prior information and combined with the likelihood from observed labels, yielding a posterior over student parameters and associated predictive uncertainty. (c) Representative KD methods, including multi-teacher, sequential, and self-distillation. (d) Representative KD methods for LLMs, including token-level, sequence-level, and rationale distillation.
Fig 2: The top panel shows three images with the lowest mean deviance, while the bottom panel shows three images with the highest mean deviance, for (a) MNIST, (b) Fashion MNIST, and (c) CIFAR-10 datasets. Reproduced from Fang et al. (2024) , licensed under CC BY 4.0.
Evaluation type
Example metrics
Statistical interpretation
Distributional
KL divergence, JS divergence
Fidelity of predictive distributions
Decision-level
Accuracy, F1 score
Quality of induced decisions
Uncertainty
ECE, reliability diagram
Calibration of predictive probabilities
Table 1: Evaluation metrics for knowledge distillation from a statistical perspective.
Fig 3: Multi-teacher Bayesian knowledge distillation for protein subcellular localization. (a) Ten eukaryotic subcellular compartments in the localization task (adapted from iStock.com/mariaflaya, used under license). (b) Immunofluorescence microscopy images from the Human Protein Atlas ( Thul et al., 2017 ) (CC BY-SA 3.0) illustrating two proteins for which MT-BKD produces contrasting uncertainty levels. In each image, green fluorescence marks the target protein (antibody staining) and blue marks nuclei (DAPI staining). (c) Model size (log scale) and accuracy: the MT-BKD student (ESM-2-8M, 8 M parameters) achieves 0.75 accuracy while being 81 × and 150 × smaller than the two teachers. Data from Fang et al. (2026a) .
Knowledge distillation is a powerful method for model compression, enabling the efficient deployment of complex deep learning models (teachers), including large language models. However, its underlying statistical mechanisms remain unclear, and uncertainty evaluation is often overlooked, especially in real-world scenarios requiring diverse teacher expertise. To address these challenges, we introduce \textit{Multi-Teacher Bayesian Knowledge Distillation} (MT-BKD), where a distilled student model learns from multiple teachers within the Bayesian framework. Our approach leverages Bayesian inference to capture inherent uncertainty in the distillation process. We introduce a teacher-informed prior, integrating external knowledge from teacher models and task-specific training data, offering better generalization, robustness, and scalability. Additionally, an entropy-based weighting mechanism adaptively adjusts each teacher's influence, allowing the student to combine multiple sources of expertise effectively. MT-BKD enhances the interpretability of the student model's learning process, improves predictive accuracy, and provides uncertainty quantification. We validate MT-BKD on both synthetic and real-world tasks, including protein subcellular location prediction and image classification. Our experiments show improved performance and robust uncertainty quantification, highlighting the strengths of our MT-BKD framework.
Luyang Fang, Yongkai Chen, Jiazhang Cai +2
Department of Statistics, University of Georgia · Department of Statistics, Harvard University
Knowledge Distillation (KD) is a central paradigm for transferring knowledge from a large teacher network to a typically smaller student model, often by leveraging soft probabilistic outputs. While KD has shown strong empirical success in numerous applications, its theoretical underpinnings remain only partially understood. In this work, we adopt a Bayesian perspective on KD to rigorously analyze the convergence behavior of students trained with Stochastic Gradient Descent (SGD). We study two regimes: (i) when the teacher provides the exact Bayes Class Probabilities (BCPs); and (ii) supervision with noisy approximations of the BCPs. Our analysis shows that learning from BCPs yields variance reduction and removes neighborhood terms in the convergence bounds compared to one-hot supervision. We further characterize how the level of noise affects generalization and accuracy. Motivated by these insights, we advocate the use of Bayesian deep learning models, which typically provide improved estimates of the BCPs, as teachers in KD. Consistent with our analysis, we experimentally demonstrate that students distilled from Bayesian teachers not only achieve higher accuracies (up to +4.27%), but also exhibit more stable convergence (up to 30% less noise), compared to students distilled from deterministic teachers.
Itai Morad, Nir Shlezinger, Yonina C. Eldar
School of ECE, Ben-Gurion University, Be’er Sheva, Israel · Faculty of Math and CS, Weizmann Institute of Science, Rehovot, Israel
Knowledge distillation is widely used to improve generalization in practice, yet its theoretical understanding remains elusive. In the standard distillation setting, a teacher model provides soft predictions to guide the training of a student model. We model teacher and student training as coupled stochastic processes and introduce a distillation divergence, defined as the Kullback-Leibler divergence between these two stochastic kernels. Within this framework, we derive two generalization bounds for the student model relative to the teacher's generalization gap: an upper bound under a sub-Gaussian assumption via algorithmic stability, and a lower bound under a central condition with sharper dependence on the distillation divergence. We further develop a loss-sharpness-aware bound with an explicit tightness regime, showing that the teacher's local flatness can strictly tighten the bound. Additionally, in a linear Gaussian case study, the distillation divergence admits an interpretable decomposition into bias, variance, and rank-bottleneck costs, yielding practical guidance for distillation design.
Bingying Li, Haiyun He
Internet of Things Thrust, Information Hub The Hong Kong University of Science and Technology (Guangzhou) Guangzhou, China