math.STMay 18, 2026

Self-Distillation is Optimal Among Spectral Shrinkage Estimators in Spiked Covariance Models

Authors: Radu LecoiuDebarghya MukherjeePragya Sur

Organizations: Department of Statistics, Harvard University · Department of Mathematics & Statistics, Boston University

Abstract

Self-distillation has emerged as a promising technique for improving model performance in modern machine learning systems. We develop the statistical foundations of self-distillation in spiked covariance models, by introducing and analyzing a broad class of estimators, namely spectral shrinkage estimators. We establish that for spiked covariance matrices with ss spikes, ss-step self-distillation achieves optimal performance among spectral shrinkage estimators, outperforming well-known estimators in statistics and machine learning. Moreover, we show that ss steps are necessary for optimality: any (sk)(s-k)-step distilled estimator is strictly suboptimal for 1ks1 \leq k \leq s. For the special subclass of isotropic covariances, we show that optimally tuned Ridge regression performs best among spectral shrinkage estimators. We also study a federated approach where multiple data centers share spectral shrinkage estimators and a common server seeks to aggregate them to achieve optimal performance. In this case, we find that the best local rule again takes the form of self-distillation, though it differs from the optimal rule when data are hosted centrally on a single server. Together, our results elucidate why self-distillation improves predictive performance and provide a broader statistical framework connecting it with classical shrinkage-based methods.

Explore similar work

CardsList
  1. Sparse Regression Distilled from a Single Robust Fit

    Sep 20, 2026Wooyoung Shin, Seunghwan ParkEstimators

  2. Covariance Shrinkage via Stochastic Interpolation

    Jun 5, 2026Mathieu Chalvidal, Florentin Coeurdoux, Eric Vanden-EijndenCovariance EstimationEmpirical Risk Minimization