cs.LGMay 29, 2026

Rethinking the Role of Temperature in Large Language Model Distillation

Authors: Hoang-Chau LuongLingwei Chen

Organizations: Golisano College of Computing and Information Sciences Rochester Institute of Technology Rochester, NY, United States

Abstract

Reverse Kullback-Leibler (RKL) divergence is widely favored over forward KL (FKL) in large language models (LLM) distillation, yet this preference is largely based on comparisons that omit the temperature ττ, overlooking its central role in softening teacher distributions and improving knowledge transfer. In this work, we revisit temperature in LLM distillation and show that it fundamentally changes the comparison between FKL and RKL. Our analysis reveals an asymmetric effect: temperature substantially enriches FKL with non-dominant token signals, whereas it mainly rescales RKL gradients, causing FKL to benefit much more from ττ scaling than RKL. This asymmetry overturns the standard empirical conclusion: although RKL outperforms FKL at τ=1τ=1, FKL consistently surpasses RKL at higher temperatures across instruction-following benchmarks. Moreover, the impact of temperature is not limited to FKL; it improves a broader family of distillation objectives, enabling simple KL-based methods to achieve competitive performance against recent state-of-the-art LLM distillation approaches.

Explore similar work

CardsList