cs.LGSep 28, 2026

First Learn, Then Memorize: The Spectral Bias of Diffusion Models

Authors: Raphaël Urfin, Tony Bonnaire, Giulio Biroli, Marc Mézard

Organizations: Laboratoire de Physique de l’École normale supérieure, ENS, Université PSL, CNRS, Sorbonne Université, Université Paris Cité, F-75005 Paris, France · Université Paris-Saclay, CNRS, Institut d’Astrophysique Spatiale, 91405 Orsay, France · Department of Computing Sciences, Bocconi University, Milano, Italy

Abstract

Diffusion models trained on a finite dataset first learn to generate novel, high-quality samples and only much later collapse onto their training set. We identify the mechanism behind this separation of timescales and the object that probes it. The training dynamics of the score function are governed---exactly, and at any width---by the Gram matrix of the Neural Tangent Kernel (NTK) evaluated on the noisy training data, so the timescales of generalization and of memorization must be encoded in its spectrum. We show that they are, and that the structure responsible has no analogue in standard kernel settings. The use of multiple noise realizations per sample (mm noised copies at a fixed noise level) in the score-matching loss is what restructures the Gram matrix spectrum into two distinct parts. The first, of large eigenvalues, carries the global features of the target distribution and is present already for m=1m=1. The second, which the repeated noising creates, consists of the smallest eigenvalues and is supported on eigenvectors aligned with the sample-specific noise directions; it sets a memorization timescale parametrically larger in the training set size nn. We establish this picture on two fronts. Analytically, we solve the spectrum in the lazy high-dimensional limit for both linear (n≍dn \asymp d) and polynomial (n≍dkn \asymp d^k) sample complexities, and prove through a bias--variance decomposition that the first bulk minimizes the approximation error while the second drives the error associated with memorization. Empirically, we show the same two-bulk structure in Convolutional NTKs on CelebA and in finite-width U-Nets trained well beyond the lazy regime, and we make the link causal: truncating the Gram matrix at rank rr tunes the generalization--memorization transition, and an L2L_2 penalty targeting the second bulk suppresses memorization in feature-learning U-Nets.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Smoothing the Score Function to Enhance Generalization in Diffusion Models

    Date pendingXinyu Zhou, Jiawei Zhang, Stephen J. WrightScore-Based Diffusion ModelGumbel-Softmax Relaxation

  2. Memorisation, convergence and generalisation in generative models

    May 20, 2026Antoine Maillard, Sebastian GoldtGenerative ModelsGeneralisation

  3. Understanding diffusion models requires rethinking (again) generalization

    May 7, 2026Pierre Marion, Yu-Han WuOverfittingDiffusion Models