Diffusion models trained on a finite dataset first learn to generate novel, high-quality samples and only much later collapse onto their training set. We identify the mechanism behind this separation of timescales and the object that probes it. The training dynamics of the score function are governed---exactly, and at any width---by the Gram matrix of the Neural Tangent Kernel (NTK) evaluated on the noisy training data, so the timescales of generalization and of memorization must be encoded in its spectrum. We show that they are, and that the structure responsible has no analogue in standard kernel settings. The use of multiple noise realizations per sample (m noised copies at a fixed noise level) in the score-matching loss is what restructures the Gram matrix spectrum into two distinct parts. The first, of large eigenvalues, carries the global features of the target distribution and is present already for m=1. The second, which the repeated noising creates, consists of the smallest eigenvalues and is supported on eigenvectors aligned with the sample-specific noise directions; it sets a memorization timescale parametrically larger in the training set size n. We establish this picture on two fronts. Analytically, we solve the spectrum in the lazy high-dimensional limit for both linear (n≍d) and polynomial (n≍dk) sample complexities, and prove through a bias--variance decomposition that the first bulk minimizes the approximation error while the second drives the error associated with memorization. Empirically, we show the same two-bulk structure in Convolutional NTKs on CelebA and in finite-width U-Nets trained well beyond the lazy regime, and we make the link causal: truncating the Gram matrix at rank r tunes the generalization--memorization transition, and an L2 penalty targeting the second bulk suppresses memorization in feature-learning U-Nets.
Figures & tables
Figure 1 : Gram matrix spectrum. Eigenvalues λr in decreasing order against their rank. ( Left ) Sketch of the spectrum predicted by Theorem 2.3 for the Gram matrix Gνα,μβ=f(Yνα⊤Yμβ/d) . ( Middle ) Spectrum of G with isotropic Σ=σ2Id , t=0.1 , m=4 , ψn=4 and f the NTK of a two-layer ReLU network (Appendix A.8 ) in the linear regime at several values of d ; the dashed black curve is the linear equivalent Glin of ( 6 ) at d=512 . ( Right ) Quadratic regime n=2d2 at d=64 ; the dashed black curve is the polynomial equivalent of Theorem 2.3 . Details in Appendix D.7 .
Figure 2 : Bias–variance decomposition. B2 , V and B2+V for the kernel f(u)=E[tanh(z1)tanh(z2)] , with (z1,z2) unit-variance Gaussians of correlation u (Appendix A.8 ), at σ2=1 , m=4 and t=0.1 . ( Left ) Vs. the rescaled ridge γ~ at ψn=8 , axis reversed to show the correspondence with training time; the inset shows the same vs. the training time τ of ( 5 ). Solid lines: Theorem 2.2 ; markers: experiments at d=256 ; in the inset, markers and dotted line are the finite- d gradient flow. ( Middle ) B2 (top) and V (bottom) vs. ψn , for the optimal ridge γ~∗=argminγ~Ltest (solid) and the ridgeless limit γ~→0+ (dashed). ( Right ) B2 vs. γ~ for several ψn , dots marking γ~∗ ; the inset shows the same vs. ψnγ~ . Details in Appendix D.10 .
Figure 3 : Structure of the generalization and memorization bulks in the CNTK spectrum on CelebA. (Left) Ordered eigenvalues for several n at m=4 . The dotted lines indicate r=n . Inset: Same, rescaled by n . (Right) Eigenvalue histogram with counts rescaled by n . Inset: collapse of the top eigenvalues under rescaling by n . s=0.05 , averaged over 10 training sets.
Figure 4 : Spectrally truncated kernel regression and generalization–memorization transition. (Left) FID (solid) and fmem (dashed) vs. truncation rank r for several n . Inset: fmem vs. r/n . (Middle) Time-averaged L2 training (solid) and test (dashed) losses; the inset shows the same curves against r/n . (Right) Top row: training images. Rows 2–6: generated samples ( n=3072 ) starting from a fixed training noise associated with the top-row clean image at r∈{2048,4096,8192,16384,N=nm} .
Figure 5 : Training-time-dependent empirical Gram eigenspectrum of a U-Net. Eigenvalues of the empirical Gram matrix computed for several training times τ at s=0.23 for (Left) the unregularized dynamics, and (Right) the L2 -regularized dynamics.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Empirical histogram of the eigenvalues of Glin for d=100 (blue) averaged over 10 runs and the analytical prediction obtained by solving the equations on the Stieltjes transform for m=5,ψn=4.0 and t=0.1 for ρΣ(λ)=21δ(λ−1)+21δ(λ−0.1) and for the kernel parameters μ1=1,μI=0.1,μB=0.2 . The left and right panels corresponds to different ranges of λ .
Figure 7 : Comparison between the linear equivalent and the non-linear NTK spectrum in finite dimension for d=128,ψn=4,m=4, and t=0.1 .
Figure 8 : Analytical solutions of the rescaled bulks ρ1,ρ2 for different values of ψn for m=5,t=0.1,Σ=Id,μ1=1.0,μB=0.2,μI=0.1,μ0=0.0 and the asymptotic limit δ(λ−1) .
Figure 9 : Bias and variance against the sample complexity. B2 , V and B2+V against mψn=mn/d for the ridgeless estimator γ~→0+ , at t=0.1 , σ2=1 , with the kernel of Figure 2 , for m=4 ( left ) and m=1 ( right ). Solid lines are the equations of Theorem 2.2 ; markers are experiments at d=256 .
Figure 10 : The ridge as a proxy for early stopping. Ltest , B2 and V against the training time τ , at m=1 ( top ) and m=4 ( bottom ). Black: the theory of Theorem 2.2 evaluated at γ~=1/τ . Green: kernel gradient flow of ( 5 ) measured at finite d=256 , ψn=8 , averaged over 16 ( m=1 ) and 8 ( m=4 ) realizations of the dataset and 512 test points. Dotted verticals are τgen and τmem . Panels are scaled independently: at m=1 the bias vanishes whereas at m=4 it saturates at Θ(1) .
Figure 11 : Scalings with t . B2 ( top ) and V ( bottom ) against t at ψn=103 and m=4 , where 1/(mψn)≪t≪1 , for the optimally ridged estimator γ~∗ (solid) and the ridgeless one γ~→0+ (dashed), with the kernel of Fig. 2 . Lines are the equations of Theorem 2.2 .
Figure 12 : Ordered eigenvalues for several m at fixed n=256 and their rescaling by m in the inset. Spectra obtained at s=0.05 and averaged over 10 realizations of the training set.
Figure 13 : Training and test losses along training for the (Left) unregularized and (Right) regularized dynamics of Fig. 5 .
Diffusion models achieve remarkable generation quality, yet face a fundamental challenge known as memorization, where generated samples can replicate training samples exactly. We develop a theoretical framework to explain this phenomenon by showing that the empirical score function (the score function corresponding to the empirical distribution) is a weighted sum of the score functions of Gaussian distributions, in which the weights are sharp softmax functions. This structure causes individual training samples to dominate the score function, resulting in sampling collapse. In practice, approximating the empirical score function with a neural network can partially alleviate this issue and improve generalization. Our theoretical framework explains why: In training, the neural network learns a smoother approximation of the weighted sum, allowing the sampling process to be influenced by local manifolds rather than single points. Leveraging this insight, we propose two novel methods to further enhance generalization: (1) Noise Unconditioning enables each training sample to adaptively determine its score function weight to increase the effect of more training samples, thereby preventing single-point dominance and mitigating collapse. (2) Temperature Smoothing introduces an explicit parameter to control the smoothness. By increasing the temperature in the softmax weights, we naturally reduce the dominance of any single training sample and mitigate memorization. Experiments across multiple datasets validate our theoretical analysis and demonstrate the effectiveness of the proposed methods in improving generalization while maintaining high generation quality.
Xinyu Zhou, Jiawei Zhang, Stephen J. Wright
Department of Computer Sciences, University of Wisconsin Madison, WI, USA.
Generative neural networks learn how to produce highly realistic images from a large, but finite number of examples - or do they simply memorise their training set? To settle this question, Kadkhodaie, Guth, Simoncelli and Mallat (ICLR '24) trained diffusion models independently on disjoint subsets of a dataset and showed that they converge to nearly the same density when the number of training images is large enough. This result raises two basic questions: how much data do you need for convergence, and what does convergence capture about learning the data distribution? Here, we address these questions by providing an exact analytical characterisation of the transition from memorisation to generalisation in linear generative models. We find that these models memorise at small load, while convergence emerges continuously when the number of samples is linear in the input dimension. Strikingly, we find that convergence is insensitive to recovery of the principal latent factors of the data, which are recovered in a sharp transition. After extending our approach to data with power-law spectra, we find the same distinction between convergence and latent recovery in our experiments with convolutional denoisers and in the data of Kadkhodaie et al. We thus show that generalisation in generative models decomposes into at least two distinct objectives: matching the bulk of the data distribution and recovering the principal latent factors. These objectives correspond to two different distances between true and learnt data distribution, and only the first one is captured by convergence.
Antoine Maillard, Sebastian Goldt
1INRIA Paris & DI ENS, PSL University, Paris, France · 2International School of Advanced Studies (SISSA), Trieste, Italy
This position paper argues that understanding generalization in diffusion models requires fundamentally new theoretical frameworks that go beyond both classical statistical learning theory and the benign overfitting paradigm developed for supervised learning. In diffusion models, unlike in supervised learning, memorization of training data and generalization to novel samples are incompatible: a model that has fully memorized its training set generates copies rather than novel data. Several theoretical explanations for why practical diffusion models nevertheless generalize have been proposed, based on capacity limitations, implicit regularization from optimization, or architectural inductive biases, but their interactions remain unclear. We argue that the field should pivot from explaining why the diffusion models do not memorize to investigating what the model actually learns during pre-memorization phase. To highlight our stance, we conduct empirical study of diffusion models trained on CIFAR-10, and we distill the findings into concrete open questions that we believe are key to improve understanding of generalization in diffusion models.
Pierre Marion, Yu-Han Wu
Inria, École Normale Supérieure, PSL Research University · LPSM, Sorbonne University & Google DeepMind