Diffusion models trained on a finite dataset first learn to generate novel, high-quality samples and only much later collapse onto their training set. We identify the mechanism behind this separation of timescales and the object that probes it. The training dynamics of the score function are governed---exactly, and at any width---by the Gram matrix of the Neural Tangent Kernel (NTK) evaluated on the noisy training data, so the timescales of generalization and of memorization must be encoded in its spectrum. We show that they are, and that the structure responsible has no analogue in standard kernel settings. The use of multiple noise realizations per sample (m noised copies at a fixed noise level) in the score-matching loss is what restructures the Gram matrix spectrum into two distinct parts. The first, of large eigenvalues, carries the global features of the target distribution and is present already for m=1. The second, which the repeated noising creates, consists of the smallest eigenvalues and is supported on eigenvectors aligned with the sample-specific noise directions; it sets a memorization timescale parametrically larger in the training set size n. We establish this picture on two fronts. Analytically, we solve the spectrum in the lazy high-dimensional limit for both linear (n≍d) and polynomial (n≍dk) sample complexities, and prove through a bias--variance decomposition that the first bulk minimizes the approximation error while the second drives the error associated with memorization. Empirically, we show the same two-bulk structure in Convolutional NTKs on CelebA and in finite-width U-Nets trained well beyond the lazy regime, and we make the link causal: truncating the Gram matrix at rank r tunes the generalization--memorization transition, and an L2 penalty targeting the second bulk suppresses memorization in feature-learning U-Nets.
Figures & tables
Figure 1 : Gram matrix spectrum. Eigenvalues λr in decreasing order against their rank. ( Left ) Sketch of the spectrum predicted by Theorem 2.3 for the Gram matrix Gνα,μβ=f(Yνα⊤Yμβ/d) . ( Middle ) Spectrum of G with isotropic Σ=σ2Id , t=0.1 , m=4 , ψn=4 and f the NTK of a two-layer ReLU network (Appendix A.8 ) in the linear regime at several values of d ; the dashed black curve is the linear equivalent Glin of ( 6 ) at d=512 . ( Right ) Quadratic regime n=2d2 at d=64 ; the dashed black curve is the polynomial equivalent of Theorem 2.3 . Details in Appendix D.7 .
Figure 2 : Bias–variance decomposition. B2 , V and B2+V for the kernel f(u)=E[tanh(z1)tanh(z2)] , with (z1,z2) unit-variance Gaussians of correlation u (Appendix A.8 ), at σ2=1 , m=4 and t=0.1 . ( Left ) Vs. the rescaled ridge γ~ at ψn=8 , axis reversed to show the correspondence with training time; the inset shows the same vs. the training time τ of ( 5 ). Solid lines: Theorem 2.2 ; markers: experiments at d=256 ; in the inset, markers and dotted line are the finite- d gradient flow. ( Middle ) B2 (top) and V (bottom) vs. ψn , for the optimal ridge γ~∗=argminγ~Ltest (solid) and the ridgeless limit γ~→0+ (dashed). ( Right ) B2 vs. γ~ for several ψn , dots marking γ~∗ ; the inset shows the same vs. ψnγ~ . Details in Appendix D.10 .
Figure 3 : Structure of the generalization and memorization bulks in the CNTK spectrum on CelebA. (Left) Ordered eigenvalues for several n at m=4 . The dotted lines indicate r=n . Inset: Same, rescaled by n . (Right) Eigenvalue histogram with counts rescaled by n . Inset: collapse of the top eigenvalues under rescaling by n . s=0.05 , averaged over 10 training sets.
Figure 4 : Spectrally truncated kernel regression and generalization–memorization transition. (Left) FID (solid) and fmem (dashed) vs. truncation rank r for several n . Inset: fmem vs. r/n . (Middle) Time-averaged L2 training (solid) and test (dashed) losses; the inset shows the same curves against r/n . (Right) Top row: training images. Rows 2–6: generated samples ( n=3072 ) starting from a fixed training noise associated with the top-row clean image at r∈{2048,4096,8192,16384,N=nm} .
Figure 5 : Training-time-dependent empirical Gram eigenspectrum of a U-Net. Eigenvalues of the empirical Gram matrix computed for several training times τ at s=0.23 for (Left) the unregularized dynamics, and (Right) the L2 -regularized dynamics.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Empirical histogram of the eigenvalues of Glin for d=100 (blue) averaged over 10 runs and the analytical prediction obtained by solving the equations on the Stieltjes transform for m=5,ψn=4.0 and t=0.1 for ρΣ(λ)=21δ(λ−1)+21δ(λ−0.1) and for the kernel parameters μ1=1,μI=0.1,μB=0.2 . The left and right panels corresponds to different ranges of λ .
Figure 7 : Comparison between the linear equivalent and the non-linear NTK spectrum in finite dimension for d=128,ψn=4,m=4, and t=0.1 .
Figure 8 : Analytical solutions of the rescaled bulks ρ1,ρ2 for different values of ψn for m=5,t=0.1,Σ=Id,μ1=1.0,μB=0.2,μI=0.1,μ0=0.0 and the asymptotic limit δ(λ−1) .
Figure 9 : Bias and variance against the sample complexity. B2 , V and B2+V against mψn=mn/d for the ridgeless estimator γ~→0+ , at t=0.1 , σ2=1 , with the kernel of Figure 2 , for m=4 ( left ) and m=1 ( right ). Solid lines are the equations of Theorem 2.2 ; markers are experiments at d=256 .
Figure 10 : The ridge as a proxy for early stopping. Ltest , B2 and V against the training time τ , at m=1 ( top ) and m=4 ( bottom ). Black: the theory of Theorem 2.2 evaluated at γ~=1/τ . Green: kernel gradient flow of ( 5 ) measured at finite d=256 , ψn=8 , averaged over 16 ( m=1 ) and 8 ( m=4 ) realizations of the dataset and 512 test points. Dotted verticals are τgen and τmem . Panels are scaled independently: at m=1 the bias vanishes whereas at m=4 it saturates at Θ(1) .
Figure 11 : Scalings with t . B2 ( top ) and V ( bottom ) against t at ψn=103 and m=4 , where 1/(mψn)≪t≪1 , for the optimally ridged estimator γ~∗ (solid) and the ridgeless one γ~→0+ (dashed), with the kernel of Fig. 2 . Lines are the equations of Theorem 2.2 .
Figure 12 : Ordered eigenvalues for several m at fixed n=256 and their rescaling by m in the inset. Spectra obtained at s=0.05 and averaged over 10 realizations of the training set.
Figure 13 : Training and test losses along training for the (Left) unregularized and (Right) regularized dynamics of Fig. 5 .