Conditional diffusion models generate diverse, novel, and high-quality samples under prescribed conditions. However, theoretical understanding of their memorization and generalization remains limited, while recent works have characterized these behaviors primarily in unconditional settings. In this work, we analyze a random-feature conditional score model in the high-dimensional proportional limit, deriving asymptotic expressions for training and test losses. By decomposing the test loss, we show that in the overparameterized regime, increasing model width improves prediction of the condition-dependent mean while reducing within-condition prediction variance, a phenomenon we term "malign generalization." Furthermore, analyzing the training loss reveals that more informative conditions lead to memorization of training samples at smaller widths. These theoretical findings are supported by experiments with U-Net architectures on realistic data.
Figures & tables
Figure 1: CelebA super-resolution to 32×32 ; each row shows a condition, four generated samples, and the training image nearest to any of them. Models in (a) and (b) are trained identically except for condition resolution. (a) With the informative 8×8 condition, samples for a training condition copy its training image, while samples for a test condition are novel and aligned with the condition but nearly identical (malign generalization). (b) With the weaker 2×2 condition, samples are diverse under both training and test conditions.
Figure 2: Test loss (ht/at2)Ltestt=ME+RE , ME ( 9 ), RE ( 10 ), PV ( 11 ), and training loss Ltraint from Theorem 4.2 against width ψp for different SNRs ( 5 ) at t=0.01 , ϱ=tanh , λ=10−5 , and ψn=32 . Top : Gaussian denoising ( ψy=1 ), varying c so that SNR(1,c)∈{4,16,64} . Bottom : super-resolution ( c=1 ), varying ψy so that SNR(ψy,1)∈{4,16,64} . Dashed lines mark the PV and the training loss of the true denoiser E[x∣xt,y] . Dots denote finite random-feature simulations at dx=256 , averaged over five realizations; their min–max bars are smaller than the dots. Appendix F.2 shows them as heatmaps over the full range of SNR and for other ψn .
Figure 3: ME and PV from Theorem 4.2 (black) split into the linear part (blue) and the nonlinear part (red) of Lemma 4.1 , at SNR=16 in the setting of Figure 2 .
Figure 4: Full reverse-process generation with random-feature models (setting of Figure 2 , with dx=64 ).
Figure 5: Full reverse-process generation with U-Nets on 32×32 CelebA, against the width W∈{16,32,64,128} at n=32768 , every model trained for 200 k steps. Markers are the mean and bars the min–max range over three seeds.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
dx,dy,n,p
Data dimension, condition dimension, sample size, and feature count.
ψn,ψp,ψy
Limits of n/dx , p/dx , and dy/dx .
U⋆,Π,Π⊥
Orthonormal condition embedding and the two projectors in the main-text model.
B,Σ
Conditional-mean map and residual covariance in x=By+Σ1/2ξ .
υΣ,d2,υΣ2
dx−1\trΣ and its limit.
Σt,τ
at2Σ+htI and its limiting normalized trace.
Appendix
Table 6
Figure 6: Error metrics at test conditions on CelebA at n=32768 , as in Figure 5 : the per-sample error, the conditional error of Figure 5 (the error of the mean of K=64 samples), the debiased error, and the pooled error (super-resolution only). Markers are the mean and bars the min–max range over three seeds.
Figure 7: Figure 5 at n=8192 , with the error metrics of Figure 6 in the last four columns: full reverse-process generation with U-Nets on 32×32 CelebA against the width W∈{16,32,64,128} , every model trained for 200 k steps. Markers are the mean and bars the min–max range over three seeds.
Figure 8: Test loss (ht/at2)Ltestt=ME+RE , ME , RE , PV and Ltraint of Theorem 4.2 at ψn=8 ; settings as in Appendix F.2 .
Figure 9: Test loss (ht/at2)Ltestt=ME+RE , ME , RE , PV and Ltraint of Theorem 4.2 at ψn=32 , as in Figures 2 and 4 ; settings as in Appendix F.2 .
Figure 10: Test loss (ht/at2)Ltestt=ME+RE , ME , RE , PV and Ltraint of Theorem 4.2 at ψn=64 ; settings as in Appendix F.2 .
Figure 11: Test loss (ht/at2)Ltestt=ME+RE , ME , RE , PV and Ltraint of Theorem 4.2 at ψn=256 ; settings as in Appendix F.2 .
Figure 12: As Figure 9 ( ψn=32 ) at t=0.001 .
Figure 13: As Figure 9 ( ψn=32 ) at t=0.1 .
Figure 14: As Figure 9 ( ψn=32 ) at t=1 .
Figure 15: Training loss Ltraint of Theorem 4.2 over SNR(1,c) and SNR(ψy,1) for ψp/ψn∈{2,8,32,64} , at ψn=256 , t=0.01 , ϱ(u)=tanh(u) and λ=10−5 . The panels share a logarithmic color scale, and dotted lines are level sets.
Figure 16: Samples at n=32768 and W=32 ( 200 k steps) for super-resolution with the 8×8 , 4×4 and 2×2 conditions. Each row is one condition: the source image, the condition made from it (nearest-upsampled), five samples generated from it, and the training image nearest to any of them. Left: training conditions; right: test conditions.
Figure 17: Samples at n=32768 and W=32 ( 200 k steps) for Gaussian denoising with σ=0.1 , 0.4 and 2 ; the noisy condition is clipped to [0,1] for display. Rows and columns as in Figure 16 .
Figure 18: Samples at n=32768 for W=16 , 32 and 64 ( 200 k steps) for super-resolution with the 8×8 condition. Rows and columns as in Figure 16 .
Figure 19: Samples at n=32768 for W=16 , 32 and 64 ( 200 k steps) for Gaussian denoising with σ=0.1 . Rows and columns as in Figure 16 .
This position paper argues that understanding generalization in diffusion models requires fundamentally new theoretical frameworks that go beyond both classical statistical learning theory and the benign overfitting paradigm developed for supervised learning. In diffusion models, unlike in supervised learning, memorization of training data and generalization to novel samples are incompatible: a model that has fully memorized its training set generates copies rather than novel data. Several theoretical explanations for why practical diffusion models nevertheless generalize have been proposed, based on capacity limitations, implicit regularization from optimization, or architectural inductive biases, but their interactions remain unclear. We argue that the field should pivot from explaining why the diffusion models do not memorize to investigating what the model actually learns during pre-memorization phase. To highlight our stance, we conduct empirical study of diffusion models trained on CIFAR-10, and we distill the findings into concrete open questions that we believe are key to improve understanding of generalization in diffusion models.
Pierre Marion, Yu-Han Wu
Inria, École Normale Supérieure, PSL Research University · LPSM, Sorbonne University & Google DeepMind
Diffusion models achieve remarkable generation quality, yet face a fundamental challenge known as memorization, where generated samples can replicate training samples exactly. We develop a theoretical framework to explain this phenomenon by showing that the empirical score function (the score function corresponding to the empirical distribution) is a weighted sum of the score functions of Gaussian distributions, in which the weights are sharp softmax functions. This structure causes individual training samples to dominate the score function, resulting in sampling collapse. In practice, approximating the empirical score function with a neural network can partially alleviate this issue and improve generalization. Our theoretical framework explains why: In training, the neural network learns a smoother approximation of the weighted sum, allowing the sampling process to be influenced by local manifolds rather than single points. Leveraging this insight, we propose two novel methods to further enhance generalization: (1) Noise Unconditioning enables each training sample to adaptively determine its score function weight to increase the effect of more training samples, thereby preventing single-point dominance and mitigating collapse. (2) Temperature Smoothing introduces an explicit parameter to control the smoothness. By increasing the temperature in the softmax weights, we naturally reduce the dominance of any single training sample and mitigate memorization. Experiments across multiple datasets validate our theoretical analysis and demonstrate the effectiveness of the proposed methods in improving generalization while maintaining high generation quality.
Xinyu Zhou, Jiawei Zhang, Stephen J. Wright
Department of Computer Sciences, University of Wisconsin Madison, WI, USA.
When do language diffusion models memorize their training data, and how to quantitatively assess their true generative regime? We address these questions by showing that Uniform-based Discrete Diffusion Models (UDDMs) fundamentally behave as Associative Memories (AMs) with emergent creative capabilities. The core idea of an AM is to reliably recover stored data points as memories by establishing distinct basins of attraction around them. Historically, models like Hopfield networks use an explicit energy function to guarantee these stable attractors. We broaden this perspective by leveraging the observation that energy is not strictly necessary, as basins of attraction can also be formed via conditional likelihood maximization. By evaluating token recovery of training and test examples, we identify in UDDMs a sharp memorization-to-generalization transition governed by the size of the training dataset: as it increases, basins around training examples shrink and basins around unseen test examples expand, until both later converge to the same level. Crucially, we can detect this transition using only the conditional entropy of predicted token sequences: memorization is characterized by vanishing conditional entropy, while in the generalization regime the conditional entropy of most tokens remains finite. Thus, conditional entropy offers a practical probe for the memorization-to-generalization transition in deployed models.
Bao Pham, Mohammed J. Zaki, Luca Ambrogioni +2
1Rensselaer Polytechnic Institute (RPI) · 2Donders Institute for Brain, Cognition, and Behaviour, Radboud University · 3Dynamical Mind +1