Conditional diffusion models generate diverse, novel, and high-quality samples under prescribed conditions. However, theoretical understanding of their memorization and generalization remains limited, while recent works have characterized these behaviors primarily in unconditional settings. In this work, we analyze a random-feature conditional score model in the high-dimensional proportional limit, deriving asymptotic expressions for training and test losses. By decomposing the test loss, we show that in the overparameterized regime, increasing model width improves prediction of the condition-dependent mean while reducing within-condition prediction variance, a phenomenon we term "malign generalization." Furthermore, analyzing the training loss reveals that more informative conditions lead to memorization of training samples at smaller widths. These theoretical findings are supported by experiments with U-Net architectures on realistic data.
Figures & tables
Figure 1: CelebA super-resolution to 32×32 ; each row shows a condition, four generated samples, and the training image nearest to any of them. Models in (a) and (b) are trained identically except for condition resolution. (a) With the informative 8×8 condition, samples for a training condition copy its training image, while samples for a test condition are novel and aligned with the condition but nearly identical (malign generalization). (b) With the weaker 2×2 condition, samples are diverse under both training and test conditions.
Figure 2: Test loss (ht/at2)Ltestt=ME+RE , ME ( 9 ), RE ( 10 ), PV ( 11 ), and training loss Ltraint from Theorem 4.2 against width ψp for different SNRs ( 5 ) at t=0.01 , ϱ=tanh , λ=10−5 , and ψn=32 . Top : Gaussian denoising ( ψy=1 ), varying c so that SNR(1,c)∈{4,16,64} . Bottom : super-resolution ( c=1 ), varying ψy so that SNR(ψy,1)∈{4,16,64} . Dashed lines mark the PV and the training loss of the true denoiser E[x∣xt,y] . Dots denote finite random-feature simulations at dx=256 , averaged over five realizations; their min–max bars are smaller than the dots. Appendix F.2 shows them as heatmaps over the full range of SNR and for other ψn .
Figure 3: ME and PV from Theorem 4.2 (black) split into the linear part (blue) and the nonlinear part (red) of Lemma 4.1 , at SNR=16 in the setting of Figure 2 .
Figure 4: Full reverse-process generation with random-feature models (setting of Figure 2 , with dx=64 ).
Figure 5: Full reverse-process generation with U-Nets on 32×32 CelebA, against the width W∈{16,32,64,128} at n=32768 , every model trained for 200 k steps. Markers are the mean and bars the min–max range over three seeds.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
dx,dy,n,p
Data dimension, condition dimension, sample size, and feature count.
ψn,ψp,ψy
Limits of n/dx , p/dx , and dy/dx .
U⋆,Π,Π⊥
Orthonormal condition embedding and the two projectors in the main-text model.
B,Σ
Conditional-mean map and residual covariance in x=By+Σ1/2ξ .
υΣ,d2,υΣ2
dx−1\trΣ and its limit.
Σt,τ
at2Σ+htI and its limiting normalized trace.
Appendix
Table 6
Figure 6: Error metrics at test conditions on CelebA at n=32768 , as in Figure 5 : the per-sample error, the conditional error of Figure 5 (the error of the mean of K=64 samples), the debiased error, and the pooled error (super-resolution only). Markers are the mean and bars the min–max range over three seeds.
Figure 7: Figure 5 at n=8192 , with the error metrics of Figure 6 in the last four columns: full reverse-process generation with U-Nets on 32×32 CelebA against the width W∈{16,32,64,128} , every model trained for 200 k steps. Markers are the mean and bars the min–max range over three seeds.
Figure 8: Test loss (ht/at2)Ltestt=ME+RE , ME , RE , PV and Ltraint of Theorem 4.2 at ψn=8 ; settings as in Appendix F.2 .
Figure 9: Test loss (ht/at2)Ltestt=ME+RE , ME , RE , PV and Ltraint of Theorem 4.2 at ψn=32 , as in Figures 2 and 4 ; settings as in Appendix F.2 .
Figure 10: Test loss (ht/at2)Ltestt=ME+RE , ME , RE , PV and Ltraint of Theorem 4.2 at ψn=64 ; settings as in Appendix F.2 .
Figure 11: Test loss (ht/at2)Ltestt=ME+RE , ME , RE , PV and Ltraint of Theorem 4.2 at ψn=256 ; settings as in Appendix F.2 .
Figure 12: As Figure 9 ( ψn=32 ) at t=0.001 .
Figure 13: As Figure 9 ( ψn=32 ) at t=0.1 .
Figure 14: As Figure 9 ( ψn=32 ) at t=1 .
Figure 15: Training loss Ltraint of Theorem 4.2 over SNR(1,c) and SNR(ψy,1) for ψp/ψn∈{2,8,32,64} , at ψn=256 , t=0.01 , ϱ(u)=tanh(u) and λ=10−5 . The panels share a logarithmic color scale, and dotted lines are level sets.
Figure 16: Samples at n=32768 and W=32 ( 200 k steps) for super-resolution with the 8×8 , 4×4 and 2×2 conditions. Each row is one condition: the source image, the condition made from it (nearest-upsampled), five samples generated from it, and the training image nearest to any of them. Left: training conditions; right: test conditions.
Figure 17: Samples at n=32768 and W=32 ( 200 k steps) for Gaussian denoising with σ=0.1 , 0.4 and 2 ; the noisy condition is clipped to [0,1] for display. Rows and columns as in Figure 16 .
Figure 18: Samples at n=32768 for W=16 , 32 and 64 ( 200 k steps) for super-resolution with the 8×8 condition. Rows and columns as in Figure 16 .
Figure 19: Samples at n=32768 for W=16 , 32 and 64 ( 200 k steps) for Gaussian denoising with σ=0.1 . Rows and columns as in Figure 16 .