Organizations: School of Mathematics, University of Minnesota · NVIDIA · Department of Mathematics, The Ohio State University · Zu Chongzhi Center and DIRC, Duke Kunshan University
Diffusion models often degrade in latent spaces, yet the formal causes remain poorly understood. We quantify latent-space diffusability via the rate of change of the Minimum Mean Squared Error (MMSE) along the diffusion trajectory. Our framework decomposes this MMSE rate into contributions from Fisher Information (FI) and Fisher Information Rate (FIR). We show that isometric embeddings preserve intrinsic FI and establish quantitative intrinsic-FI bounds for a broader class of bi-Lipschitz encoders with controlled weak volume distortion, whereas FIR is governed by the interplay between encoder and data geometries. Our analysis separates four geometric contributions in local stability bounds for Gaussian-smoothed FIR: dimensional compression, tangential distortion, high-frequency encoder curvature, and curvature of data manifold. Experiments across diverse autoencoding architectures provide qualitative support for the geometric mechanisms identified by the theory and show that empirical FI and FIR track several measures of generation quality and latent-space geometry in the settings tested. We establish FI and FIR as a comprehensive analytical framework for understanding latent diffusability.
Figures & tables
Figure 1: Values of I (left) and R (right) plotted against the noise variance τ . (a,b) Gaussian toy experiments using tiny diffusion models. Pixel denotes models trained on x∼N(0,I2) , while latent denotes models trained on z=E(x) , with the pointwise activation E indicated in the legend. (c,d) Results for recent image generative models. JiT operates directly in pixel space, whereas SiT, RAE, RAEv2, and REPA-E operate in latent space. For (c,d), we show τ∈[0.01,80] , excluding smaller noise levels because of numerical instability.
Figure 2: FIR deviation DR vs. (a) δ0 , (b) d , and (c) ε0 in toy settings. Data y∼N(0,I2) are embedded as x=(y1,y2,0,…,0)∈RD and encoded as z=E(x)∈Rd . We compute DR from R(D)(μτ) and R(d)((μZ)τ) using diffusion models trained on x and z . Curves correspond to fixed noise variance τ . The solid lines y=1.25δ0 in (a) and y=(D−d)/(D−2) in (b) serve as linear references. In (a), D=d=2 and E(x)=Ax , where A=diag(1+δ0,1−δ0) . In (b), D=512 and z=(y1,y2,0,…,0)∈Rd . In (c), D=d=3 and E((x1,x2,0),ε0)=(sin(ε0x1)/ε0,x2,(1−cos(ε0x1))/ε0) .
Generation Metrics
Geometric and Spectral Metrics
Model
FID ↓
sFID ↓
IS ↑
Precision ↑
Recall ↑
Straightness ↑
Efficiency ↑
SEC 0.25↓
SEC 0.5↓
RAE
2.432
8.340
243.794
0.720
0.640
0.999
0.982
0.277
0.112
RAEv2
2.370
7.458
225.068
0.730
0.633
0.999
0.949
0.309
0.137
SiT
7.945
9.442
123.962
0.594
0.678
1.000
0.851
0.411
0.209
REPA-E
3.496
7.973
159.359
0.692
0.630
1.000
0.875
0.312
0.102
JiT
3.169
8.449
65.185
0.692
0.631
–
–
–
–
Table 1: Generation, geometric, and spectral evaluation results. Generation quality is evaluated using FID, spatial FID (sFID), Inception Score (IS), precision, and recall. Straightness and efficiency measure the geometry of ODE sampling trajectories, while SEC measures the fraction of latent spectral energy outside the low-frequency region, with subscripts indicating the frequency threshold. Geometric and spectral metrics are reported only for the latent diffusion models.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Geometric interpretation of encoder-induced distortions.
Figure 4: Values of (c) I and (d) R plotted versus the noise variance τ for the ReLU-based experiments. Pixel curves correspond to models trained on x∼N(0,I2) , while latent curves correspond to models trained on z=E(x) . For Leaky ReLU, α denotes the negative slope.
Figure 5: FIR deviation DR vs. (a) τ and (b) δ0 in toy settings. Data y∼N(0,I2) are embedded as x=(y1,y2,0,…,0)∈RD and encoded to z=E(x)∈Rd . We consider D=d=2 and E(x)=Ax with A=diag(1+δ0,1−δ0) . We compute DR from R(D)(μτ) and R(d)((μZ)τ) using diffusion models trained on x and z . Solid line y=1.25δ0 in (b) serves as a linear reference.
Figure 6: FIR deviation DR vs. (a) τ and (b) d in toy settings. Data y∼N(0,I2) are embedded as x=(y1,y2,0,…,0)∈RD with D=512 , and z=(y1,y2,0,…,0)∈Rd . We compute DR from R(D)(μτ) and R(d)((μZ)τ) using diffusion models trained on x and z . Solid line y=D−2D−d in (b) serves as a linear reference.
Figure 7: FIR deviation DR vs. (a) τ and (b) ε0 in toy settings. Data y∼N(0,I2) are embedded as x=(y1,y2,0,…,0)∈RD and encoded to z=E(x)∈Rd . We consider D=d=3 and E((x1,x2,0),ε0)=(sin(ε0x1)/ε0,x2,(1−cos(ε0x1))/ε0) . We compute DR from R(D)(μτ) and R(d)((μZ)τ) using diffusion models trained on x and z .
Figure 8: Values of (a) I and (b) R plotted versus the noise variance τ , computed from diffusion models trained on different data representations. The pixel curves correspond to models trained directly on FFHQ images. The latent curves correspond to models trained on latent representations of an image encoder (VAE or KL-AE) pretrained on FFHQ. We show τ∈[0.01,80] , excluding smaller τ due to numerical instability.
Figure 9: Dimension-normalized FIR deviation DRsc vs. noise variance τ on FFHQ. Here, D=64×64×3 , m=20 , and the latent dimension is d=256 for the VAE and d=3×16×16 for the KL-AE. FFHQ images x are encoded as latents z=E(x) using the VAE or KL-AE.
Figure 10: (a) Samples from a diffusion model trained directly on 64×64 FFHQ. (b, c) Samples from latent diffusion models trained on (b) VAE and (c) KL-AE representations. Generated latents are mapped back to image space using their respective decoders.
Model
Reconstruction MSE
VAE
0.0094
KL-AE
0.41
Appendix
Table 2: Reconstruction MSE (mean squared error) for the VAE and KL-AE trained on FFHQ.
Model
FID vs. FFHQ
FID vs. Reconstructions
Pixel Diffusion
2.56
–
VAE (Latent Diffusion)
62.51
–
KL-AE (Latent Diffusion)
14.97
3.58
Appendix
Table 3: Generation quality on FFHQ. FID (Fréchet Inception Distance) is computed using two reference distributions: the original FFHQ images and the reconstructions produced by the corresponding autoencoder. The first FID column measures generation quality relative to the original image distribution, while the second measures generation quality relative to the reconstruction distribution of the corresponding latent model. For the pixel model, images are generated by sampling a diffusion model trained directly on FFHQ. For the VAE and KL-AE models, images are generated by sampling diffusion models trained in the latent space and decoding the generated latents using the corresponding decoder. All diffusion models use the U-Net architecture within the EDM framework.
Model
c^
C^
C^/c^
VAE
0.20
3.86
19.67
SD-VAE
1.60
4.83
3.02
Appendix
Table 4: Empirical estimates (c^,C^) of the constants (c,C) in Proposition B.1 for encoders of the VAE and SD-VAE trained on FFHQ. The ratio C^/c^ provides an empirical estimate of the bi-Lipschitz constant.
Figure 11: Sensitivity of MMSE to τ , plotted against the noise variance τ , computed from diffusion models trained on different data representations. The pixel curve corresponds to a model trained directly on FFHQ images, while the latent curves correspond to models trained on VAE and KL-AE latent representations. We show τ∈[0.01,80] , excluding smaller noise levels due to numerical instability. For visualization on a logarithmic scale, we plot dτdMMSE+ε with ε=100 .
Figure 12: Values of (a) I and (b) R plotted versus the noise variance τ for diffusion models trained on different data representations. Pixel and latent curves denote models trained on FFHQ images and their pretrained NVAE latents, respectively. NVAE was pretrained on FFHQ, with spatial size 20×dz×dz . We show τ∈[0.01,80] , excluding smaller τ due to numerical instability.
Figure 13: (a) Power spectrum of FFHQ images. (b)(c) Power spectra of NVAE latent representations with spatial resolutions dz=16 and 8 . (d)(e) Power spectra of VAE and KL-AE latent representations. The NVAE, VAE, and KL-AE models are pretrained on FFHQ images. For FFHQ images, NVAE latents, and KL-AE latents, the spectrum is computed using the two-dimensional FFT over the spatial dimensions and averaged over channels. For the one-dimensional latent vectors produced by the VAE encoder, the spectrum is computed using a one-dimensional FFT along the latent dimension. The spectra are averaged over 10,000 samples drawn from the dataset. Both the frequency axis k and the power values P are normalized by the signal length (or spatial resolution for the 2D case). The zero-frequency (DC) component corresponding to the mean signal level is excluded.