Generative denoising models, such as diffusion and flow-matching, learn to sample from complex distributions by training a deep neural network denoiser to recover clean data from noise-corrupted samples. While such models are typically compared on the quality of their synthesized samples, these metrics provide limited insight into how the underlying denoiser, which drives generation, differs. In this work, we propose to analyze the spectrum of the denoiser Jacobian as a tool to characterize these differences. Across pre-trained denoising models, we observe that better generative performance is associated with larger Jacobian eigenvalues. Motivated by this, we introduce a regularization scheme that controls the Jacobian spectrum by training the denoiser on perturbed inputs, with perturbations suppressing or amplifying Jacobian responses. On ImageNet, we test whether directly modifying the Jacobian spectral properties leads to improved generations. Our findings suggest that denoisers benefit from both strengthening responses along data-relevant principal eigen-directions and suppressing the noisy, data-irrelevant ones. This establishes the denoiser Jacobian as a useful tool for identifying differences between generative denoising models.
Figures & tables
Figure 1: Using Algorithm 1 ( n=10 , K=10 ), we compute eigenvalues of Jacobians for different SiT denoisers. Top : We observe a distinct ordering that aligns with generative performance (SiT-S < SiT-XL < SiT-XL + REPA), whereas denoising capabilities are indistinguishable. Bottom : When using classifer-free guidance with w=4.0 the differences between the models are accentuated.
Figure 2: Examples of a denoised image and the eigenvectors computed around it for pre-trained SiT models. We show the top-10 eigenvectors for the same image at t=0.4 . Better generative models capture more and sharper variability in their top components.
Figure 3: Training a denoising generative model on a mixture of 2D Gaussians. The grayscale color represents the maximum eigenvalue of the Jacobian at t=0.5 for each point on the grid. (a) Baseline model trained without regularization. (b) Jacobian regularization using a perturbation that minimizes variation along the orthogonal axes. (c) Jacobian regularization using the residual perturbation, which increases eigenvalues (showing λmax ), and results in fewer samples falling between modes.
Figure 4: Eigenvalue analysis for SiT-S models trained with residual regularization. By varying the target gain 1/δ , we impose larger eigenvalues, with diminishing effects at 1/δ=10 . Using classifier-free guidance with w=4.0 (bottom) amplifies the differences between the models.
Figure 5: Eigenvalue analysis for SiT-S models trained with stochastic regularization. Using τ=0.1 does not alter the top-10 eigenvalues of the trained model. In contrast, with τ=1.0 , the trained model is over-constrained, exhibiting smaller eigenvalues than the baseline model when using guidance with scale w=4.0 (bottom). We also include the model combining both regularization signals ( 1/δ=5 , τ=0.1 ), which we show obtains similar spectra to the residual-only regularized models.
Figure 6
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Using Algorithm 1 , we measure eigenvalues of the Jacobians of different flow-matching denoisers. We find an ordering that correlates with the model’s expected performance (SiT-S < SiT-XL < SiT-XL + REPA). Top : Using n=20 and K=20 . Bottom : Using classifier-free guidance with scale w=4.0 .
Figure 8: Comparing the spectrum of the top-100 eigenvalues of SiT-S (solid) and SiT-S with residual regularization (dashed).
Figure 9: SiT-B and UNet denoiser eigenvalue comparison between baseline and models trained with the proposed residual regularization.
Figure 10
Figure 11: We synthesize images from the same noise, using 50 steps and classifier-free guidance scale w=4.0 to amplify differences. We observe that the model using the residual perturbation performs larger corrections to the baseline-generated images, indicating a stronger overall effect.
Figure 12: We compare trained model eigenvalues when using the masked and full residual as the perturbation direction. Top : Comparison for n=10 , K=10 . Bottom : Same comparison using classifier-free guidance with scale w=4.0 .
Figure 13
Figure 15: Training a denoising generative model on a mixture of 2D Gaussians. The grayscale color represents the maximum eigenvalue of the Jacobian at t=0.5 for each point on a 2D grid. (a) Baseline model trained without regularization. (b) Jacobian regularization using the residual perturbation, pushing λmax higher. (c) Direct regularization (Appendix D ) maximizing all eigenvalues simultaneously. The resulting model is overly sensitive and fails to learn the target distribution.
Figure 16: Gain and Rayleigh quotient R during training for Jacobian-regularized SiT-S/B models. The gain overestimates the effect of the residual perturbation gain. We use R to set the gain target during training.
Figure 17: Examples of a denoised image and the eigenvectors computed around it: (a) Different training iterations in the SiT-B + residual regularization model. As training progresses, the model captures wider and sharper variations in its top components. (b) The SiT-B model with and without residual regularization. The residual-regularized model improves the spectrum by making larger changes in the top components.
Figure 18: (a) Eigenvectors at t=0.4 , using classifier-free guidance with scale w=4.0 . Increasing guidance leads to amplified differences between the eigenvectors and larger eigenvalues. (b) Eigenvectors at t=0.2 without classifier-free guidance. The variations in lower timesteps capture larger-scale structures in the images.
Figure 19: Feature comparison (reduced to RGB using PCA) between the baseline and the Jacobian-regularized SiT-B. At low timesteps ( t={0.3,0.5} ), the image structures emerge in earlier blocks in the Jacobian-regularized model.
Figure 20: Examples of images generated with the baseline and Jacobian-regularized SiT-B models. We use the Euler sampler with 50 inference steps and guidance scale w=4.0 .