Generative denoising models, such as diffusion and flow-matching, learn to sample from complex distributions by training a deep neural network denoiser to recover clean data from noise-corrupted samples. While such models are typically compared on the quality of their synthesized samples, these metrics provide limited insight into how the underlying denoiser, which drives generation, differs. In this work, we propose to analyze the spectrum of the denoiser Jacobian as a tool to characterize these differences. Across pre-trained denoising models, we observe that better generative performance is associated with larger Jacobian eigenvalues. Motivated by this, we introduce a regularization scheme that controls the Jacobian spectrum by training the denoiser on perturbed inputs, with perturbations suppressing or amplifying Jacobian responses. On ImageNet, we test whether directly modifying the Jacobian spectral properties leads to improved generations. Our findings suggest that denoisers benefit from both strengthening responses along data-relevant principal eigen-directions and suppressing the noisy, data-irrelevant ones. This establishes the denoiser Jacobian as a useful tool for identifying differences between generative denoising models.
Figures & tables
Figure 1: Using Algorithm 1 ( n=10 , K=10 ), we compute eigenvalues of Jacobians for different SiT denoisers. Top : We observe a distinct ordering that aligns with generative performance (SiT-S < SiT-XL < SiT-XL + REPA), whereas denoising capabilities are indistinguishable. Bottom : When using classifer-free guidance with w=4.0 the differences between the models are accentuated.
Figure 2: Examples of a denoised image and the eigenvectors computed around it for pre-trained SiT models. We show the top-10 eigenvectors for the same image at t=0.4 . Better generative models capture more and sharper variability in their top components.
Figure 3: Training a denoising generative model on a mixture of 2D Gaussians. The grayscale color represents the maximum eigenvalue of the Jacobian at t=0.5 for each point on the grid. (a) Baseline model trained without regularization. (b) Jacobian regularization using a perturbation that minimizes variation along the orthogonal axes. (c) Jacobian regularization using the residual perturbation, which increases eigenvalues (showing λmax ), and results in fewer samples falling between modes.
Figure 4: Eigenvalue analysis for SiT-S models trained with residual regularization. By varying the target gain 1/δ , we impose larger eigenvalues, with diminishing effects at 1/δ=10 . Using classifier-free guidance with w=4.0 (bottom) amplifies the differences between the models.
Figure 5: Eigenvalue analysis for SiT-S models trained with stochastic regularization. Using τ=0.1 does not alter the top-10 eigenvalues of the trained model. In contrast, with τ=1.0 , the trained model is over-constrained, exhibiting smaller eigenvalues than the baseline model when using guidance with scale w=4.0 (bottom). We also include the model combining both regularization signals ( 1/δ=5 , τ=0.1 ), which we show obtains similar spectra to the residual-only regularized models.
Figure 6
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Using Algorithm 1 , we measure eigenvalues of the Jacobians of different flow-matching denoisers. We find an ordering that correlates with the model’s expected performance (SiT-S < SiT-XL < SiT-XL + REPA). Top : Using n=20 and K=20 . Bottom : Using classifier-free guidance with scale w=4.0 .
Figure 8: Comparing the spectrum of the top-100 eigenvalues of SiT-S (solid) and SiT-S with residual regularization (dashed).
Figure 9: SiT-B and UNet denoiser eigenvalue comparison between baseline and models trained with the proposed residual regularization.
Figure 10
Figure 11: We synthesize images from the same noise, using 50 steps and classifier-free guidance scale w=4.0 to amplify differences. We observe that the model using the residual perturbation performs larger corrections to the baseline-generated images, indicating a stronger overall effect.
Figure 12: We compare trained model eigenvalues when using the masked and full residual as the perturbation direction. Top : Comparison for n=10 , K=10 . Bottom : Same comparison using classifier-free guidance with scale w=4.0 .
Figure 13
Figure 15: Training a denoising generative model on a mixture of 2D Gaussians. The grayscale color represents the maximum eigenvalue of the Jacobian at t=0.5 for each point on a 2D grid. (a) Baseline model trained without regularization. (b) Jacobian regularization using the residual perturbation, pushing λmax higher. (c) Direct regularization (Appendix D ) maximizing all eigenvalues simultaneously. The resulting model is overly sensitive and fails to learn the target distribution.
Figure 16: Gain and Rayleigh quotient R during training for Jacobian-regularized SiT-S/B models. The gain overestimates the effect of the residual perturbation gain. We use R to set the gain target during training.
Figure 17: Examples of a denoised image and the eigenvectors computed around it: (a) Different training iterations in the SiT-B + residual regularization model. As training progresses, the model captures wider and sharper variations in its top components. (b) The SiT-B model with and without residual regularization. The residual-regularized model improves the spectrum by making larger changes in the top components.
Figure 18: (a) Eigenvectors at t=0.4 , using classifier-free guidance with scale w=4.0 . Increasing guidance leads to amplified differences between the eigenvectors and larger eigenvalues. (b) Eigenvectors at t=0.2 without classifier-free guidance. The variations in lower timesteps capture larger-scale structures in the images.
Figure 19: Feature comparison (reduced to RGB using PCA) between the baseline and the Jacobian-regularized SiT-B. At low timesteps ( t={0.3,0.5} ), the image structures emerge in earlier blocks in the Jacobian-regularized model.
Figure 20: Examples of images generated with the baseline and Jacobian-regularized SiT-B models. We use the Euler sampler with 50 inference steps and guidance scale w=4.0 .
Diffusion models have shown remarkable performance on diverse generation tasks. Recent work finds that imposing representation alignment on the hidden states of diffusion networks can both facilitate training convergence and enhance sampling quality, yet the mechanism driving this synergy remains insufficiently understood. In this paper, we investigate the connection between self-supervised spectral representation learning and diffusion generative models through a shared perspective on perturbation kernels. On the diffusion side, samples (e.g., images, videos) are produced by reversing a stochastic noise-injection process specified by Gaussian kernels; on the spectral representation side, spectral embeddings emerge from contrasting positive and negative relations induced by random perturbation kernels. Motivated by this, we propose a self-supervised spectral representation alignment method to facilitate diffusion model training. In addition, we clarify how joint spectral learning can benefit diffusion training from a geometric perspective. Furthermore, we find that the optimization of the spectral alignment objective is in an equivalent form of diffusion score distillation in the representation space. Building on these findings, we integrate a spectral regularizer into diffusion training objectives to improve the performance of diffusion models on multiple datasets. Experiments across images and 3D point clouds show consistent gains in generation quality. Code is released at https://github.com/yuehaowang/spectral-reg-diffusion.
Yuehao Wang, Peihao Wang, Hanwen Jiang +3
University of Texas at Austin, USA · 2Adobe Research, USA
Pixel-space diffusion models are trained on full-bandwidth noisy images, yet the useful signal available to the denoiser is strongly frequency dependent. Under rectified-flow diffusion and natural-image power-law spectra, the per-band data-to-noise contour k∗(t)=(1−t)−2/α separates a signal-bearing low-frequency region from a noise-dominated high-frequency region at each time t. We show that this implicit coarse-to-fine structure is not merely descriptive: it induces a capacity-allocation problem. A standard pixel-space denoiser must discover the moving bandwidth boundary internally and can spend computation on frequency-time regions where the optimal prediction collapses to deterministic baselines rather than data-distribution modeling. To make this boundary explicit, we introduce Spectral Forcing, a parameter-free, time-conditional 2D-DCT low-pass operator applied to the noisy input before the patch embedder. Its cutoff expands monotonically with the diffusion time and becomes the identity at the data endpoint. Through controlled synthetic experiments, we identify the regime in which the operator is beneficial: coarse patch tokenization and data whose high-frequency content is predominantly noise rather than essential signal. On ImageNet-256 with JiT-700M/32, Spectral Forcing consistently improves both FID and Inception Score across different training epochs, demonstrating robust gains throughout training; at finer tokenization, the spectral forcing is still competitive. We further insert the unchanged operator into SenseNova-U1, a unified text-to-image model, where it improves DPG-Bench and GenEval, showing that the input-side spectral prior transfers beyond class-conditional generation. These results suggest a route to capacity-efficient pixel-space diffusion by showing the signal and hiding the noise.
Recent work has shown that diffusion models trained with the denoising score matching (DSM) objective often violate the Fokker--Planck (FP) equation that governs the evolution of the true data density. Directly penalizing these deviations in the objective function reduces their magnitude but introduces a significant computational overhead. It is also observed that enforcing strict adherence to the FP equation does not necessarily lead to improvements in the quality of the generated samples, as often the best results are obtained with weaker FP regularization. In this paper, we investigate whether simpler penalty terms can provide similar benefits. We empirically analyze several lightweight regularizers, study their effect on FP residuals and generation quality, and show that the benefits of FP regularization are available at substantially lower computational cost. Our code is available at https://github.com/OnnoNiemann/fp_diffusion_analysis.
Onno Niemann, Gonzalo Martínez Muñoz, Alberto Suárez Gonzalez