Organizations: Department of Artificial Intelligence, School of Engineering, Westlake University · Zhejiang University · Research School of Astronomy and Astrophysics, Australian National University
Diffusion models generate data through a continuum of denoising problems, and are widely observed to reveal coarse structure before fine detail. Yet, this intuition is mostly empirical and qualitative. We introduce feature information dynamics, an information-theoretic framework for localizing when a feature is generated during diffusion. Using the I-MMSE identity, we connect the rate of feature mutual information change to a gap between optimal unconditional and feature-conditional denoising losses, yielding practical estimators for feature information density. We further develop a chained decomposition that separates shared from incremental information in a feature hierarchy. We use this framework first to quantitatively confirm spectral autoregression in pixel diffusion, and then to extend the analysis beyond frequency: under a class → mask → Canny conditioning chain, the per-feature information densities differ across pixel, SDVAE, VAVAE, and RAE, exposing fundamental differences between these representations and suggesting that ordered generation could be beneficial for training diffusion models. Our code is available at https://github.com/AI4Science-WestlakeU/feature-information-dynamics.
Figures & tables
Figure 1 : Pixel-space example of chained feature information dynamics. Top: a clean image X paired with a nested feature hierarchy of increasing spatial specificity — class label ( Y≤1 ), segmentation mask ( Y≤2 ), and Canny edges ( Y≤3 ). Conditioning on the cumulative bundle progressively constrains the diffusion model’s generated samples ( bottom row ): the unconditional model produces an arbitrary natural image; class fixes semantic identity; adding mask fixes object-level spatial support and pose; adding Canny additionally fixes local boundary structure. Right: at each SNR γ , conditioning on a richer bundle reduces mk(γ) ( top right ); the chained MMSE gap Δkchain in Eq. ( 17 ) ( middle right ) yields the per-level information density Dk∣k−1(logγ) ( bottom right ), which localizes where along the trajectory each feature contributes its incremental information. Curves peaking at distinct log-SNR values are the operational signature of a hierarchically organized denoising trajectory.
Figure 2 : Spectral information densities. In the top row of (a), single-band conditioning produces densities DYk(logγ) that overlap heavily because shared low-frequency information is counted in every band; in the bottom row of (a) and in (b), forward chaining removes information already explained by coarser bands, revealing a clean coarse-to-fine progression along logγ . Color: low (purple) to high (yellow) frequency band index k .
Figure 3 : Unconditional FID trajectories under a matched recipe. Best FID reached over 200 XL epochs: RAE 8.39 (ep. 180) < VAVAE 21.16< SDVAE 30.83< pixel 78.51 (all at ep. 200). Hyperparameters (αt,αs,lr) are selected on a smaller LDiT-B model and then frozen across the XL trajectories.
Figure 4 : Layer-wise feature information dynamics across representations. Per-feature log-SNR information densities Dk∣k−1(logγ) of the class, mask, and Canny increments for pixel, SDVAE, VAVAE, and RAE diffusion. Full panoramas in Figure 9 .
Model / checkpoint
Conditioning order
Class
Mask
Canny
RAE-S, 20 epochs
class–mask–Canny
−5.00
−3.50
−2.00
RAE-S, 40 epochs
class–mask–Canny
−3.75
−2.75
−1.50
RAE-S, 60 epochs
class–mask–Canny
−3.75
−2.75
−1.75
RAE-S, 80 epochs
class–mask–Canny
−3.75
−2.75
−1.75
RAE/DiTDH-XL
class–mask–Canny
−3.75
−2.75
−1.75
RAE/DiTDH-XL
mask–class–Canny
−4.00
−3.25
−1.75
Table 1: Peak log-SNRs for the model-size, checkpoint, and chain-order checks. RAE-S has approximately 132M parameters and the DiTDH-XL model approximately 842M.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Estimator
Advantage
Disadvantage
Single-feature MMSE gap
Directly reuses denoiser training; clear information-theoretic meaning
Requires two independent models; large-number subtraction amplifies variance
Denoiser gap
Equivalent to the single-feature MMSE gap but computed as a squared ℓ2 distance between two denoisers; numerically more stable
Still requires two denoiser models
Score gap
Natural interface with score-based models; equivalent to denoiser gap by Tweedie
Score training is less stable; the 1/γ factor amplifies errors at low SNR
Likelihood score
Only one model needed (a y -conditional likelihood-score predictor)
Requires a reverse-direction model; 1/γ singularity as γ→0
Appendix
Table 2: Comparison of estimators for the feature information density.
Figure 5 : Empirical finite-SNR summaries are close for many bands, but the analytic tail is not invariant to the feature-noise level α . CIFAR-10 multiband chained MMSE diagnostics (12 frequency bands) at α=0.5 (top row) and α=0 (bottom row), with all other settings held fixed. From left to right: per-band MMSE mk , chained gap Δkchain=mk−1−mk , normalized log- γ information density Dk∣k−1(logγ) , and cumulative chained information. The rows are close on most finite-SNR summaries, while the analytic high-SNR tail fit is unstable at α=0 (Table 3 ).
Figure 6 : Why the analytic Lorentzian estimator requires α>0 . Two representative CIFAR-10 bands at α=0.5 (blue) and α=0 (red): the mid band b5 , where the chained MMSE gap decays cleanly with γ , and the high band b10 , where the gap is small and noisy. Solid lines/markers show the empirical chained MMSE gap Δkchain(γ) and dashed lines show the corresponding double-pole fits L(γ)=a/(1+bγ)2 . On b5 , both fits track the empirical curve and the fitted scales agree to within 0.001 in μlg=log10(1/b) (Table 3 ). On b10 , the α=0.5 fit cleanly tracks the empirical decay, but the α=0 fit collapses to a near-constant baseline— b pinned at the optimizer lower bound 10−12 , apparent peak at μlg=+12 , far outside the data support which ends at +1.46 . This degeneracy is the analytic-estimator manifestation of the divergent I(Y;X) for deterministic continuous features.
Band
Setting
a
b
μlg=log10(1/b)
b5
α=0.5
3.65×101
7.51×10−1
+0.124
b5
α=0
3.68×101
7.50×10−1
+0.125
b10
α=0.5
1.37×10−1
3.23×10−2
+1.491
b10
α=0
6.01×10−2
1.00×10−12
+12.00
Appendix
Table 3: Double-pole Lorentzian fit parameters L(γ)=a/(1+bγ)2 for the two bands and two α settings shown in Figure 6 . Bold cells flag the α=0b10 degeneracy where b is pinned at the optimizer lower bound and the apparent peak location μlg is far outside the empirical log10γ support (which ends at +1.46 ).
Figure 7 : Double-pole Lorentzian fits to empirical single-feature MMSE gaps. Dots show empirical per-band single-feature MMSE gaps and black curves show K=1 double-pole fits. The vertical dashed line marks the adaptive reliable cutoff used for fitting. Despite the simplicity of the parametric form, the fits track the empirical gaps surprisingly well across bands, including the decay region that dominates the mutual-information integral.
Figure 8 : Detailed comparison of single-band and chained frequency decompositions on MNIST. The full diagnostic includes MMSE, MMSE gap or chained MMSE gap, log- γ information density, and cumulative information. The main text retains only the density row; the complete four-panel view is collected here for reference.
Representation
Model
Prediction
Learning rate
αt
αs
Pixel
LDiT-XL/16
x prediction
5×10−5
1.0
1.0
SDVAE
LDiT-XL/2
velocity
1×10−4
1.0
1.0
VAVAE
LDiT-XL/1
velocity
1×10−4
1.78
1.78
RAE
LDiT-XL/1
x prediction
2×10−4
0.10
0.10
Appendix
Table 4: Selected XL cells used for the representation learnability trajectories. RAE required an additional boundary-extension check beyond the original coarse shift grid, which selected αt=αs=0.10 .
Representation
Generator
Data space
Batch size
Conditioning stages
Pixel
JiT-L/16
raw images
1024
class, class+mask, class+mask+canny
SDVAE
SiT-XL/2
32×32×4 latents
256
class, class+mask, class+mask+canny
VAVAE
LDiT-XL/1
16×16×32 latents
1024
class, class+mask, class+mask+canny
RAE
DiTDH-XL
DINOv2-B RAE tokens
1024
class, class+mask, class+mask+canny
Appendix
Table 5: Backbone and data-space settings for the chained hierarchy measurements. Each representation is trained as a chain of three nested-condition stages—class, class+mask, class+mask+canny—initialized from the previous stage’s checkpoint. The conditioning artifacts are shared across representations; only the native representation encoder/backbone and the required optimizer scale differ.
Representation
Phase
Prior best val. loss
FiLM best val. loss
Δ
Δ (%)
Outcome
Pixel
mask
0.053708 (45 ep.)
0.053690 (51 ep.)
−0.000018
−0.03
tied
Pixel
canny
0.051556 (54 ep.)
0.051283 (66 ep.)
−0.000273
−0.53
FiLM slightly better
SDVAE
mask
0.719400 (42 ep.)
0.723878 (30 ep.)
+0.004478
+0.62
prior better
SDVAE
canny
0.692970 (60 ep.)
0.698930 (30 ep.)
+0.005960
+0.86
prior better
VAVAE
mask
0.274291 (30 ep.)
0.270818 (51 ep.)
−0.003473
−1.27
FiLM better
VAVAE
canny
0.268236 (15 ep.)
0.265441 (60 ep.)
−0.002795
−1.04
FiLM better
Appendix
Table 6: Validation-loss comparison between the prior conditioning implementation and a FiLM-style adaptive conditioning adapter for the mask and canny phases. Negative Δ means FiLM improves validation loss.
Figure 9 : Full chained-hierarchy panoramas across representations. Each row reports the same four diagnostics for one representation: MMSE curves, chained MMSE gaps, log-SNR information densities, and cumulative information for class, mask, and canny increments. The main text summarizes the density column in Figure 4 ; this appendix figure keeps the full diagnostic view for all four representations.
Diffusion models have become the dominant tool for high-fidelity image and video generation, yet are critically bottlenecked by their inference speed due to the numerous iterative passes of Diffusion Transformers. To reduce the exhaustive compute, recent works resort to the feature caching and reusing scheme that skips network evaluations at selected diffusion steps by using cached features in previous steps. However, their preliminary design solely relies on local approximation, causing errors to grow rapidly with large skips and leading to degraded sample quality at high speedups. In this work, we propose spectral diffusion feature forecaster (Spectrum), a training-free approach that enables global, long-range feature reuse with tightly controlled error. In particular, we view the latent features of the denoiser as functions over time and approximate them with Chebyshev polynomials. Specifically, we fit the coefficient for each basis via ridge regression, which is then leveraged to forecast features at multiple future diffusion steps. We theoretically reveal that our approach admits more favorable long-horizon behavior and yields an error bound that does not compound with the step size. Extensive experiments on various state-of-the-art image and video diffusion models consistently verify the superiority of our approach. Notably, we achieve up to 4.79× speedup on FLUX.1 and 4.67× speedup on Wan2.1-14B, while maintaining much higher sample quality compared with the baselines.
Despite the empirical success of Diffusion Models (DMs) and Variational Autoencoders (VAEs), their generalization performance remains theoretically underexplored, especially lacking a full consideration of the shared encoder-generator structure. Leveraging recent information-theoretic tools, we propose a unified theoretical framework that provides guarantees for the generalization of both the encoder and generator by treating them as randomized mappings. This framework further enables (1) a refined analysis for VAEs, accounting for the generator's generalization, which was previously overlooked; (2) illustrating an explicit trade-off in generalization terms for DMs that depends on the diffusion time T; and (3) providing computable bounds for DMs based solely on the training data, allowing the selection of the optimal T and the integration of such bounds into the optimization process to improve model performance. Empirical results on both synthetic and real datasets illustrate the validity of the proposed theory.
Qi Chen, Jierui Zhu, Florian Shkurti
Department of Computer Science, University of Toronto · Data Science Institute · Vector Institute +2
Diffusion models are widely used to generate high-quality images and videos, but their iterative denoising process remains computationally intensive. A growing class of training-free accelerators reduces this cost by reusing cached intermediate features or forecasting future ones. To control draft drift, these methods sometimes compute an exact block feature for verification. Yet the resulting exact feature is typically used only to measure discrepancy or guide a later decision and is then discarded. We find that this previously computed feature can instead be reused for correction. Forwarding it at the verification site resets the local draft residual and reduces downstream feature error. Based on this observation, we introduce FeatFix, a local exact-feature correction method for cached diffusion inference. FeatFix operates at a fixed sparse set of layer--timestep sites. At each selected site, it replaces the complete draft block output with the exact output computed from the same incoming state, avoiding token- or channel-level partial replacement and full-timestep recomputation. Experiments across four image and video backbones show that FeatFix consistently accelerates generation, achieving a speedup of up to 6.70× over Vanilla while maintaining competitive output quality.
Hanshuai Cui, Zhiqing Tang, Zhi Yao +3
School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China · Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai 519087, China