Organizations: Department of Artificial Intelligence, School of Engineering, Westlake University · Zhejiang University · Research School of Astronomy and Astrophysics, Australian National University
Diffusion models generate data through a continuum of denoising problems, and are widely observed to reveal coarse structure before fine detail. Yet, this intuition is mostly empirical and qualitative. We introduce feature information dynamics, an information-theoretic framework for localizing when a feature is generated during diffusion. Using the I-MMSE identity, we connect the rate of feature mutual information change to a gap between optimal unconditional and feature-conditional denoising losses, yielding practical estimators for feature information density. We further develop a chained decomposition that separates shared from incremental information in a feature hierarchy. We use this framework first to quantitatively confirm spectral autoregression in pixel diffusion, and then to extend the analysis beyond frequency: under a class → mask → Canny conditioning chain, the per-feature information densities differ across pixel, SDVAE, VAVAE, and RAE, exposing fundamental differences between these representations and suggesting that ordered generation could be beneficial for training diffusion models. Our code is available at https://github.com/AI4Science-WestlakeU/feature-information-dynamics.
Figures & tables
Figure 1 : Pixel-space example of chained feature information dynamics. Top: a clean image X paired with a nested feature hierarchy of increasing spatial specificity — class label ( Y≤1 ), segmentation mask ( Y≤2 ), and Canny edges ( Y≤3 ). Conditioning on the cumulative bundle progressively constrains the diffusion model’s generated samples ( bottom row ): the unconditional model produces an arbitrary natural image; class fixes semantic identity; adding mask fixes object-level spatial support and pose; adding Canny additionally fixes local boundary structure. Right: at each SNR γ , conditioning on a richer bundle reduces mk(γ) ( top right ); the chained MMSE gap Δkchain in Eq. ( 17 ) ( middle right ) yields the per-level information density Dk∣k−1(logγ) ( bottom right ), which localizes where along the trajectory each feature contributes its incremental information. Curves peaking at distinct log-SNR values are the operational signature of a hierarchically organized denoising trajectory.
Figure 2 : Spectral information densities. In the top row of (a), single-band conditioning produces densities DYk(logγ) that overlap heavily because shared low-frequency information is counted in every band; in the bottom row of (a) and in (b), forward chaining removes information already explained by coarser bands, revealing a clean coarse-to-fine progression along logγ . Color: low (purple) to high (yellow) frequency band index k .
Figure 3 : Unconditional FID trajectories under a matched recipe. Best FID reached over 200 XL epochs: RAE 8.39 (ep. 180) < VAVAE 21.16< SDVAE 30.83< pixel 78.51 (all at ep. 200). Hyperparameters (αt,αs,lr) are selected on a smaller LDiT-B model and then frozen across the XL trajectories.
Figure 4 : Layer-wise feature information dynamics across representations. Per-feature log-SNR information densities Dk∣k−1(logγ) of the class, mask, and Canny increments for pixel, SDVAE, VAVAE, and RAE diffusion. Full panoramas in Figure 9 .
Model / checkpoint
Conditioning order
Class
Mask
Canny
RAE-S, 20 epochs
class–mask–Canny
−5.00
−3.50
−2.00
RAE-S, 40 epochs
class–mask–Canny
−3.75
−2.75
−1.50
RAE-S, 60 epochs
class–mask–Canny
−3.75
−2.75
−1.75
RAE-S, 80 epochs
class–mask–Canny
−3.75
−2.75
−1.75
RAE/DiTDH-XL
class–mask–Canny
−3.75
−2.75
−1.75
RAE/DiTDH-XL
mask–class–Canny
−4.00
−3.25
−1.75
Table 1: Peak log-SNRs for the model-size, checkpoint, and chain-order checks. RAE-S has approximately 132M parameters and the DiTDH-XL model approximately 842M.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Estimator
Advantage
Disadvantage
Single-feature MMSE gap
Directly reuses denoiser training; clear information-theoretic meaning
Requires two independent models; large-number subtraction amplifies variance
Denoiser gap
Equivalent to the single-feature MMSE gap but computed as a squared ℓ2 distance between two denoisers; numerically more stable
Still requires two denoiser models
Score gap
Natural interface with score-based models; equivalent to denoiser gap by Tweedie
Score training is less stable; the 1/γ factor amplifies errors at low SNR
Likelihood score
Only one model needed (a y -conditional likelihood-score predictor)
Requires a reverse-direction model; 1/γ singularity as γ→0
Appendix
Table 2: Comparison of estimators for the feature information density.
Figure 5 : Empirical finite-SNR summaries are close for many bands, but the analytic tail is not invariant to the feature-noise level α . CIFAR-10 multiband chained MMSE diagnostics (12 frequency bands) at α=0.5 (top row) and α=0 (bottom row), with all other settings held fixed. From left to right: per-band MMSE mk , chained gap Δkchain=mk−1−mk , normalized log- γ information density Dk∣k−1(logγ) , and cumulative chained information. The rows are close on most finite-SNR summaries, while the analytic high-SNR tail fit is unstable at α=0 (Table 3 ).
Figure 6 : Why the analytic Lorentzian estimator requires α>0 . Two representative CIFAR-10 bands at α=0.5 (blue) and α=0 (red): the mid band b5 , where the chained MMSE gap decays cleanly with γ , and the high band b10 , where the gap is small and noisy. Solid lines/markers show the empirical chained MMSE gap Δkchain(γ) and dashed lines show the corresponding double-pole fits L(γ)=a/(1+bγ)2 . On b5 , both fits track the empirical curve and the fitted scales agree to within 0.001 in μlg=log10(1/b) (Table 3 ). On b10 , the α=0.5 fit cleanly tracks the empirical decay, but the α=0 fit collapses to a near-constant baseline— b pinned at the optimizer lower bound 10−12 , apparent peak at μlg=+12 , far outside the data support which ends at +1.46 . This degeneracy is the analytic-estimator manifestation of the divergent I(Y;X) for deterministic continuous features.
Band
Setting
a
b
μlg=log10(1/b)
b5
α=0.5
3.65×101
7.51×10−1
+0.124
b5
α=0
3.68×101
7.50×10−1
+0.125
b10
α=0.5
1.37×10−1
3.23×10−2
+1.491
b10
α=0
6.01×10−2
1.00×10−12
+12.00
Appendix
Table 3: Double-pole Lorentzian fit parameters L(γ)=a/(1+bγ)2 for the two bands and two α settings shown in Figure 6 . Bold cells flag the α=0b10 degeneracy where b is pinned at the optimizer lower bound and the apparent peak location μlg is far outside the empirical log10γ support (which ends at +1.46 ).
Figure 7 : Double-pole Lorentzian fits to empirical single-feature MMSE gaps. Dots show empirical per-band single-feature MMSE gaps and black curves show K=1 double-pole fits. The vertical dashed line marks the adaptive reliable cutoff used for fitting. Despite the simplicity of the parametric form, the fits track the empirical gaps surprisingly well across bands, including the decay region that dominates the mutual-information integral.
Figure 8 : Detailed comparison of single-band and chained frequency decompositions on MNIST. The full diagnostic includes MMSE, MMSE gap or chained MMSE gap, log- γ information density, and cumulative information. The main text retains only the density row; the complete four-panel view is collected here for reference.
Representation
Model
Prediction
Learning rate
αt
αs
Pixel
LDiT-XL/16
x prediction
5×10−5
1.0
1.0
SDVAE
LDiT-XL/2
velocity
1×10−4
1.0
1.0
VAVAE
LDiT-XL/1
velocity
1×10−4
1.78
1.78
RAE
LDiT-XL/1
x prediction
2×10−4
0.10
0.10
Appendix
Table 4: Selected XL cells used for the representation learnability trajectories. RAE required an additional boundary-extension check beyond the original coarse shift grid, which selected αt=αs=0.10 .
Representation
Generator
Data space
Batch size
Conditioning stages
Pixel
JiT-L/16
raw images
1024
class, class+mask, class+mask+canny
SDVAE
SiT-XL/2
32×32×4 latents
256
class, class+mask, class+mask+canny
VAVAE
LDiT-XL/1
16×16×32 latents
1024
class, class+mask, class+mask+canny
RAE
DiTDH-XL
DINOv2-B RAE tokens
1024
class, class+mask, class+mask+canny
Appendix
Table 5: Backbone and data-space settings for the chained hierarchy measurements. Each representation is trained as a chain of three nested-condition stages—class, class+mask, class+mask+canny—initialized from the previous stage’s checkpoint. The conditioning artifacts are shared across representations; only the native representation encoder/backbone and the required optimizer scale differ.
Representation
Phase
Prior best val. loss
FiLM best val. loss
Δ
Δ (%)
Outcome
Pixel
mask
0.053708 (45 ep.)
0.053690 (51 ep.)
−0.000018
−0.03
tied
Pixel
canny
0.051556 (54 ep.)
0.051283 (66 ep.)
−0.000273
−0.53
FiLM slightly better
SDVAE
mask
0.719400 (42 ep.)
0.723878 (30 ep.)
+0.004478
+0.62
prior better
SDVAE
canny
0.692970 (60 ep.)
0.698930 (30 ep.)
+0.005960
+0.86
prior better
VAVAE
mask
0.274291 (30 ep.)
0.270818 (51 ep.)
−0.003473
−1.27
FiLM better
VAVAE
canny
0.268236 (15 ep.)
0.265441 (60 ep.)
−0.002795
−1.04
FiLM better
Appendix
Table 6: Validation-loss comparison between the prior conditioning implementation and a FiLM-style adaptive conditioning adapter for the mask and canny phases. Negative Δ means FiLM improves validation loss.
Figure 9 : Full chained-hierarchy panoramas across representations. Each row reports the same four diagnostics for one representation: MMSE curves, chained MMSE gaps, log-SNR information densities, and cumulative information for class, mask, and canny increments. The main text summarizes the density column in Figure 4 ; this appendix figure keeps the full diagnostic view for all four representations.
Jul 30, 2026·Hanshuai Cui, Zhiqing Tang, Zhi Yao +3DrafterCorrection
School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China · Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai 519087, China