Residual-Stream Burden Shapes Representation Learning in Diffusion Transformers
Organizations: UC San Diego · Shanghai Jiao Tong University · Aether AI
Abstract
In diffusion-based generation, a neural network can be trained to predict the clean data, the noise, or the velocity from a noisy input. These prediction targets are interconvertible and describe the same generative process, yet plain Diffusion Transformers operating on large pixel patches succeed with clean prediction and fail with noise or velocity prediction. We argue that this asymmetry arises because noisy targets require the residual stream to preserve noise-dependent input variation through depth for the final readout, forcing subsequent layers to compute on noisy representations. A spectrally concentrated clean target imposes a lighter demand, leaving greater freedom to organize hidden representations for subsequent computation. We call this preservation requirement residual-stream burden and show how it shapes representation learning in Diffusion Transformers. Controlled experiments indicate that the exploitable structure is spectral concentration in patch space and that the bandwidth of the persistent residual state is a key resource for noisy prediction. We further show that this account is consistent with recent decoupled pixel-space architectures, whose diverse designs all reduce the residual-stream burden on the main pathway. To examine this understanding from a complementary direction, we expand and reorganize the residual-stream bandwidth directly, introducing Spatially Indexed Hyper-Connections (SiHC) that reach FID 1.71 on ImageNet . Together, these results identify residual-stream burden as a mechanism through which prediction targets and architecture jointly shape representation learning in Diffusion Transformers.
Figures & tables
| Hierarchical design | Decoupling via long skip | Decoupling via cross-attention |
| ADM ( Dhariwal and Nichol, 2021 ) , SiD / SiD2 ( Hoogeboom et al., 2023 ; Hoogeboom et al., 2025 ) , PixelFlow ( Chen et al., 2025b ) | PixNerd ( Wang et al., 2026a ) , DiP ( Chen et al., 2026b ) , DeCo ( Ma et al., 2026 ) , PixelDiT ( Yu et al., 2026 ) | RIN ( Jabri et al., 2023 ) , HyperDiT ( He et al., 2026 ) , DuSPiT ( Bai et al., 2026 ) |
| Prediction target | Clean | Velocity | Clean |
| Data space | Pixel | Pixel | Whitened |
| FID | 10.19 | 139.83 | 132.95 |
| Effective rank | 4.80 | 367.37 | 768 |
| Stable rank | 1.41 | 5.94 | 768 |
| variance rank | 8 | 668 | 692 |
| Prediction | Plain FID | mHC FID |
| Pixel, | 10.19 | 9.78 |
| Pixel, | 139.83 | 25.36 |
| DINOv2-B, | 5.50 | 3.93 |
| DINOv2-B, | 16.56 | 3.98 |
| Design | What changes | State region | Streams | FID |
| Plain DiT | 1 | 139.83 | ||
| mHC, copied input | Four persistent states | 4 | 25.36 | |
| mHC, spatial input | Local input and prediction per state | 4 | 13.25 | |
| + identity carry, scalar maps | Static maps without state mixing | 4 | 11.50 | |
| + feature-wise maps (SiHC) | Channel-specific read and write | 4 | 10.10 | |
| + subpatches | More states, smaller regions | 16 | 8.07 |
| Model | Prediction | Params (M) | GFLOPs | FID | IS |
| JiT-H/16 ( Li and He, 2026 ) | 953 | 364 | 1.86 | 303.4 | |
| JiT-G/16 ( Li and He, 2026 ) | 2000 | 767 | 1.82 | 292.6 | |
| PixelREPA-H/16 ( Shin et al., 2026 ) | 953 | 364 | 1.81 | 317.2 | |
| PixNerd-XL/16 † ( Wang et al., 2026a ) | 700 | 268 | 1.89 | 309.0 | |
| DiP-XL/16 † ( Chen et al., 2026b ) | 631 | 255 | 1.75 | 273.0 | |
| DeCo-XL/16 † ( Ma et al., 2026 ) | 682 | 245 | 1.69 | 304.0 |
Appendix figures & tables41 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Noisy-input path | Connection to backbone | Prediction computation |
| Analytic output paths | |||
| JiT [ Li and He, 2026 ] | Noisy input enters both the patch embedding and an analytic output skip. | Final patch features predict the clean image. | Linear patch head, then analytic conversion to velocity. |
| AsymFlow [ Chen et al., 2026a ] | An analytic input path recovers velocity outside a chosen noise subspace. | A JiT backbone predicts full-dimensional data minus low-rank noise. | Patch head with projection-based analytic velocity recovery. |
| Decoders conditioned on final backbone features | |||
| DiP [ Chen et al., 2026b ] | Noisy RGB patches enter local U-Nets. | Context tokens enter U-Net bottlenecks. | Patch-local convolutions with down/up blocks and skips. |
| DeCo [ Ma et al., 2026 ] | Noisy RGB and positions form pixel queries. | Upsampled coarse features supply AdaLN-zero conditioning. | Pixel-wise residual MLPs without attention. |
| Model | Params (M) | GFLOPs | FID | IS |
| Latent diffusion with VAE | ||||
| DiT-XL/2 [ Peebles and Xie, 2023 ] | 675 | 238 | 2.27 | 278.2 |
| SiT-XL/2 [ Ma et al., 2024 ] | 675 | 238 | 2.06 | 277.5 |
| SiT-XL/2 + REPA [ Yu et al., 2025 ] | 675 | 238 | 1.42 | 305.7 |
| LightningDiT-XL/2 [ Yao et al., 2025 ] | 675 | 238 | 1.35 | 295.3 |
| DDT-XL/2 [ Wang et al., 2026c ] | 675 | 238 | 1.26 | 310.6 |
| Paper | 0.1 | 0.3 | 0.5 | 0.7 | 0.75 | 0.9 |
| Affine clean | 1 | 2 | 3 | 5 | 5 | 7 |
| Affine velocity | 674 | 656 | 638 | 603 | 588 | 488 |
| Gain share (%) | Pixel, | Pixel, | DeCo | PixelDiT | DiP | HyperDiT |
| Top 256 | 41.7 | 95.3 | 96.9 | 98.0 | 86.5 | 97.9 |
| Bottom 256 | 28.0 | 0.2 | 0.2 | 0.3 | 5.4 | 0.3 |
| Model | |||||
| Pixel, | 0.903 | 0.925 | 0.921 | 0.909 | 0.949 |
| Pixel, | 0.076 | 0.095 | 0.268 | 0.530 | 0.652 |
| DeCo-B | 0.948 | 0.946 | 0.930 | 0.927 | 0.929 |
| PixelDiT-B | 0.920 | 0.877 | 0.885 | 0.891 | 0.935 |
| DiP-B | 0.924 | 0.929 | 0.884 | 0.726 | 0.770 |
| HyperDiT-B | 0.925 | 0.867 | 0.874 | 0.891 | 0.912 |
| Target | ER | SR | |||
| DINO, | 579.6 | 66.84 | 563 | 649 | 735 |
| DINO, | 708.3 | 121.31 | 651 | 708 | 755 |
| Setting | Value |
| Covariance fitting set | 100,000 ImageNet training images, 100 per class |
| Selection / flip seed | 20260806 / 20260807 |
| Preprocessing | ADM center crop to , RGB in ; one deterministic horizontal-flip draw per selected image |
| Patch coordinates | Non-overlapping RGB; channel, patch row, patch column order |
| Observations | M patches of dimension 768 |
| Centering | One global 768-dimensional mean over all patches |
| Epoch | Clean FID | Clean IS | Velocity FID | Velocity IS |
| 40 | 298.46 | 2.60 | 238.82 | 2.85 |
| 80 | 166.89 | 5.82 | 243.77 | 2.29 |
| 120 | 148.95 | 7.64 | 204.79 | 2.55 |
| 160 | 140.39 | 8.61 | 209.37 | 3.01 |
| 200 | 132.95 | 9.54 | 212.54 | 3.35 |
| Setting | Value |
| Weights | Raw epoch 200 / step 250,200 |
| Images | 2,048 ImageNet training images, uniform sampling without replacement |
| Corruption | 128 Gaussian draws per image at each |
| Pairing | Identical image indices and deterministic Gaussian draws across models/times |
| Input / conditioning | Pixel or whitened training coordinates (Appendix B.9 ); null class 1000; no image flips |
| Observation | Attention/MLP input before normalization and adaptive modulation |
| Model | |||||||||
| Pixel, | 0.164 | 0.081 | 1.04 | 0.133 | 0.142 | 1.89 | 0.071 | 0.221 | 3.90 |
| Pixel, | 0.116 | 0.137 | 1.50 | 0.045 | 0.192 | 4.39 | 0.014 | 0.234 | 17.69 |
| DeCo | 0.090 | 0.098 | 1.41 | 0.035 | 0.149 | 4.44 | 0.010 | 0.196 | 19.07 |
| PixelDiT | 0.086 | 0.197 | 2.78 | 0.032 | 0.224 | 7.87 | 0.010 | 0.241 | 28.83 |
| DiP | 0.083 | 0.121 | 1.72 | 0.049 | 0.162 | 3.56 | 0.036 | 0.182 | 11.24 |
| Setting | Value |
| Dataset / ambient dimension | 8,192 points / 512; dataset seed 42 |
| Construction | Keep coordinates 0 and 2 of the Swiss roll, project with a Gaussian-QR basis, standardize each ambient coordinate |
| Standardization | Subtract empirical mean, divide by empirical standard deviation plus |
| Training | 200,000 updates, batch 256, AdamW, constant learning rate |
| Optimizer | , , weight decay 0, gradient norm clip 1 |
| Time | Sigmoid-normal; implementation noise time , minimum |
| Setting | Value |
| Image resolution | RGB |
| Training duration | 200 epochs (controlled/scaling); 600 (extended H/XL) |
| Optimizer | AdamW, , |
| Global batch size | 1,024 |
| Learning rate | , constant after warmup |
| Weight decay | 0 |
| SiHC-B | SiHC-L | SiHC-XL | SiHC-H | |
| Depth | 12 | 24 | 40 | 32 |
| Hidden dimension | 768 | 1,024 | 1,024 | 1,280 |
| Attention heads | 12 | 16 | 16 | 16 |
| First block with context | 5 | 9 | 9 | 9 |
| REPA | No | No | Yes | No |
| Parameters (M) | 131.0 | 459.5 | 762.8 | 953.8 |
| Experiment | Solver | EMA | CFG | Interval |
| Pixel B-size / ablations | Heun50 | 0.9999 | 2.9 | |
| SiHC-L | Heun50 | 0.9999 | 2.4 | |
| SiHC-H | Heun50 | 0.9999 | 2.1 | |
| SiHC-XL + REPA | Heun50 | 0.9999 | 2.4 | |
| Whitened / | Heun50 | 0.9999 | 2.9 | |
| DINO/RAE plain / mHC | Heun50 | 0.9999 | 2.9 |
| Training epochs | |||||
| Model | 40 | 80 | 120 | 160 | 200 |
| DeCo-B | 63.84 | 19.42 | 12.64 | 9.54 | 8.07 |
| DiP-B | 82.13 | 27.74 | 19.54 | 15.89 | 13.31 |
| PixelDiT-B | 67.44 | 14.83 | 9.44 | 7.35 | 6.33 |
| Configuration | Read/write | Carry | FID | ||
| mHC copied input | 16 | 4 | mHC | mixed | 25.36 |
| mHC spatial input | 8 | 4 | mHC | mixed | 13.25 |
| Static scalar access | 8 | 4 | scalar | identity | 11.50 |
| SiHC, spatial input | 8 | 4 | feature-wise | identity | 10.10 |
| SiHC, direct input | 4 | 16 | feature-wise | identity | 8.07 |
| + 32 in-context class tokens | 4 | 16 | feature-wise | identity | 6.17 |
| Model | Depth width | Params | GFLOPs | FID | IS |
| DiT-only configurations reported by DiP | |||||
| DiT-only | 629M | 5.28 | 243.8 | ||
| DiT-only | 772M | 4.91 | 251.7 | ||
| DiT-only | 776M | 4.28 | 249.6 | ||
| DiT-only | 1.1B | 2.83 | 285.6 | ||
| SiHC, 200 epochs, without REPA | |||||
| Operation | Computation / implementation | Output shape |
| Coefficient contraction | Reduce over streams for | |
| All base reads | Read with every ; separate outputs feed compiled consumers | tensors |
| Active triangular sum | At update , accumulate the previous updates with | |
| Nonlinear update | Attention/MLP, normalization and conditioning on the compact workspace | |
| Final accumulation | Reduce all writes, add , store the final spatial states |
| Training | Inference | ||||
| Size | Routing | ms | GiB | ms | GiB |
| Blockwise | |||||
| B | Torch | 123.42 | 26.82 | 35.90 | 3.05 |
| B | Fused | 97.05 | 18.90 | 26.54 | 3.01 |
| L | Torch | 365.24 | 70.83 | 102.17 | 5.05 |
| L | Fused | 292.77 | 50.33 | 76.76 | 4.99 |
| Setting | Hidden-state probes |
| Frozen backbone | Raw epoch 200; null-class conditioning |
| Data | All 1,281,167 training and 50,000 validation images |
| Image preprocessing | ADM center crop; no random horizontal flips |
| Noise coefficient | , with paper time |
| Feature / head | Per-token RMS ( ) then GAP; affine |
| Training | 40 epochs; batch size 4,096 |
| Model | 25% noise | 50% noise | 75% noise |
| Pixel, | 25.32 / 46.45 | 27.54 / 49.64 | 22.82 / 43.91 |
| Pixel, | 5.38 / 14.51 | 4.56 / 12.68 | 3.47 / 10.17 |
| Pixel SiHC, | 25.51 / 46.68 | 27.81 / 50.32 | 22.88 / 43.98 |
| DINO, | 80.15 / 95.47 | 80.09 / 95.41 | 80.06 / 95.36 |
| DINO, | 81.09 / 95.74 | 81.07 / 95.81 | 80.93 / 95.77 |
| Matched input noise | Top-1 (%) | Top-5 (%) |
| 0% | 79.372 | 95.036 |
| 25% | ||
| 50% | ||
| 75% |