cs.CVSep 27, 2026

Residual-Stream Burden Shapes Representation Learning in Diffusion Transformers

Authors: Tongtong Liang, Siqi Kou, Ziqiao Xi, Esha Singh, Kun Zhou, Zhijie Deng, Alexander Cloninger, Yu-Xiang Wang, +1 more

Organizations: UC San Diego · Shanghai Jiao Tong University · Aether AI

Abstract

In diffusion-based generation, a neural network can be trained to predict the clean data, the noise, or the velocity from a noisy input. These prediction targets are interconvertible and describe the same generative process, yet plain Diffusion Transformers operating on large pixel patches succeed with clean prediction and fail with noise or velocity prediction. We argue that this asymmetry arises because noisy targets require the residual stream to preserve noise-dependent input variation through depth for the final readout, forcing subsequent layers to compute on noisy representations. A spectrally concentrated clean target imposes a lighter demand, leaving greater freedom to organize hidden representations for subsequent computation. We call this preservation requirement residual-stream burden and show how it shapes representation learning in Diffusion Transformers. Controlled experiments indicate that the exploitable structure is spectral concentration in patch space and that the bandwidth of the persistent residual state is a key resource for noisy prediction. We further show that this account is consistent with recent decoupled pixel-space architectures, whose diverse designs all reduce the residual-stream burden on the main pathway. To examine this understanding from a complementary direction, we expand and reorganize the residual-stream bandwidth directly, introducing Spatially Indexed Hyper-Connections (SiHC) that reach FID 1.71 on ImageNet 2562256^2. Together, these results identify residual-stream burden as a mechanism through which prediction targets and architecture jointly shape representation learning in Diffusion Transformers.

Figures & tables

Appendix figures & tables41 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Equivalent Flows, Unequal Learning: Clean-Latent Prediction in Transformers

    May 26, 2026Funing Fu, Tenghui Wang, Guanyu Zhou +2Latent Diffusion ModelLatent Prediction

  2. PixelDiT2: Representation-Grounded Pixel Diffusion Transformers

    Sep 21, 2026Yongsheng Yu, Wei Xiong, Yichen Sheng +2Contextual Grounding