cs.CVSep 28, 2026

What Visual Generators Need from Teachers: Rethinking Representation Alignment

Authors: Yongcong Wang, Hingchin Chen, Mingyu Fan, Shuo Jiang, Teer Zhang, Yucong Sun, Zijia Wang, Yiming Lu, +1 more

Organizations: Central South University · The Hong Kong University of Science and Technology · Tsinghua University · The Chinese University of Hong Kong, Shenzhen · SenseTime Research · Shandong University · Imperial College London · University of Oxford · Dell Technologies · University of International Relations

Abstract

Representation alignment speeds up diffusion transformer training by pulling an intermediate block of the model (student) toward features of a frozen pretrained encoder (teacher). Which teacher layer to align, and for how long, is still set by convention, and each alternative costs a training run. We find that alignment helps where the student cannot linearly recover the teacher's features, not where it already resembles them. Since a deep teacher layer is largely predictable from the one below, we isolate what each layer adds, its increment, and measure how much of it an unaligned student recovers. The student fills the teacher's hierarchy from the bottom up and stalls near the top, which we call hierarchy filling: even after 400K steps it recovers almost none of the deepest. The recoverability gap is the unrecovered share of an increment, read from one unaligned checkpoint. In short runs that each align one teacher layer at one block, the gap nearly reproduces their ranking by FID improvement, and CKA, a measure of feature similarity, largely reverses it. Representation Alignment and Recoverability Estimation (RARE) picks the teacher layer with the largest gap before training. During training, it tracks each token's remaining distance to that layer, the online counterpart of the gap, weights tokens by it, and phases out the loss once the average distance stops falling. With SiT-B/2 on ImageNet 256×256256\times256, RARE reaches an FID of 18.02 without guidance and 4.46 with it, ahead of seven alignment baselines including REPA, iREPA and HASTE. It also trains in 14% fewer GPU-hours than iREPA. Its FID stays below iREPA's across model scales, teachers, datasets and backbones.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training

    Jun 7, 2026Lianyu Pang, Tianlin Pan, Cheng Da +5Diffusion AlignmentDiffusion Transformers

  2. Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers

    May 16, 2026Shaodong Xu, Zhendong Wang, Litong Gong +4Diffusion AlignmentDiffusion Transformers

  3. Improving Visual Representation Alignment Generation with GRPO

    May 30, 2026Shentong Mo, Sukmin YunDiffusion AlignmentDiffusion Transformers