cs.CVSep 28, 2026

Scaffold Then Internalize: Representation Injection for Diffusion Transformers

Authors: Han Fu, Jiacheng Chen, Baoquan Zhao, Weidong Chen, Wei Liu, Qing Li, Xudong Mao

Organizations: Sun Yat-sen University · Video Rebirth · The Hong Kong Polytechnic University

Abstract

Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direction to REPA: rather than projecting diffusion representations into the encoder's space, we inject encoder representations into the diffusion transformer, allowing them to actively participate in the denoising process. To this end, we introduce \textit{REPresentation Injection} (REPI), a training framework based on a scaffold-to-internalization strategy, in which projected encoder representations initially serve as a temporary scaffold and are then progressively internalized by the diffusion transformer. REPI outperforms REPA across a wide range of backbones and is highly complementary to it: combining the two yields substantial gains over either alone. Notably, with only 160K training steps, REPI + REPA matches vanilla SiT trained for 7M steps, a speedup of over 43.5×43.5\times. Code will be available at https://jeneveuxpas.github.io/REPI

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers

    May 16, 2026Shaodong Xu, Zhendong Wang, Litong Gong +4Diffusion AlignmentDiffusion Transformers

  2. What Visual Generators Need from Teachers: Rethinking Representation Alignment

    Sep 28, 2026Yongcong Wang, Hingchin Chen, Mingyu Fan +6Diffusion AlignmentDiffusion Transformers

  3. Improving Visual Representation Alignment Generation with GRPO

    May 30, 2026Shentong Mo, Sukmin YunDiffusion AlignmentDiffusion Transformers