cs.LGJun 14, 2026

Distilling Drifting Transformers with Representation Autoencoders

Authors: Jiawei ZhangMengfei XiaGen LiYuantao Gu

Organizations: Tsinghua University · Ant Group · CUHK

Abstract

Representation Autoencoders (RAEs) have improved diffusion and flow models by semantically richer latent space owing to the strongly label-wise clustered DINO features in the pretrained encoders. Yet in the distillation stage, the severe anisotropy and large curvatures caused by the rich semantic representations would hinder the convergence and performance, making the trajectory-based distillation unstable. In this work, we argue that the RAE latent space is compatible with distillation via the newly proposed Drifting Models. We first quantitatively study the curvatures and isotropy statistics across different autoencoders, and theoretically reveal that Drifting Model itself is highly likely to fail on extremely scattered spaces like reconstruction-based VAEs. These motivate us to apply the drifting paradigm directly to representation autoencoders. Our proposed method, Drift-RAE, distills pretrained flow models in RAE latent spaces using Drifting, together with insightful modifications that improve training stability by thereotically aligning drifting fields with other frameworks. Regarding the experimental evidences, we achieve 1.77 FID on ImageNet 256 dataset using only 10k distillation steps, surpassing state-of-the-art RAE distillation methods and appearing comparative with the original Drifting Model without requiring an auxiliary MAE feature extractor. The code will be made publicly available.

Explore similar work

CardsList
  1. Improved Baselines with Representation Autoencoders

    May 18, 2026Jaskirat Singh, Boyang Zheng, Zongze Wu +3Masked AutoencodersImagenet

  2. On the Diffusibility of High-Dimensional Latents

    Sep 23, 2026Chao Feng, Zhiyang Xu, Bowei Chen +7