cs.CVSep 29, 2026

LDM-is-AE: Latent Diffusion Model is an Auto-Encoder for End-to-End Image Generation

Authors: Zhengqiang Zhang, Lingchen Sun, Rongyuan Wu, Qiaosi Yi, Xiangtao Kong, Chaodong Xiao, Lei Zhang

Organizations: The Hong Kong Polytechnic University · OPPO Research Institute

Abstract

Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusion model is trained to perform denoising within it. Such a two-stage design introduces a representation mismatch, as the latent space is optimized for reconstruction rather than adapting the denoising dynamics. We reveal that the LDM itself is an AE, and consequently present LDM-is-AE, an end-to-end one-stage LDM training framework that eliminates the need for a separately trained tokenizer. Our key observation is that the LDM backbone actually performs a latent-to-feature-to-latent transformation at each denoising step, which can be interpreted as an internal decoding--encoding process. Leveraging this structure, we split the DiT backbone into two reciprocal components, DiT-E (i.e., DiT Encoding) and DiT-D (i.e., DiT Decoding), and impose image-space supervision on the intermediate features across all timesteps. Our model encourages the internal representation to align with the image domain throughout denoising, thereby establishing an explicit latent-to-image-to-latent path. At the zero-noise timestep, our model further performs an image-to-latent-to-image mapping, corresponding to an auto-encoding process. As a result, LDM-is-AE jointly learns latent representations and denoising dynamics in an end-to-end manner, yielding a diffusion-native latent space tailored to the generation process. Experiments demonstrate that LDM-is-AE exhibits highly competitive generation performance, achieving an FID of 1.80 and 1.90 on 256x256 and 512x512 class-conditional image generation, respectively.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How to Train Your Latent Diffusion Language Model Jointly With the Latent Space

    May 8, 2026Viacheslav Meshchaninov, Alexander Shabalin, Egor Chimbulatov +4Diffusion Language ModelsLatent Diffusion Model

  2. TextLDM: Language Modeling with Continuous Latent Diffusion

    May 8, 2026Jiaxiu Jiang, Jingjing Ren, Wenbo Li +10Identity-Conditioned Latent Diffusion ModelsMultimodal Generation

  3. DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence

    Sep 30, 2026Xu Huang, Ye Huang, Zijun Liao +6Generative Image CompressionToken Compression