Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusion model is trained to perform denoising within it. Such a two-stage design introduces a representation mismatch, as the latent space is optimized for reconstruction rather than adapting the denoising dynamics. We reveal that the LDM itself is an AE, and consequently present LDM-is-AE, an end-to-end one-stage LDM training framework that eliminates the need for a separately trained tokenizer. Our key observation is that the LDM backbone actually performs a latent-to-feature-to-latent transformation at each denoising step, which can be interpreted as an internal decoding--encoding process. Leveraging this structure, we split the DiT backbone into two reciprocal components, DiT-E (i.e., DiT Encoding) and DiT-D (i.e., DiT Decoding), and impose image-space supervision on the intermediate features across all timesteps. Our model encourages the internal representation to align with the image domain throughout denoising, thereby establishing an explicit latent-to-image-to-latent path. At the zero-noise timestep, our model further performs an image-to-latent-to-image mapping, corresponding to an auto-encoding process. As a result, LDM-is-AE jointly learns latent representations and denoising dynamics in an end-to-end manner, yielding a diffusion-native latent space tailored to the generation process. Experiments demonstrate that LDM-is-AE exhibits highly competitive generation performance, achieving an FID of 1.80 and 1.90 on 256x256 and 512x512 class-conditional image generation, respectively.
Figures & tables
Figure 1 : Illustration of the auto-encoding nature of LDM. (a) The LDM backbone (i.e., a DiT) naturally performs a latent-to-feature-to-latent transformation. (b) Image-space supervision can align intermediate features with the image domain across timesteps. (c) At the zero-noise timestep ( t=1 ), an explicit latent-to-image-to-latent path is established, which in turn enables the reverse image-to-latent-to-image auto-encoding process.
Figure 2 : Architecture and training design of LDM-is-AE. (a) Network architecture and training pipeline. (b) Time-aware auxiliary feature mixing.
Method
Models
Repr.
Total Params (M)
Epochs
Aux. Data
Training FLOPs
w/ CFG
FID ↓
IS ↑
Two-stage
DiT-XL/2 [ 26 ]
Gen.+AE
Fixed
759
1400
[ 17 ]
45.4
2.27
278
SiT-XL/2 [ 23 ]
Gen.+AE
Fixed
759
1400
[ 17 ]
45.4
2.06
270
LightningDiT [ 40 ]
Gen.+AE+VFM
Fixed
745
800
-
19.1
1.35
295
REPA-SiT [ 42 ]
Gen.+AE+VFM
Fixed
759
800
[ 17 ]
25.9
1.29
306
DDT-XL/2 [ 37 ]
Gen.+AE+VFM
Fixed
759
400
[ 17 ]
19.2
1.26
311
Table 1 : Class-conditional ImageNet generation at 256×256 resolution. For Models , Gen. means generator, AE means autoencoder, Dec. means decoder, and VFM indicates the vision foundation model. For Repr. , Pixel means pixel diffusion, Fixed means fixed latent representation in training, and Dynamic means dynamically evolved representation in training. Aux. Data denotes external training data beyond ImageNet, and Training FLOPs report the generator-only training cost ( ×1019 ).
Figure 3 : Class-conditional ImageNet samples generated by LDM-is-AE at 256×256 resolution.
Method
Models
Repr.
Total Params (M)
w/ CFG
FID ↓
IS ↑
Two- stage
DiT-XL/2 [ 26 ]
Gen.+AE
Fixed
759
3.04
241
SiT-XL/2 [ 23 ]
Gen.+AE
Fixed
759
2.62
252
REPA-SiT-XL/2 [ 42 ]
Gen.+AE+VFM
Fixed
759
2.08
275
One- stage
ADM-G [ 8 ]
Gen.
Pixel
559
7.72
173
RIN [ 14 ]
Gen.
Pixel
320
3.95
216
Table 2 : Class-conditional ImageNet generation at 512×512 resolution.
Figure 6
Figure 5 : Ablations on the three main design choices in LDM-is-AE.
Method
FID ↓
IS ↑
sFID ↓
Prec. ↑
Recall ↑
SiT-B/2
45.19
37.40
26.16
0.471
0.613
+ LDM-is-AE
41.68
40.70
27.34
0.485
0.626
JiT-B/16
30.27
54.80
22.01
0.5249
0.6753
+ LPIPS
27.57
61.09
21.71
0.5623
0.6662
+ LDM-is-AE
25.95
67.45
20.38
0.5545
0.6836
Table 4 : Ablation on architecture backbone and LPIPS supervision. All models are evaluated without classifier-free guidance. Bold marks our method and the best result in each column.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Sin.Obj.
Two.Obj
Counting
Colors
Pos
Color.Attr.
Overall ↑
PixArt- α [ 5 ]
0.98
0.50
0.44
0.80
0.08
0.07
0.48
SD3 [ 10 ]
0.98
0.84
0.66
0.74
0.40
0.43
0.68
PixNerd [ 36 ]
0.97
0.86
0.44
0.83
0.71
0.53
0.73
DeCo [ 24 ]
1.00
0.92
0.72
0.91
0.80
0.79
0.86
LDM-is-AE
0.99
0.95
0.75
0.93
0.62
0.74
0.83
Appendix
Table 5 : Text-to-image generation on the GenEval benchmark.
Figure 6 : Additional class-conditional ImageNet samples generated by LDM-is-AE. All images are produced at a resolution of 256×256 .
Figure 7 : Reconstruction examples and latent visualizations for the auto-encoding path.
Scales (w(t),wlpips)
FID ↓
IS ↑
(1.0,1.0)
17.69
119
(1.0,0.3)
19.67
108
(1.0,3.0)
18.28
96
(0.3,1.0)
26.59
68
(3.0,1.0)
17.94
91
Appendix
Table 6 : Sensitivity to the loss weights w(t) and wlpips of Eq. 7 . All models are evaluated without classifier-free guidance.