High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff between reconstruction fidelity and generation efficiency: high compression image encoder always increases the learning difficulty of diffusion training, resulting in slow model convergence. Recent representation autoencoders speed up the diffusion training by improving the latent feature's expressive capability by replacing VAE encoders with pretrained semantic encoders, yet they are typically limited to moderate compression and lose pixel-level details necessary for faithful reconstruction. To achieve both high compression and fast diffusion training, we propose DC-SAE, a Decoupled Compact Semantic Autoencoder designed for high-compression image generation with accelerated diffusion model convergence. DC-SAE consists of two key components: (1) a macro-level architecture design that leverages semantic encoders to enable higher compression ratios, and (2) a pixel-level encoder that preserves low-level details, ensuring high-fidelity image reconstruction. We empirically demonstrate that DC-SAE performs strongly on image generation tasks, achieving both compact latent representations and efficient training dynamics. Specifically, on the ImageNet dataset with 512×512 resolution, DC-SAE achieves 32× spatial compression, with 29.79 PSNR and 3.37 gFID, substantially outperforming the previous state-of-the-art high-compression tokenizer baselines DC-AE by 13.5% and 54.9% on PSNR and gFID, respectively, maintaining comparable throughput and faster diffusion model training convergence. Beyond class-conditional generation, a 1.6B-parameter DiT using DC-SAE achieves 0.84 on GenEval and 86.007 on DPG-Bench for text-to-image generation at 1024×1024 resolution.
Figures & tables
Figure 1 : Overview of DC-SAE under 32× spatial compression. (a) DC-SAE preserves reconstruction details missed by RAE and improves over DC-AE. (b) On ImageNet 512×512 , DC-SAE achieves better gFID than DC-AE. (c) DC-SAE also provides substantially higher autoencoder throughput.
Figure 2 : Pipeline of DC-SAE. We first adapt a frozen semantic encoder to a 32× token grid. Since the resulting semantic autoencoder suffers from poor reconstruction under aggressive compression, we introduce an unconstrained pixel branch to preserve local image details. The semantic and pixel branches are fused and decoded by a ViT decoder with Spatial Demerger.
Figure 3 : (a) Validation PSNR comparison between DeMerger and No-DeMerger variants. (b) DiT training convergence on joint semantic–pixel, semantic-only, and pixel-only latents under 32× compression. (c) Illustration of pre-merge and post-merge semantic compression for constructing a 32× semantic tokenizer.
Metric
RAE [ 52 ]
SD-VAE [ 35 ]
VA-VAE [ 44 ]
SVG [ 37 ]
REPA [ 46 ]
DC-SAE-16x
Reconstruction
PSNR ↑
19.20
24.08
27.96
23.89
–
27.80
rFID ↓
0.62
0.87
0.28
0.65
–
0.26
Generation
gFID ↓
2.16
–
5.96
6.57
7.90
3.09
IS ↑
214.80
–
128.00
137.90
118.60
189.95
Table 1 : Comparison of 16× semantic autoencoders on reconstruction quality and ImageNet 256×256 generation after 80 epochs. DC-SAE-16x achieves strong reconstruction while retaining competitive generation efficiency, making it a flexible baseline for high-compression tokenizer design. Notice: The 16× DC-SAE setting is used only in this ablation; all other DC-SAE results in this paper use the 32× setting.
Figure 4 : Qualitative results on ImageNet. The images are generated with 512×512 resolutions in the latent space of DC-SAE under 32× compression, with randomly sampled class conditions.
ImageNet 512×512
Image Generative Model
Autoencoder
Compression Rate
Params (B)
gFID ↓
Inception Score ↑
PSNR ↑
rFiD ↓
w/o CFG
w/ CFG
DiT-XL [ 32 ]
Flux-VAE-f8c16 [ 20 ]
8×
0.68
27.35
8.72
53.09
–
–
DiT-XL [ 32 ]
SD-VAE-f8c4 [ 35 ]
8×
0.67
12.03
3.04
105.25
–
–
SiT-XL [ 27 ]
SD-VAE-f8c4 [ 35 ]
8×
0.67
–
2.62
–
MAGVIT-v2 [ 45 ]
–
–
–
3.07
1.91
213.10
–
–
Table 2 : Comparison with state-of-the-art image generative models on ImageNet 512×512 class-conditional generation. DC-SAE uses a 32× high-compression semantic tokenizer and trains a DiT-XL generator in the learned latent space.
Image Generative Model
Autoencoder
Compression Rate
Params (B)
GenEval ↑
DPG-Bench ↑
DC-Gen-FLUX.1-Krea-12B [ 14 ]
DC-AE f32c32 [ 6 ]
32×
12
0.72
87.073
DC-Gen-FLUX.1-Krea-12B [ 14 ]
DC-AE-1.5 f64c128 [ 7 ]
64×
12
0.59
75.439
1.6B-DiT
DC-SAE
32×
1.6
0.84
86.007
Table 3 : Text-to-image generation results with CFG. The DC-SAE model is evaluated at 1024×1024 resolution. Generator sizes and tokenizer compression ratios are listed explicitly; the baselines use different model and training configurations. Higher scores are better for both benchmarks.
Semantic Encoder
PSNR ↑
rFID ↓
gFID ↓
IS ↑
Post-merge
Qwen-ViT Frozen Merger
23.40
0.62
25.96
66.83
Qwen-ViT AvgPool Merger
23.07
0.61
17.71
84.45
Qwen-ViT Learnable Merger
23.85
0.58
16.84
78.34
DINOv2 Learnable Merger
24.03
0.35
8.60
108.11
Pre-merge
Table 4 : Comparison between post-merge and pre-merge strategies for constructing a 32× semantic tokenizer.
ImageNet 256×256
Diffusion Model
Setting
Autoencoder
rFID ↓
gFID ↓
Inception Score ↑
DiT-XL [ 32 ]
f32c32
DC-AE [ 6 ]
0.69
10.18
107.49
DiT-XL [ 32 ]
f32c128
DC-AE [ 6 ]
0.26
26.44
53.41
DiT-XL [ 32 ]
f32c32
DC-AE-1.5 [ 7 ]
–
10.50
107.99
DiT-XL [ 32 ]
f32c128
DC-AE-1.5 [ 7 ]
0.26
17.31
80.38
DiT-XL [ 32 ]
f32c64
DC-SAE (Ours)
0.45
5.67
156.68
Table 5 : Comparison with DC-AE and DC-AE-1.5 on ImageNet 256×256 class-conditional generation.
Metric
Joint
Late-Concat
Late-Concat-E
PSNR ↑
29.79
31.91
29.15
rFID ↓
0.17
0.09
0.17
20 ep.
8.12 / 134.23
10.95 / 122.48
10.35 / 127.78
40 ep.
5.23 / 166.76
7.41 / 153.16
6.91 / 156.97
60 ep.
4.33 / 180.31
6.22 / 168.13
6.01 / 167.48
80 ep.
4.02 / 185.46
5.64 / 174.52
5.40 / 173.00
Table 6 : Standalone autoencoder ablation. Entries are FID / IS after DiT training.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Model
256×256
512×512
1024×1024
DC-AE-f32c32
265.6
69.7
17.4
DC-SAE-f32c64
1233.5
330.8
76.8
DC-SAE-f32c256
1054.5
284.4
67.9
Appendix
Table 7 : End-to-end throughput comparison between DC-AE and DC-SAE (imgs/s).
Semantic encoder setting
Spatial DeMerger
w/o DeMerger
w/ DeMerger
QwenViT w/ original merger
15.39
17.52
QwenViT w/ 2×2 avg pooling
15.50
17.64
Appendix
Table 8 : Additional analysis of Spatial DeMerger with QwenViT-based semantic encoders. We compare the original QwenViT merger and a 2×2 average pooling variant, with and without Spatial DeMerger. Both semantic encoders are frozen, and no additional pixel encoder is used.