High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff between reconstruction fidelity and generation efficiency: high compression image encoder always increases the learning difficulty of diffusion training, resulting in slow model convergence. Recent representation autoencoders speed up the diffusion training by improving the latent feature's expressive capability by replacing VAE encoders with pretrained semantic encoders, yet they are typically limited to moderate compression and lose pixel-level details necessary for faithful reconstruction. To achieve both high compression and fast diffusion training, we propose DC-SAE, a Decoupled Compact Semantic Autoencoder designed for high-compression image generation with accelerated diffusion model convergence. DC-SAE consists of two key components: (1) a macro-level architecture design that leverages semantic encoders to enable higher compression ratios, and (2) a pixel-level encoder that preserves low-level details, ensuring high-fidelity image reconstruction. We empirically demonstrate that DC-SAE performs strongly on image generation tasks, achieving both compact latent representations and efficient training dynamics. Specifically, on the ImageNet dataset with 512×512 resolution, DC-SAE achieves 32× spatial compression, with 29.79 PSNR and 3.37 gFID, substantially outperforming the previous state-of-the-art high-compression tokenizer baselines DC-AE by 13.5% and 54.9% on PSNR and gFID, respectively, maintaining comparable throughput and faster diffusion model training convergence. Beyond class-conditional generation, a 1.6B-parameter DiT using DC-SAE achieves 0.84 on GenEval and 86.007 on DPG-Bench for text-to-image generation at 1024×1024 resolution.
Figures & tables
Figure 1 : Overview of DC-SAE under 32× spatial compression. (a) DC-SAE preserves reconstruction details missed by RAE and improves over DC-AE. (b) On ImageNet 512×512 , DC-SAE achieves better gFID than DC-AE. (c) DC-SAE also provides substantially higher autoencoder throughput.
Figure 2 : Pipeline of DC-SAE. We first adapt a frozen semantic encoder to a 32× token grid. Since the resulting semantic autoencoder suffers from poor reconstruction under aggressive compression, we introduce an unconstrained pixel branch to preserve local image details. The semantic and pixel branches are fused and decoded by a ViT decoder with Spatial Demerger.
Figure 3 : (a) Validation PSNR comparison between DeMerger and No-DeMerger variants. (b) DiT training convergence on joint semantic–pixel, semantic-only, and pixel-only latents under 32× compression. (c) Illustration of pre-merge and post-merge semantic compression for constructing a 32× semantic tokenizer.
Metric
RAE [ 52 ]
SD-VAE [ 35 ]
VA-VAE [ 44 ]
SVG [ 37 ]
REPA [ 46 ]
DC-SAE-16x
Reconstruction
PSNR ↑
19.20
24.08
27.96
23.89
–
27.80
rFID ↓
0.62
0.87
0.28
0.65
–
0.26
Generation
gFID ↓
2.16
–
5.96
6.57
7.90
3.09
IS ↑
214.80
–
128.00
137.90
118.60
189.95
Table 1 : Comparison of 16× semantic autoencoders on reconstruction quality and ImageNet 256×256 generation after 80 epochs. DC-SAE-16x achieves strong reconstruction while retaining competitive generation efficiency, making it a flexible baseline for high-compression tokenizer design. Notice: The 16× DC-SAE setting is used only in this ablation; all other DC-SAE results in this paper use the 32× setting.
Figure 4 : Qualitative results on ImageNet. The images are generated with 512×512 resolutions in the latent space of DC-SAE under 32× compression, with randomly sampled class conditions.
ImageNet 512×512
Image Generative Model
Autoencoder
Compression Rate
Params (B)
gFID ↓
Inception Score ↑
PSNR ↑
rFiD ↓
w/o CFG
w/ CFG
DiT-XL [ 32 ]
Flux-VAE-f8c16 [ 20 ]
8×
0.68
27.35
8.72
53.09
–
–
DiT-XL [ 32 ]
SD-VAE-f8c4 [ 35 ]
8×
0.67
12.03
3.04
105.25
–
–
SiT-XL [ 27 ]
SD-VAE-f8c4 [ 35 ]
8×
0.67
–
2.62
–
MAGVIT-v2 [ 45 ]
–
–
–
3.07
1.91
213.10
–
–
Table 2 : Comparison with state-of-the-art image generative models on ImageNet 512×512 class-conditional generation. DC-SAE uses a 32× high-compression semantic tokenizer and trains a DiT-XL generator in the learned latent space.
Image Generative Model
Autoencoder
Compression Rate
Params (B)
GenEval ↑
DPG-Bench ↑
DC-Gen-FLUX.1-Krea-12B [ 14 ]
DC-AE f32c32 [ 6 ]
32×
12
0.72
87.073
DC-Gen-FLUX.1-Krea-12B [ 14 ]
DC-AE-1.5 f64c128 [ 7 ]
64×
12
0.59
75.439
1.6B-DiT
DC-SAE
32×
1.6
0.84
86.007
Table 3 : Text-to-image generation results with CFG. The DC-SAE model is evaluated at 1024×1024 resolution. Generator sizes and tokenizer compression ratios are listed explicitly; the baselines use different model and training configurations. Higher scores are better for both benchmarks.
Semantic Encoder
PSNR ↑
rFID ↓
gFID ↓
IS ↑
Post-merge
Qwen-ViT Frozen Merger
23.40
0.62
25.96
66.83
Qwen-ViT AvgPool Merger
23.07
0.61
17.71
84.45
Qwen-ViT Learnable Merger
23.85
0.58
16.84
78.34
DINOv2 Learnable Merger
24.03
0.35
8.60
108.11
Pre-merge
Table 4 : Comparison between post-merge and pre-merge strategies for constructing a 32× semantic tokenizer.
ImageNet 256×256
Diffusion Model
Setting
Autoencoder
rFID ↓
gFID ↓
Inception Score ↑
DiT-XL [ 32 ]
f32c32
DC-AE [ 6 ]
0.69
10.18
107.49
DiT-XL [ 32 ]
f32c128
DC-AE [ 6 ]
0.26
26.44
53.41
DiT-XL [ 32 ]
f32c32
DC-AE-1.5 [ 7 ]
–
10.50
107.99
DiT-XL [ 32 ]
f32c128
DC-AE-1.5 [ 7 ]
0.26
17.31
80.38
DiT-XL [ 32 ]
f32c64
DC-SAE (Ours)
0.45
5.67
156.68
Table 5 : Comparison with DC-AE and DC-AE-1.5 on ImageNet 256×256 class-conditional generation.
Metric
Joint
Late-Concat
Late-Concat-E
PSNR ↑
29.79
31.91
29.15
rFID ↓
0.17
0.09
0.17
20 ep.
8.12 / 134.23
10.95 / 122.48
10.35 / 127.78
40 ep.
5.23 / 166.76
7.41 / 153.16
6.91 / 156.97
60 ep.
4.33 / 180.31
6.22 / 168.13
6.01 / 167.48
80 ep.
4.02 / 185.46
5.64 / 174.52
5.40 / 173.00
Table 6 : Standalone autoencoder ablation. Entries are FID / IS after DiT training.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Model
256×256
512×512
1024×1024
DC-AE-f32c32
265.6
69.7
17.4
DC-SAE-f32c64
1233.5
330.8
76.8
DC-SAE-f32c256
1054.5
284.4
67.9
Appendix
Table 7 : End-to-end throughput comparison between DC-AE and DC-SAE (imgs/s).
Semantic encoder setting
Spatial DeMerger
w/o DeMerger
w/ DeMerger
QwenViT w/ original merger
15.39
17.52
QwenViT w/ 2×2 avg pooling
15.50
17.64
Appendix
Table 8 : Additional analysis of Spatial DeMerger with QwenViT-based semantic encoders. We compare the original QwenViT merger and a 2×2 average pooling variant, with and without Spatial DeMerger. Both semantic encoders are frozen, and no additional pixel encoder is used.
Existing text-to-image diffusion models excel at generating high-quality images, but face significant efficiency challenges when scaled to high resolutions, like 4K image generation. While previous research accelerates diffusion models in various aspects, it seldom handles the inherent redundancy within the latent space. To bridge this gap, this paper introduces DC-Gen, a general framework that accelerates text-to-image diffusion models by leveraging a deeply compressed latent space. Rather than a costly training-from-scratch approach, DC-Gen uses an efficient post-training pipeline to preserve the quality of the base model. A key challenge in this paradigm is the representation gap between the base model's latent space and a deeply compressed latent space, which can lead to instability during direct fine-tuning. To overcome this, DC-Gen first bridges the representation gap with a lightweight embedding alignment training. Once the latent embeddings are aligned, only a small amount of LoRA fine-tuning is needed to unlock the base model's inherent generation quality. We verify DC-Gen's effectiveness on SANA and FLUX.1-Krea. The resulting DC-Gen-SANA and DC-Gen-FLUX models achieve quality comparable to their base models but with a significant speedup. Specifically, DC-Gen-FLUX reduces the latency of 4K image generation by 53x on the NVIDIA H100 GPU. When combined with NVFP4 SVDQuant, DC-Gen-FLUX generates a 4K image in just 3.5 seconds on a single NVIDIA 5090 GPU, achieving a total latency reduction of 138x compared to the base FLUX.1-Krea model. Code: https://github.com/dc-ai-projects/DC-Gen.
Tokenizers are a key component of state-of-the-art generative image models, extracting the most important features from the signal while reducing data dimension and redundancy. Most current tokenizers are based on KL-regularized variational autoencoders (KL-VAE), trained with reconstruction, perceptual and adversarial losses. Diffusion decoders have been proposed as a more principled alternative to model the distribution over images conditioned on the latent. However, matching the performance of KL-VAE still requires adversarial losses, as well as a higher decoding time due to iterative sampling. To address these limitations, we introduce a new pixel diffusion decoder architecture for improved scaling and training stability, benefiting from transformer components and GAN-free training. We use distillation to replicate the performance of the diffusion decoder in an efficient single-step decoder. This makes SSDD the first diffusion decoder optimized for single-step reconstruction trained without adversarial losses, reaching higher reconstruction quality and faster sampling than KL-VAE. In particular, SSDD improves reconstruction FID from 0.87 to 0.46 with 1.4× higher throughput and preserve generation quality of DiTs with 3.8× faster sampling. As such, SSDD can be used as a drop-in replacement for KL-VAE, and for building higher-quality and faster generative models.
Théophane Vallaeys, Jakob Verbeek, Matthieu Cord
Meta Fundamental AI Research · Sorbonne University
We present Qwen-Image-VAE-2.0, a suite of high-compression Variational Autoencoders (VAEs) that achieve significant advances in both reconstruction fidelity and diffusability. To address the reconstruction bottlenecks of high compression, we adopt an improved architecture featuring Global Skip Connections (GSC) and expanded latent channels. Moreover, we scale training to billions of images and incorporate a synthetic rendering engine to improve performance in text-rich scenarios. To tackle the convergence challenges of high-dimensional latent space, we implement an enhanced semantic alignment strategy to make the latent space highly amenable to diffusion modeling. To optimize computational efficiency, we leverage an asymmetric and attention-free encoder-decoder backbone to minimize encoding overhead. We present a comprehensive evaluation of Qwen-Image-VAE-2.0 on public reconstruction benchmarks. To evaluate performance in text-rich scenarios, we propose OmniDoc-TokenBench, a new benchmark comprising a diverse collection of real-world documents coupled with specialized OCR-based evaluation metrics. Qwen-Image-VAE-2.0 achieves state-of-the-art reconstruction performance, demonstrating exceptional capabilities in both general domains and text-rich scenarios at high compression ratio. Furthermore, downstream DiT experiments reveal our models possess superior diffusability, significantly accelerating convergence compared to existing high-compression baselines. These establish Qwen-Image-VAE-2.0 as a leading model with high compression, superior reconstruction, and exceptional diffusability.