Rethinking Generative Image Compression at Extremely Low Bitrates
Organizations: University of Science and Technology of China
Abstract
Generative image compression produces visually plausible reconstructions at low bitrates, yet their behavior as the rate approaches zero remains largely unexplored. When pushed below normal operating rates, representative codecs undergo semantic collapse: rather than gracefully losing source-specific detail, they produce malformed or unrecognizable content. Our analysis identifies two factors. As the bitrate decreases, reconstruction losses increasingly conflict with semantic objectives on gradients and visual results, while pixel-space and reconstruction-oriented VAE diffusion models become less efficient on semantic preservation. Guided by these findings, we introduce RAE-CoD, a compression-oriented diffusion (CoD) built in a representation autoencoder (RAE) space with direct alignment between compressed and source representations, preserving recognizable, naturally structured content for a image with as few as 16 bits. We evaluate this framework using five vision foundation models (VFM) and a blinded vision-language model protocol. On MSCOCO-30K, RAE-CoD stands out from all evaluation. At 0.001-0.008 bpp, it reduces relative VFM feature MSE and Fréchet Distance ratio by at least 25.7% and 69.1% over the best competitors. Meanwhile, semantic recognizability and quality of the reconstructions remain nearly constant while source consistency falls smoothly, replacing abrupt semantic collapse with a graceful transition toward unconditional generation. Code will be released at https://github.com/LuizScarlet/RAE-CoD.
Figures & tables
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Checkpoint identifier | Arch. | Input | Dim. |
| Inception-v3 ( Szegedy et al., 2016 ) | inception_v3 | CNN | 2048 | |
| ConvNeXt-v2 ( Woo et al., 2023 ) | convnextv2_base.fcmae_ft_in22k_in1k | CNN | 1024 | |
| DINOv2 ( Oquab et al., 2023 ) | vit_large_patch14_dinov2.lvd142m | ViT | 1024 | |
| SigLIP2 ( Tschannen et al., 2025 ) | vit_so400m_patch16_siglip_256.v2_webli | ViT | 1152 | |
| CLIP ( Radford et al., 2021 ) | vit_large_patch14_clip_224.openai | ViT | 1024 |
| Representation | Feature | Top-1 (%) | Top-5 (%) |
| ConvNeXt-v2 | spatial average | 86.47 | 97.98 |
| DINOv2 | class token | 85.99 | 97.44 |
| SigLIP2 | attention pool | 85.56 | 97.78 |
| CLIP | class token | 83.27 | 97.14 |
| Inception-v3 | spatial average | 77.73 | 93.89 |
| LPIPS-VGG ( Zhang et al., 2018 ) | pooled multilevel features | 66.09 | 86.73 |
| Stage | Input to the VLM | Structured output | Role |
| A | Source image only | Global interpretation and at most 12 nonredundant units, each with importance | Defines source semantics independently for all metrics. |
| B | Reconstruction only | Visible-content description and at most 12 units with importance, identity confidence, and four intrinsic-quality ratings | Produces source-blind evidence for SR and SQ without reference. |
| C | Text inventories from A and B only | Global similarity and one match for every source and reconstruction semantic unit | Measures source coverage in both directions; images and the SR/SQ confidence and quality are withheld. |
| Case | Unit | Reconstruction-only semantic description | Quality tuple | ||
| (b) | C1 | A living room interior with furniture arranged around a central coffee table | 5 | 95 | |
| C2 | A patterned armchair with a floral or paisley design in earth tones | 4 | 90 | ||
| C3 | A glass-topped coffee table with a metal base in the center of the room | 4 | 90 | ||
| C4 | A large area rug with a geometric pattern of blue, red, and white squares | 3 | 85 | ||
| C5 | A small, round, dark-colored side table with curved legs near the window | 3 | 85 | ||
| C6 | A television set on a low cabinet in the background right | 3 | 80 |
| Component | Construction | Output |
| RAEv2 encoder | DINOv3-L/16 with seven-layer aggregation and normalization | |
| Pixel encoder | Four downsampling stages with channels | |
| Latent encoder | Concatenate both maps, project, and downsample once | |
| Hyper encoder | Two downsampling stages | |
| VQ bottleneck | 4-bit (16-entry) vector codebook | |
| Hyper decoder | Two upsampling stages to produce a hyperprior |
| Setting | Stage I | Stage II / inference |
| Initialization | Pretrained RAEv2 encoder, DDT and decoder; new codec, DDT LoRAs and condition projection | Corresponding Stage-I rate checkpoint |
| Trainable parameters | Codec, LoRA ( ), condition projection, and output interfaces | Codec and all DDT parameters |
| Frozen modules | RAEv2 encoder, decoder and the base DDT parameters | RAEv2 encoder and decoder |
| Rate weight | Undergoes 0.1, 2, 12, 16, 24, 32, and 48 | Fixed in |
| Rate schedule | 0.1 at step 0, 2 at 20K, then 12, 16, 24, 32, 48 at 30K intervals starting at 30K | 100K steps for each operating point |
| Learning rate |