Progressive autoregressive image codecs provide an appealing paradigm for generative compression by quantizing continuous latents into discrete tokens, transmitting coarse-to-fine prefix tokens and generating the remaining suffix tokens at the decoder. However, their reconstruction quality is fundamentally limited by two residuals introduced along this pipeline: the quantization residual, arising from information loss during discrete tokenization, and the generation residual, resulting from imperfect autoregressive generation of the suffix tokens. To address these limitations, we introduce ResARC, a residual-aware autoregressive codec that explicitly compensates for both residuals at the decoder. Specifically, we generate the quantization residual with a diffusion transformer conditioned on the autoregressive decoding context, while requiring no additional side information. In parallel, we compute the generation residual at the encoder and employ a learned Generation Residual Codec to efficiently compress and transmit it for decoder-side correction. The recovered residuals are then integrated with the reconstructed latent representation and decoded through an adapted VAE decoder. Extensive experiments demonstrate that ResARC achieves competitive perceptual similarity while substantially improving distributional fidelity over leading generative codecs across the ultra-low bitrate regime. Code and models will be released soon.
Figures & tables
Figure 1: Qualitative comparison with baselines at comparable ultra-low bitrates. ResARC better preserves fine-grained textures and image structures than competing generative codecs.
Figure 2: Residual-aware design of ResARC. (a) Prior autoregressive codecs introduce quantization and generation residuals . (b) ResARC generates the former and transmits the latter.
Figure 3: Overview of the ResARC architecture. ResARC decomposes the discrepancy in the autoregressive codec into a quantization residual and a generation residual. The quantization residual is synthesized at the decoder without additional bits, while the generation residual is compressed and transmitted for decoder-side correction.
Figure 4: Rate-perception comparison on DIV2K (left) and CLIC2020 (right). ResARC achieves strong overall performance across perceptual similarity and distributional fidelity metrics in the ultra-low bitrate regime.
Figure 5: Qualitative comparison with other generative codecs on DIV2K and CLIC2020.
Figure 6Figure 7Table 8
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: Architecture of the Generation Residual Codec. Prefix-depth conditioning and autoregressive context guide the analysis, entropy modeling, and synthesis stages of generation residual coding.
Stage
Steps
Global batch
LR
1. Autoregressive prior fine-tuning
2,000
64
6×10−5
2. Mixed-latent VAE decoder fine-tuning
35,000
16
2×10−5
3. Quantization residual generator training
Flow-matching pretraining
3,000
64
1×10−4
Image-space perceptual refinement
2,000
32
1×10−5
4. Generation residual codec training
Appendix
Table 7: Training hyperparameters across stages.
Figure 11: Rate allocation on DIV2K. The left panel shows the bitrate of each transmitted component across prefix depths, while the right panel shows the fraction of the total bitrate contributed by the generation-residual stream.
k
Text
AR Prefix
Gen. Residual
Total
5
3.75×10−4
2.85×10−3
2.30×10−4
3.46×10−3
6
3.75×10−4
6.16×10−3
2.30×10−4
6.76×10−3
7
3.75×10−4
1.18×10−2
2.29×10−4
1.25×10−2
8
3.75×10−4
2.05×10−2
2.79×10−4
2.11×10−2
9
3.75×10−4
3.24×10−2
3.14×10−4
3.31×10−2
10
3.75×10−4
5.51×10−2
1.98×10−4
5.56×10−2
Appendix
Table 8: Rate decomposition on DIV2K. Rates are reported in bpp and averaged over the DIV2K validation set.
k
bpp
KID mean ± std
10-seed range
5
0.00346
5.477±1.056
-
6
0.00676
3.177±0.946
-
7
0.01245
1.380±0.830
-
8
0.02110
0.675±0.854
[0.487,0.835]
9
0.03312
−0.397±0.846
[−0.501,−0.243]
10
0.05562
−1.140±0.782
[−1.191,−0.903]
Appendix
Table 9: ResARC KID on DIV2K. All KID entries are in 10−4 units.
Figure 12: Additional rate-quality comparison with leading generative codecs on the DIV2K validation and CLIC2020 test sets.
Methods
LPIPS ↓
DISTS ↓
FID ↓
DIV2K
CLIC2020
DIV2K
CLIC2020
DIV2K
CLIC2020
ARPC (ICLR’26) ( Zhang et al., 2026b )
0.00
0.00
0.00
0.00
0.00
0.00
DiffEIC (TCSVT’25) ( Li et al., 2024 )
+106.16
+104.11
+520.97
+350.87
+423.82
+466.27
DLF (ICCV’25) ( Xue et al., 2025a )
-32.50
-31.46
+22.63
+3.75
+42.75
+42.42
PerCo (ICLR’24) ( Careil et al., 2024 )
+190.35
+243.61
+639.33
+567.67
+770.14
+1207.55
GLC (CVPR’24) ( Jia et al., 2024 )
-34.72
-33.83
+43.01
+24.56
+76.14
+95.65
Appendix
Table 10: Signed BD-rate (%) relative to ARPC. Bold pink and blue denote the lowest and second-lowest reported values for each metric and dataset.
Figure 13: Progressive reconstruction across different prefix depths. Increasing k transmits more ground-truth prefix scales and progressively improves reconstruction quality.
Figure 14: More visual comparisons on CLIC2020. Labels report per-image bitrate (bpp) and DISTS.
Figure 15: More visual comparisons on DIV2K. Labels report per-image bitrate (bpp) and DISTS.
Generative image compression produces visually plausible reconstructions at low bitrates, yet their behavior as the rate approaches zero remains largely unexplored. When pushed below normal operating rates, representative codecs undergo semantic collapse: rather than gracefully losing source-specific detail, they produce malformed or unrecognizable content. Our analysis identifies two factors. As the bitrate decreases, reconstruction losses increasingly conflict with semantic objectives on gradients and visual results, while pixel-space and reconstruction-oriented VAE diffusion models become less efficient on semantic preservation. Guided by these findings, we introduce RAE-CoD, a compression-oriented diffusion (CoD) built in a representation autoencoder (RAE) space with direct alignment between compressed and source representations, preserving recognizable, naturally structured content for a 256×256 image with as few as 16 bits. We evaluate this framework using five vision foundation models (VFM) and a blinded vision-language model protocol. On MSCOCO-30K, RAE-CoD stands out from all evaluation. At 0.001-0.008 bpp, it reduces relative VFM feature MSE and Fréchet Distance ratio by at least 25.7% and 69.1% over the best competitors. Meanwhile, semantic recognizability and quality of the reconstructions remain nearly constant while source consistency falls smoothly, replacing abrupt semantic collapse with a graceful transition toward unconditional generation. Code will be released at https://github.com/LuizScarlet/RAE-CoD.
Diffusion-based image compression has achieved strong perceptual quality at ultra-low bitrates. However, existing codecs are often tied to specific backbones and specialized components, making diverse, rapidly evolving generative models difficult to reuse. This raises a natural question: Can modern generative foundation models be connected to image compression through a simple and extensible interface? Two insights guide our design: stronger generative priors make a simpler codec interface viable, and generation and compression can be intrinsically linked through latent transport. We therefore propose Gen2-IC with two stages: (1) Latent Compression maps clean image latents to entropy-constrained latents; and (2) Latent Transport refines them with one near-terminal update based on the pretrained model. Gen2-IC requires neither auxiliary conditioning signals nor task-specific backbone modifications. With lightweight adaptation and no distillation, it supports fast encoding and one-step decoding across multiple bitrates. We validate Gen2-IC on SD-2.1, SANA-1.5, FLUX.1-dev, and Qwen-Image-2512, spanning U-Net and Transformer architectures as well as diffusion and flow-matching formulations. With stronger priors, Gen2-IC delivers gains below 0.05 bpp: the Qwen variant leads diffusion-based generative codecs in reconstruction fidelity (PSNR), perceptual similarity (LPIPS and DISTS), and recognizer-based semantic fidelity (OCR CER/WER and face-ROI similarity) across four benchmarks.
DiffC provides a principled way to reuse pre-trained diffusion models for lossy compression, but its encoding and decoding procedures remain slow because they require many discretized forward and reverse steps. We study whether few-step generative models -- Rectified Flow, Consistency Trajectory Models (CTM), and MeanFlow -- can be cast as codecs within the same reverse channel coding (RCC) framework. The main challenge is that RCC requires posterior and shared distribution parameters, whereas these models do not explicitly parameterize intermediate conditional distributions. For Rectified Flow and MeanFlow, we use the equivalence between velocity parameterization and diffusion-style denoising parameterization to derive the quantities required by RCC. For CTM, which is distilled from EDM, we adopt the EDM noise parameterization together with local Gaussian approximations of the sender and shared distributions at intermediate states. This yields a proof-of-concept probabilistic formulation that enables compression with pre-trained few-step generative models without retraining. On low-resolution benchmarks, the resulting codecs reduce encoding and decoding time and improve realism in the low-bit-rate regime.