Progressive autoregressive image codecs provide an appealing paradigm for generative compression by quantizing continuous latents into discrete tokens, transmitting coarse-to-fine prefix tokens and generating the remaining suffix tokens at the decoder. However, their reconstruction quality is fundamentally limited by two residuals introduced along this pipeline: the quantization residual, arising from information loss during discrete tokenization, and the generation residual, resulting from imperfect autoregressive generation of the suffix tokens. To address these limitations, we introduce ResARC, a residual-aware autoregressive codec that explicitly compensates for both residuals at the decoder. Specifically, we generate the quantization residual with a diffusion transformer conditioned on the autoregressive decoding context, while requiring no additional side information. In parallel, we compute the generation residual at the encoder and employ a learned Generation Residual Codec to efficiently compress and transmit it for decoder-side correction. The recovered residuals are then integrated with the reconstructed latent representation and decoded through an adapted VAE decoder. Extensive experiments demonstrate that ResARC achieves competitive perceptual similarity while substantially improving distributional fidelity over leading generative codecs across the ultra-low bitrate regime. Code and models will be released soon.
Figures & tables
Figure 1: Qualitative comparison with baselines at comparable ultra-low bitrates. ResARC better preserves fine-grained textures and image structures than competing generative codecs.
Figure 2: Residual-aware design of ResARC. (a) Prior autoregressive codecs introduce quantization and generation residuals . (b) ResARC generates the former and transmits the latter.
Figure 3: Overview of the ResARC architecture. ResARC decomposes the discrepancy in the autoregressive codec into a quantization residual and a generation residual. The quantization residual is synthesized at the decoder without additional bits, while the generation residual is compressed and transmitted for decoder-side correction.
Figure 4: Rate-perception comparison on DIV2K (left) and CLIC2020 (right). ResARC achieves strong overall performance across perceptual similarity and distributional fidelity metrics in the ultra-low bitrate regime.
Figure 5: Qualitative comparison with other generative codecs on DIV2K and CLIC2020.
Figure 6Figure 7Table 8
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: Architecture of the Generation Residual Codec. Prefix-depth conditioning and autoregressive context guide the analysis, entropy modeling, and synthesis stages of generation residual coding.
Stage
Steps
Global batch
LR
1. Autoregressive prior fine-tuning
2,000
64
6×10−5
2. Mixed-latent VAE decoder fine-tuning
35,000
16
2×10−5
3. Quantization residual generator training
Flow-matching pretraining
3,000
64
1×10−4
Image-space perceptual refinement
2,000
32
1×10−5
4. Generation residual codec training
Appendix
Table 7: Training hyperparameters across stages.
Figure 11: Rate allocation on DIV2K. The left panel shows the bitrate of each transmitted component across prefix depths, while the right panel shows the fraction of the total bitrate contributed by the generation-residual stream.
k
Text
AR Prefix
Gen. Residual
Total
5
3.75×10−4
2.85×10−3
2.30×10−4
3.46×10−3
6
3.75×10−4
6.16×10−3
2.30×10−4
6.76×10−3
7
3.75×10−4
1.18×10−2
2.29×10−4
1.25×10−2
8
3.75×10−4
2.05×10−2
2.79×10−4
2.11×10−2
9
3.75×10−4
3.24×10−2
3.14×10−4
3.31×10−2
10
3.75×10−4
5.51×10−2
1.98×10−4
5.56×10−2
Appendix
Table 8: Rate decomposition on DIV2K. Rates are reported in bpp and averaged over the DIV2K validation set.
k
bpp
KID mean ± std
10-seed range
5
0.00346
5.477±1.056
-
6
0.00676
3.177±0.946
-
7
0.01245
1.380±0.830
-
8
0.02110
0.675±0.854
[0.487,0.835]
9
0.03312
−0.397±0.846
[−0.501,−0.243]
10
0.05562
−1.140±0.782
[−1.191,−0.903]
Appendix
Table 9: ResARC KID on DIV2K. All KID entries are in 10−4 units.
Figure 12: Additional rate-quality comparison with leading generative codecs on the DIV2K validation and CLIC2020 test sets.
Methods
LPIPS ↓
DISTS ↓
FID ↓
DIV2K
CLIC2020
DIV2K
CLIC2020
DIV2K
CLIC2020
ARPC (ICLR’26) ( Zhang et al., 2026b )
0.00
0.00
0.00
0.00
0.00
0.00
DiffEIC (TCSVT’25) ( Li et al., 2024 )
+106.16
+104.11
+520.97
+350.87
+423.82
+466.27
DLF (ICCV’25) ( Xue et al., 2025a )
-32.50
-31.46
+22.63
+3.75
+42.75
+42.42
PerCo (ICLR’24) ( Careil et al., 2024 )
+190.35
+243.61
+639.33
+567.67
+770.14
+1207.55
GLC (CVPR’24) ( Jia et al., 2024 )
-34.72
-33.83
+43.01
+24.56
+76.14
+95.65
Appendix
Table 10: Signed BD-rate (%) relative to ARPC. Bold pink and blue denote the lowest and second-lowest reported values for each metric and dataset.
Figure 13: Progressive reconstruction across different prefix depths. Increasing k transmits more ground-truth prefix scales and progressively improves reconstruction quality.
Figure 14: More visual comparisons on CLIC2020. Labels report per-image bitrate (bpp) and DISTS.
Figure 15: More visual comparisons on DIV2K. Labels report per-image bitrate (bpp) and DISTS.