Despite recent progress in learned image compression, current image codecs still struggle to maintain realistic reconstructions at low bit-rates, often producing structured artifacts such as grid patterns or repetitive textures, even when trained with perceptual or adversarial losses. We introduce PerCoV2, an ultra-low bit-rate perceptual image compression system that unifies semantic tokenization, flow-based generation, and learned entropy modeling within a single framework. Building on the fully open flow-based SANA architecture, PerCoV2 introduces a novel resolution-adaptive 1D query-based tokenizer that produces compact semantic image tokens with a dual role in flow matching: providing a data-dependent reconstruction prior for initialization and a conditioning signal for flow-based refinement. By explicitly decoupling semantic representation from perceptual generation, our dual representation simplifies the flow-based learning objective, leading to more stable optimization and improved perceptual compression performance. PerCoV2 further introduces a dedicated 1D masked entropy model to improve rate efficiency and optional decoder-side multimodal enhancement via a vision-language model (Molmo) without increasing the transmitted bit budget. On MSCOCO-30k, PerCoV2 achieves state-of-the-art statistical fidelity, measured by FID and KID, across ultra-low and extreme bit-rates (0.0015-0.025 bpp). When trained solely on the general-purpose SA-1B dataset, PerCoV2 further demonstrates strong zero-shot generalization to widely adopted high-resolution benchmarks, including DIV2K and CLIC 2020, achieving competitive statistical fidelity with the current leading method, AEIC-ME. Finally, we introduce PerCoV2-distilled, a practical single-step variant derived from multi-step flow matching that accelerates decoding by 5.37x over PerCoV1, while preserving perceptual compression performance.
Figures & tables
Figure 1 : Perceptual compression performance of PerCoV2 on MSCOCO-30k. A larger area indicates better performance.
Figure 2 : Visual comparison of PerCoV2 on the Kodak dataset. Bit-rate increases relative to our method are indicated by (×) . Best viewed electronically.
Table 1 : Comparison of generative backbones in terms of AE compression design and model size. F denotes the spatial downsampling factor, C the number of latent channels, and P2 indicates 2×2 latent patchification.
Figure 4 : PerCoV1 preliminary investigation
Figure 5 : PerCoV2 architecture overview. For visualization clarity, the semantic reconstruction prior used for flow initialization is omitted and the flow process is illustrated from random noise.
Figure 6 : Tokenizer output (left), implicit residual refinement visualized as a normalized per-pixel difference map (middle), and the final reconstructed image (right). Bright regions indicate stronger refinement.
Figure 7 : Fixed-resolution TiTok (top) vs. resolution-adaptive TiTok (bottom)
Figure 8 : Inter-window communication
Figure 9 : Comparison of encoder representations between PerCoV1 (top) and PerCoV2 (bottom) through token-wise effective receptive field (ERF) visualization. The ERFs are computed from the continuous latent representations before vector quantization. Best viewed electronically.
Figure 10 : Performance of BLIP-2 (OPT-2.7B) on the Karpathy test set under corruptions
Figure 11 : Quantitative comparison of PerCoV2 on MSCOCO-30k and Kodak at 512×512
Figure 12 : Zero-shot comparison of PerCoV2 on high-resolution benchmarks
Figure 13 : Visual comparison of PerCoV2 and competing methods on MSCOCO-30k; bit-rate increases relative to PerCoV2 are shown as (×) . Best viewed zoomed-in.
Figure 14 : Extended comparison of PerCoV2 and competing methods on MSCOCO-30k; bit-rate increases relative to PerCoV2 are shown as (×) . Best viewed zoomed-in.
Figure 15 : Extended comparison of PerCoV2 and competing methods on MSCOCO-30k; bit-rate increases relative to PerCoV2 are shown as (×) . Best viewed zoomed-in.
Figure 16 : Visual comparison on CLIC 2020. We show 768×576 crops taken from high-resolution images of size 2048×1152 ; bit-rate increases relative to PerCoV2 are shown as (×) . Best viewed zoomed-in.
Figure 17 : Visual comparison on CLIC 2020. We show 768×576 crops taken from high-resolution images of size 2048×1152 ; bit-rate increases relative to PerCoV2 are shown as (×) . Best viewed zoomed-in.
Method
FID
KID
LPIPS
MS-SSIM
mIoU
CLIP
MS-ILLM
2708.8
2173.8
–
−29.5
318.0
181.9
DLF
–
–
−31.5
−20.3
34.0
51.9
CoD-Lite
5018.2
–
−19.8
−32.1
52.8
216.0
OneDC
712.3
874.6
–
–
−37.6
36.8
StableCodec
454.4
154.3
−25.2
−43.5
27.9
188.7
AEIC
303.5
79.2
−28.8
−42.5
44.2
197.3
Table 2 : BD-Rate (%) ↓ comparison on MSCOCO-30k, grouped into VAE-based, one-step diffusion, and multi-step diffusion methods. PerCoV2 (Ours) is used as the anchor.
Config / Component
Bpp ↓
Sav. (%) ↑
FID ↓
Δ%
KID ↓
Δ%
Base (Image Tokenizer)
0.00586
–
2.771
–
0.00081
–
+ Flow Model (SANA)
0.00586
0.00
2.498
9.85
0.00059
27.16
+ Multim. Enh.
0.00586
0.00
2.198
20.6
0.00035
56.8
+ Entropy Model
0.00565
3.58
2.198
0.00
0.00035
0.00
Base (Image Tokenizer)
0.01172
–
1.824
–
0.00037
–
+ Flow Model (SANA)
0.01172
0.00
1.697
6.96
0.00015
59.46
Table 3 : Ablation of PerCoV2 components across complexity tiers/ bit-rates
Figure 18 : Semantic preservation on MSCOCO-30k. Bit-rate increases are indicated by (×) . Best viewed electronically.
Figure 19 : Visualization of the generative capabilities of our hybrid 1D MIM. Revealed tokens are shown, while the remaining tokens are generated.
Strategy
FID ↓
Δ%
KID ↓
Δ%
LPIPS ↓
Δ%
SANA (No Text)
3.636
–
0.00008
–
0.486
–
+ Fixed Text ( Zhang et al., 2025 )
3.703
−5.89
0.00078
2.50
0.489
−0.62
+ Multim. Enh. (Ours)
3.496
3.85
0.00067
16.25
0.488
−0.41
+ GT Text (Up. Bound)
2.544
30.03
0.00023
71.25
0.480
1.23
SANA (No Text)
2.473
–
0.00028
–
0.417
–
+ Fixed Text ( Zhang et al., 2025 )
2.474
−0.04
0.00029
−3.57
0.420
−0.72
Table 4 : Global semantic guidance ablation. Gains ( Δ% ) are w.r.t. the No Text baseline.
Method
FID ↓
Δ (%)
VQ + CFM+ (noise, N2N)
50.165
–
VQ + Proxy + CFM+ (noise, N2N)
5.339
89.36
CFM+ (noise, frozen VQ)
4.771
90.49
CFM+ (semantic, frozen VQ)
4.186
91.65
Table 5 : Loss ablation. Relative improvements ( Δ% ) are reported w.r.t. VQ + CFM+ .
Figure 20 : ERF visualization of inter-window communication at native 2048×1536 resolution. Best viewed zoomed-in.
Figure 21 : Visual impressions at increasing bit rates. Best viewed electronically.
Diffusion-based image compression has achieved strong perceptual quality at ultra-low bitrates. However, existing codecs are often tied to specific backbones and specialized components, making diverse, rapidly evolving generative models difficult to reuse. This raises a natural question: Can modern generative foundation models be connected to image compression through a simple and extensible interface? Two insights guide our design: stronger generative priors make a simpler codec interface viable, and generation and compression can be intrinsically linked through latent transport. We therefore propose Gen2-IC with two stages: (1) Latent Compression maps clean image latents to entropy-constrained latents; and (2) Latent Transport refines them with one near-terminal update based on the pretrained model. Gen2-IC requires neither auxiliary conditioning signals nor task-specific backbone modifications. With lightweight adaptation and no distillation, it supports fast encoding and one-step decoding across multiple bitrates. We validate Gen2-IC on SD-2.1, SANA-1.5, FLUX.1-dev, and Qwen-Image-2512, spanning U-Net and Transformer architectures as well as diffusion and flow-matching formulations. With stronger priors, Gen2-IC delivers gains below 0.05 bpp: the Qwen variant leads diffusion-based generative codecs in reconstruction fidelity (PSNR), perceptual similarity (LPIPS and DISTS), and recognizer-based semantic fidelity (OCR CER/WER and face-ROI similarity) across four benchmarks.
Generative image compression produces visually plausible reconstructions at low bitrates, yet their behavior as the rate approaches zero remains largely unexplored. When pushed below normal operating rates, representative codecs undergo semantic collapse: rather than gracefully losing source-specific detail, they produce malformed or unrecognizable content. Our analysis identifies two factors. As the bitrate decreases, reconstruction losses increasingly conflict with semantic objectives on gradients and visual results, while pixel-space and reconstruction-oriented VAE diffusion models become less efficient on semantic preservation. Guided by these findings, we introduce RAE-CoD, a compression-oriented diffusion (CoD) built in a representation autoencoder (RAE) space with direct alignment between compressed and source representations, preserving recognizable, naturally structured content for a 256×256 image with as few as 16 bits. We evaluate this framework using five vision foundation models (VFM) and a blinded vision-language model protocol. On MSCOCO-30K, RAE-CoD stands out from all evaluation. At 0.001-0.008 bpp, it reduces relative VFM feature MSE and Fréchet Distance ratio by at least 25.7% and 69.1% over the best competitors. Meanwhile, semantic recognizability and quality of the reconstructions remain nearly constant while source consistency falls smoothly, replacing abrupt semantic collapse with a graceful transition toward unconditional generation. Code will be released at https://github.com/LuizScarlet/RAE-CoD.
Most existing video compression algorithms follow a paradigm of transformation and quantization, optimizing the trade-off between distortion and bitrate. However, extremely low-bitrate compression remains an underexplored frontier where perceptual quality optimization under severely constrained coding resources has not been adequately addressed. In this paper, we propose a unified generative framework that leverages pre-trained Diffusion Transformer (DiT) priors to achieve high perceptual quality at extremely low bitrates. We first introduce a flexible Group-of-Latents (GoL) strategy within the latent space of a causal tokenizer, explicitly partitioning the latent stream into intra I-latents and inter P-latents. The Deep Compression Module (I-DCM) then encodes key I-latents to preserve perceptual anchors with minimal overhead. Building upon these anchors, the DiT-based Unified Latent Denoising Module (U-LDM) refines intra-frame textures and synthesizes P-latents from noise, reconstructing temporal dynamics at zero additional bitrate cost. Extensive experiments demonstrate that our method uniquely operates in the extreme-low-bitrate regime (e.g., (<0.005) bpp), achieving state-of-the-art perceptual fidelity with rich spatial details and robust temporal consistency. The code will be made publicly available.
Shaokang Wang, Jinchang Xu, Peidong Jia +9
1State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China · 2State Key Discipline Laboratory of Wide Band-Gap Semiconductor Technology, School of Microelectronics, Xidian University, Xi'an, China · 3X-humanoid, Beijing, China