PerCoV2: Ultra-Low Bit-Rate Perceptual Image Compression via Query-Based 1D Multimodal Image Tokens
Organizations: Technical University of Munich · University of Applied Sciences Landshut · TU Wien
Abstract
Despite recent progress in learned image compression, current image codecs still struggle to maintain realistic reconstructions at low bit-rates, often producing structured artifacts such as grid patterns or repetitive textures, even when trained with perceptual or adversarial losses. We introduce PerCoV2, an ultra-low bit-rate perceptual image compression system that unifies semantic tokenization, flow-based generation, and learned entropy modeling within a single framework. Building on the fully open flow-based SANA architecture, PerCoV2 introduces a novel resolution-adaptive 1D query-based tokenizer that produces compact semantic image tokens with a dual role in flow matching: providing a data-dependent reconstruction prior for initialization and a conditioning signal for flow-based refinement. By explicitly decoupling semantic representation from perceptual generation, our dual representation simplifies the flow-based learning objective, leading to more stable optimization and improved perceptual compression performance. PerCoV2 further introduces a dedicated 1D masked entropy model to improve rate efficiency and optional decoder-side multimodal enhancement via a vision-language model (Molmo) without increasing the transmitted bit budget. On MSCOCO-30k, PerCoV2 achieves state-of-the-art statistical fidelity, measured by FID and KID, across ultra-low and extreme bit-rates (0.0015-0.025 bpp). When trained solely on the general-purpose SA-1B dataset, PerCoV2 further demonstrates strong zero-shot generalization to widely adopted high-resolution benchmarks, including DIV2K and CLIC 2020, achieving competitive statistical fidelity with the current leading method, AEIC-ME. Finally, we introduce PerCoV2-distilled, a practical single-step variant derived from multi-step flow matching that accelerates decoding by 5.37x over PerCoV1, while preserving perceptual compression performance.
Figures & tables
| Backbone | AE | Model Size | Open Source |
|---|---|---|---|
| GLIDE | F8C4 | B | ✗ |
| SD v2.1 | F8C4 | M | ✓ |
| SD v3.0 | F8C16-P2 | B | ✓ |
| FLUX.2 | F8C32-P2 | B | ✓ |
| SANA 1.5 | F32C32 | B | ✓ |
| Method | FID | KID | LPIPS | MS-SSIM | mIoU | CLIP |
|---|---|---|---|---|---|---|
| MS-ILLM | – | |||||
| DLF | – | – | ||||
| CoD-Lite | – | |||||
| OneDC | – | – | ||||
| StableCodec | ||||||
| AEIC |
| Config / Component | Bpp | Sav. (%) | FID | KID | ||
|---|---|---|---|---|---|---|
| Base (Image Tokenizer) | – | – | – | |||
| + Flow Model (SANA) | ||||||
| + Multim. Enh. | ||||||
| + Entropy Model | ||||||
| Base (Image Tokenizer) | – | – | – | |||
| + Flow Model (SANA) |
| Strategy | FID | KID | LPIPS | |||
|---|---|---|---|---|---|---|
| SANA (No Text) | – | – | – | |||
| + Fixed Text ( Zhang et al., 2025 ) | 3.703 | |||||
| + Multim. Enh. (Ours) | ||||||
| + GT Text (Up. Bound) | ||||||
| SANA (No Text) | – | – | – | |||
| + Fixed Text ( Zhang et al., 2025 ) |
| Method | FID | (%) |
|---|---|---|
| VQ + CFM+ (noise, N2N) | 50.165 | – |
| VQ + Proxy + CFM+ (noise, N2N) | 5.339 | |
| CFM+ (noise, frozen VQ) | 4.771 | |
| CFM+ (semantic, frozen VQ) | 4.186 |
| Model | Encoding (s) | Decoding (s) |
|---|---|---|
| MS-ILLM | ||
| DLF | ||
| CoD-Lite | ||
| OneDC | ||
| AEIC-ME | ||
| DiffEIC |