Vector quantization, which discretizes a continuous vector space into a finite set of representative vectors (a codebook), has been widely adopted in modern machine learning. Despite its effectiveness, vector quantization poses a fundamental challenge: the non-differentiable quantization step blocks gradient backpropagation. Smoothed vector quantization addresses this issue by relaxing the discrete selection of a codebook vector into a weighted combination of codebook entries, represented as the matrix product of a simplex vector and the codebook. Effective smoothing requires two properties: (1) smoothed code-selection vectors should remain close to a onehot vector, ensuring tight approximation, and (2) all codebook entries should be utilized, preventing code collapse. Existing methods typically address these desiderata separately. By contrast, the present study introduces a simple and intuitive regularization that promotes both simultaneously by minimizing the distance between each simplex vertex and its K-nearest smoothed code-selectors. Representative benchmarks on image encoding demonstrate that the proposed method achieves more effective codebook utilization and improves performance over prior approaches.
Figures & tables
Figure 1 : ( 1 ) Four different distributions on the simplex Δ3−1 . For effective smoothed vector quantization, samples should be concentrated near the vertices of the simplex (i.e., onehot-like vectors; orange), rather than centered (dark gray) or uniformly spread across the simplex (light gray). At the same time, each vertex must have some samples in its neighborhood to avoid code collapse (blue). ( 1 ) Maximizing the perplexity of the sample mean ( Baevski et al., 2020b, ) penalizes code collapse but cannot discriminate among the other three distributions. ( 1 ) The proposed K -nearest neighbor (KNN) distance minimization ( K=8 ) favors the desired vertex-concentrated distribution while also preventing code collapse.
Feature Map Size; Codebook Size
16 × 16 × 32; 1024
64 × 64 × 3; 8196
64 × 64 × 64; 8196
Method
K/4
Usage ( ↑ )
rMSE ( ↓ )
FID ( ↓ )
IS ( ↑ )
Usage ( ↑ )
rMSE ( ↓ )
FID ( ↓ )
IS ( ↑ )
Usage ( ↑ )
rMSE ( ↓ )
FID ( ↓ )
IS ( ↑ )
STE
Euclid
—
4.5%
0.108
101.84
10.67
100.0%
0.045
10.41
142.12
1.7%
0.067
25.65
82.61
Cos
—
3.0%
0.102
108.09
10.03
70.9%
0.053
18.13
107.38
1.7%
0.054
20.08
101.74
SimVQ
—
100.0%
0.091
86.94
15.93
100.0%
0.046
10.82
141.17
100.0%
0.045
12.54
134.84
RE
Euclid
—
3.1%
0.123
127.40
7.40
78.8%
0.046
14.04
123.30
0.2%
0.077
42.12
51.86
Table 1 : Performance of discrete autoencoding on the ImageNet validation set. Reported metrics are codebook usage and reconstruction quality scores: root mean squared error (rMSE), Fréchet Inception Distance (FID), Inception Score (IS), and Structural Similarity Index Measure (SSIM). The proposed method is denoted as “KNN-L2/CE”. The baseline methods are abbreviated as follows: “Cos” denotes cosine-distance quantization (contrasted with Euclidean distance, ’Euclid’). “HG” and “SG” denote Hard-Gumbel and Soft-Gumbel, i.e., Gumbel-softmax sampling with and without straight-through hard quantization in the forward pass, respectively. Best scores across all methods are highlighted in boldface, while underlined values indicate the best-performing value of K/4 .
Feature Map Size; Codebook Size
16 × 16 × 32; 1024
64 × 64 × 3; 8192
64 × 64 × 64; 8192
Method
K/4
75%
90%
99%
Max
75%
90%
99%
Max
75%
90%
99%
Max
PPL
—
910.40
933.12
963.29
995.91
7559.89
7561.99
7563.10
7565.02
8186.32
8186.40
8186.55
8186.76
KNN-L2
1
1.17
1.59
2.36
6.26
1.00
1.00
1.00
4.00
9.63
18.10
51.04
123.62
2
1.00
1.08
1.86
27.37
1.00
1.00
1.06
4.65
3.85
6.36
16.53
43.40
4
1.00
1.06
1.86
62.99
1.00
1.00
1.00
4.61
1.06
1.06
1.06
1.07
Table 2 : Tightness of softmax-based smoothing (without Gumbel sampling), measured by the individual perplexity of smoothed code-selection vectors, exp(−∑m=1Mpmlogpm) . Reported values are the 75th, 90th, and 99th percentiles, as well as the maximum, computed across all feature-map pixels in the ImageNet validation set.
Figure 2 : (Top) Histogram of per-image codebook utilization on the validation dataset, defined as the percentage of codes used at least once within each image. (Bottom) Histogram of code popularity, defined as the proportion of validation images in which each code occurs.
Method
Evaluated in
D/G
MG
Usage % ( ↑ )
FID ( ↓ )
IS ( ↑ )
STE Euclid
Esser et al., (2021)
256/1
10241
—
7.94
—
256/1
163841
—
4.98
—
This Work
256/8
10248
0.1–0.2
398.52
1.00
RE Euclid
Fifty et al., (2025)
256/1
10241
98
4.6
146.5
This Work
256/8
10248
0.2–0.8
506.60
1.00
SimVQ
This Work 7 7 7 The original SimVQ paper evaluated the method on 128×128 inputs while keeping a 16×16 latent map ( Zhu et al.,, 2025 ) , which severely reduces the compression rate and makes reconstruction significantly easier. Since those scores are not directly comparable to the standard 256×256 VQGAN benchmark, SimVQ is evaluated here on the same 256×256 pipeline to ensure a rigorous and direct comparison.
32/1
10241
1.0
436.34
1.00
Table 3 : Performance of VQGAN on the ImageNet validation set. Reported metrics are codebook usage and reconstruction quality scores. Best scores across all methods are highlighted in boldface. For the single-codebook STE/RE Euclid baselines ( G=1 ), scores from the original publications are reported, since attempts to reproduce these results were unsuccessful. 6 6 6 Attempts to replicate STE Euclid results using the official repository yielded highly degraded scores; for example, the FID scores were 172.78 for M=1024 and 349.18 for M=16384 . These results are consistent with community-reported issues (e.g., https://github.com/CompVis/taming-transformers/issues/87 ). Furthermore, the pre-trained weights for the RE Euclid model ( Fifty et al.,, 2025 , for replicating Table 2 in) are unavailable in the official repository ( https://github.com/cfifty/rotation_trick/tree/main ). Due to these reproducibility challenges, the “optimal scores” published in the original papers are reported for the single-codebook baselines. All product-quantized configurations ( G=8 ), including these baselines, were newly implemented and evaluated.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
GAN (§ 4.2 )
Autoencoding (§ 4.1 )
Feature Map Size
16×16×32
16×16×32
64×64×3
64×64×64
Codebook Size
1024
8196
8196
8196
Latent Channels
128→128→64→64→32→32
128→64→32→{3,64}
Height & Width
256→128→64→32→16→16
256→128→64→64
Batch Size
64
64
64
64
Training Epochs
60
25
20
20
Appendix
Table 4 : Hyperparameters for discrete autoencoding/GAN.
Feature Map Size; Codebook Size
16 × 16 × 32; 1024
64 × 64 × 3; 8196
64 × 64 × 64; 8196
Wall-Clock Time
Wall-Clock Time
Wall-Clock Time
Method
K/4
Entire
Quant.
VRAM
Entire
Quant.
VRAM
Entire
Quant.
VRAM
(ms)
(ms)
(GB)
(ms)
(ms)
(GB)
(ms)
(ms)
(GB)
STE
STE Euclid
—
330.39
1.35
34.20
508.49
1.35
54.83
502.20
1.30
54.93
SimVQ
—
330.18
1.31
34.20
504.28
1.36
54.83
502.25
1.37
54.92
Appendix
Table 5 : Wall-clock time and peak VRAM consumption per training iteration. The models were run on four NVIDIA A100 GPUs (80GB VRAM per GPU) using DistributedDataParallel of PyTorch. The reported scores represent computational load per GPU on average over a single epoch of the ImageNet training split, with batch size set to 64. The “Entire” wall-clock time includes the forward and backward passes through the encoder, quantizer, and decoder modules, while the ‘Quant.’ reports the wall-clock time consumed by the forward pass through the quantizer alone.
Figure 3 : Dirichlet distributions on the simplex Δ3−1 with concentration parameters α1=α2=α3=α , where α∈{0.5,1.0,2.0} .
The effectiveness of modern visual representation learning and autoregressive models critically depends on vector quantization (VQ), which discretizes continuous feature representations using a learnable codebook. Despite its widespread use, existing VQ methods often suffer from training instability and codebook collapse, arising from gradient mismatch induced by the straight-through estimator and the under-utilization of code vectors. In this work, we show that both issues can be traced to a fundamental mismatch between the distributions of feature vectors and code vectors, leading to inefficient representation and information loss. Building on this observation, we propose a distributional matching framework for vector quantization. We introduce principled criteria for desirable VQ behavior and demonstrate through theoretical analysis and empirical evaluation that aligning feature and code vector distributions provides a unifying mechanism for mitigating training instability and codebook collapse. We instantiate this framework using a Wasserstein-based objective with an efficient closed-form under a mild Gaussian approximation, and further show that a nonparametric alternative based on maximum mean discrepancy yields comparable performance. Extensive experiments on visual tokenization benchmarks support the effectiveness and robustness of the proposed approach.
Xianghong Fang, Litao Guo, Hengchao Chen +8
University of Toronto · The Hong Kong University of Science and Technology · Lehigh University +2
Vector quantization is central to modern generative modeling pipelines, but large-codebook VQ models often suffer from codebook collapse. We identify encoder drift as a key driver of this failure: as the encoder moves the latent distribution, sparsely updated code vectors can lag behind, lose assignments, and increase quantization error, creating a feedback loop through the straight-through estimator. We propose NSVQ, a non-stationary-aware VQ training strategy that combines a dense non-stationary embedding loss, codebook replacement, and stage-wise encoder freezing. NSVQ first helps the codebook track encoder drift during early training, then freezes the encoder to consolidate the codebook under a fixed latent geometry, and finally reintroduces adversarial refinement. Experiments on ImageNet-1k show that NSVQ improves reconstruction quality while maintaining full codebook utilization. On ImageNet-1k at 128×128 with 65,536 codes, NSVQ reduces rFID from 2.39 to 2.10 compared with SimVQ, while both methods maintain 100% utilization. Additional latent diffusion experiments show that NSVQ also improves downstream ImageNet generation FID.
Hao Lu, Yongxin Guo, Onur Koyun +3
Wake Forest University School of Medicine / Advocate Health · Wake Forest University School of Medicine · Advocate Health
Vector quantization is a fundamental tool for compressing high-dimensional embeddings, yet existing multi-codebook methods rely on static codebooks that limit expressiveness under heterogeneous data geometry. While recent dynamic quantizers like QINCo adapt codebooks to individual inputs and improve expressiveness, their strict sequential dependencies create decoding bottlenecks. We propose Residual Quantization via Mixture of Experts (RQ-MoE), a framework combining a two-level MoE with dual-stream quantization to enable input-dependent codebook adaptation for efficient vector quantization. RQ-MoE enables dynamic codebook construction and decouples instruction from quantization, facilitating parallel decoding. Theoretically, we show that standard Residual Quantization and QINCo can be recovered as constrained special cases of RQ-MoE, and derive a guideline for setting expert dimensionality in RQ-MoE. Extensive experiments show that RQ-MoE achieves state-of-the-art or on-par performance in reconstruction and retrieval, while providing 6x-14x faster decoding than prior vector quantization methods. The implementation is available at https://github.com/KDEGroup/RQ-MoE.
Zhengjia Zhong, Shuyan Ke, Zaizhou Lin +3
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, Xiamen, China.