Vector quantization, which discretizes a continuous vector space into a finite set of representative vectors (a codebook), has been widely adopted in modern machine learning. Despite its effectiveness, vector quantization poses a fundamental challenge: the non-differentiable quantization step blocks gradient backpropagation. Smoothed vector quantization addresses this issue by relaxing the discrete selection of a codebook vector into a weighted combination of codebook entries, represented as the matrix product of a simplex vector and the codebook. Effective smoothing requires two properties: (1) smoothed code-selection vectors should remain close to a onehot vector, ensuring tight approximation, and (2) all codebook entries should be utilized, preventing code collapse. Existing methods typically address these desiderata separately. By contrast, the present study introduces a simple and intuitive regularization that promotes both simultaneously by minimizing the distance between each simplex vertex and its K-nearest smoothed code-selectors. Representative benchmarks on image encoding demonstrate that the proposed method achieves more effective codebook utilization and improves performance over prior approaches.
Figures & tables
Figure 1 : ( 1 ) Four different distributions on the simplex Δ3−1 . For effective smoothed vector quantization, samples should be concentrated near the vertices of the simplex (i.e., onehot-like vectors; orange), rather than centered (dark gray) or uniformly spread across the simplex (light gray). At the same time, each vertex must have some samples in its neighborhood to avoid code collapse (blue). ( 1 ) Maximizing the perplexity of the sample mean ( Baevski et al., 2020b, ) penalizes code collapse but cannot discriminate among the other three distributions. ( 1 ) The proposed K -nearest neighbor (KNN) distance minimization ( K=8 ) favors the desired vertex-concentrated distribution while also preventing code collapse.
Feature Map Size; Codebook Size
16 × 16 × 32; 1024
64 × 64 × 3; 8196
64 × 64 × 64; 8196
Method
K/4
Usage ( ↑ )
rMSE ( ↓ )
FID ( ↓ )
IS ( ↑ )
Usage ( ↑ )
rMSE ( ↓ )
FID ( ↓ )
IS ( ↑ )
Usage ( ↑ )
rMSE ( ↓ )
FID ( ↓ )
IS ( ↑ )
STE
Euclid
—
4.5%
0.108
101.84
10.67
100.0%
0.045
10.41
142.12
1.7%
0.067
25.65
82.61
Cos
—
3.0%
0.102
108.09
10.03
70.9%
0.053
18.13
107.38
1.7%
0.054
20.08
101.74
SimVQ
—
100.0%
0.091
86.94
15.93
100.0%
0.046
10.82
141.17
100.0%
0.045
12.54
134.84
RE
Euclid
—
3.1%
0.123
127.40
7.40
78.8%
0.046
14.04
123.30
0.2%
0.077
42.12
51.86
Table 1 : Performance of discrete autoencoding on the ImageNet validation set. Reported metrics are codebook usage and reconstruction quality scores: root mean squared error (rMSE), Fréchet Inception Distance (FID), Inception Score (IS), and Structural Similarity Index Measure (SSIM). The proposed method is denoted as “KNN-L2/CE”. The baseline methods are abbreviated as follows: “Cos” denotes cosine-distance quantization (contrasted with Euclidean distance, ’Euclid’). “HG” and “SG” denote Hard-Gumbel and Soft-Gumbel, i.e., Gumbel-softmax sampling with and without straight-through hard quantization in the forward pass, respectively. Best scores across all methods are highlighted in boldface, while underlined values indicate the best-performing value of K/4 .
Feature Map Size; Codebook Size
16 × 16 × 32; 1024
64 × 64 × 3; 8192
64 × 64 × 64; 8192
Method
K/4
75%
90%
99%
Max
75%
90%
99%
Max
75%
90%
99%
Max
PPL
—
910.40
933.12
963.29
995.91
7559.89
7561.99
7563.10
7565.02
8186.32
8186.40
8186.55
8186.76
KNN-L2
1
1.17
1.59
2.36
6.26
1.00
1.00
1.00
4.00
9.63
18.10
51.04
123.62
2
1.00
1.08
1.86
27.37
1.00
1.00
1.06
4.65
3.85
6.36
16.53
43.40
4
1.00
1.06
1.86
62.99
1.00
1.00
1.00
4.61
1.06
1.06
1.06
1.07
Table 2 : Tightness of softmax-based smoothing (without Gumbel sampling), measured by the individual perplexity of smoothed code-selection vectors, exp(−∑m=1Mpmlogpm) . Reported values are the 75th, 90th, and 99th percentiles, as well as the maximum, computed across all feature-map pixels in the ImageNet validation set.
Figure 2 : (Top) Histogram of per-image codebook utilization on the validation dataset, defined as the percentage of codes used at least once within each image. (Bottom) Histogram of code popularity, defined as the proportion of validation images in which each code occurs.
Method
Evaluated in
D/G
MG
Usage % ( ↑ )
FID ( ↓ )
IS ( ↑ )
STE Euclid
Esser et al., (2021)
256/1
10241
—
7.94
—
256/1
163841
—
4.98
—
This Work
256/8
10248
0.1–0.2
398.52
1.00
RE Euclid
Fifty et al., (2025)
256/1
10241
98
4.6
146.5
This Work
256/8
10248
0.2–0.8
506.60
1.00
SimVQ
This Work 7 7 7 The original SimVQ paper evaluated the method on 128×128 inputs while keeping a 16×16 latent map ( Zhu et al.,, 2025 ) , which severely reduces the compression rate and makes reconstruction significantly easier. Since those scores are not directly comparable to the standard 256×256 VQGAN benchmark, SimVQ is evaluated here on the same 256×256 pipeline to ensure a rigorous and direct comparison.
32/1
10241
1.0
436.34
1.00
Table 3 : Performance of VQGAN on the ImageNet validation set. Reported metrics are codebook usage and reconstruction quality scores. Best scores across all methods are highlighted in boldface. For the single-codebook STE/RE Euclid baselines ( G=1 ), scores from the original publications are reported, since attempts to reproduce these results were unsuccessful. 6 6 6 Attempts to replicate STE Euclid results using the official repository yielded highly degraded scores; for example, the FID scores were 172.78 for M=1024 and 349.18 for M=16384 . These results are consistent with community-reported issues (e.g., https://github.com/CompVis/taming-transformers/issues/87 ). Furthermore, the pre-trained weights for the RE Euclid model ( Fifty et al.,, 2025 , for replicating Table 2 in) are unavailable in the official repository ( https://github.com/cfifty/rotation_trick/tree/main ). Due to these reproducibility challenges, the “optimal scores” published in the original papers are reported for the single-codebook baselines. All product-quantized configurations ( G=8 ), including these baselines, were newly implemented and evaluated.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
GAN (§ 4.2 )
Autoencoding (§ 4.1 )
Feature Map Size
16×16×32
16×16×32
64×64×3
64×64×64
Codebook Size
1024
8196
8196
8196
Latent Channels
128→128→64→64→32→32
128→64→32→{3,64}
Height & Width
256→128→64→32→16→16
256→128→64→64
Batch Size
64
64
64
64
Training Epochs
60
25
20
20
Appendix
Table 4 : Hyperparameters for discrete autoencoding/GAN.
Feature Map Size; Codebook Size
16 × 16 × 32; 1024
64 × 64 × 3; 8196
64 × 64 × 64; 8196
Wall-Clock Time
Wall-Clock Time
Wall-Clock Time
Method
K/4
Entire
Quant.
VRAM
Entire
Quant.
VRAM
Entire
Quant.
VRAM
(ms)
(ms)
(GB)
(ms)
(ms)
(GB)
(ms)
(ms)
(GB)
STE
STE Euclid
—
330.39
1.35
34.20
508.49
1.35
54.83
502.20
1.30
54.93
SimVQ
—
330.18
1.31
34.20
504.28
1.36
54.83
502.25
1.37
54.92
Appendix
Table 5 : Wall-clock time and peak VRAM consumption per training iteration. The models were run on four NVIDIA A100 GPUs (80GB VRAM per GPU) using DistributedDataParallel of PyTorch. The reported scores represent computational load per GPU on average over a single epoch of the ImageNet training split, with batch size set to 64. The “Entire” wall-clock time includes the forward and backward passes through the encoder, quantizer, and decoder modules, while the ‘Quant.’ reports the wall-clock time consumed by the forward pass through the quantizer alone.
Figure 3 : Dirichlet distributions on the simplex Δ3−1 with concentration parameters α1=α2=α3=α , where α∈{0.5,1.0,2.0} .