Organizations: WeChat AI, Tencent Inc., China · School of Intelligence Science and Technology, Peking University · College of Computer Science and Artificial Intelligence, Fudan University
Most evaluations of generative models rely on feature-distribution metrics such as FID, which operate on continuous recognition features that are explicitly trained to be invariant to appearance variations, and thus discard cues critical for perceptual quality. We instead evaluate models in the space of discrete visual tokens, where modern 1D image tokenizers compactly encode both semantic and perceptual information and quality manifests as predictable token statistics. We introduce Codebook Histogram Distance (CHD), a training-free distribution metric in token space, and Code Mixture Model Score (CMMS), a no-reference quality metric learned from synthetic degradations of token sequences. To stress-test metrics under broad distribution shifts, we further propose VisForm, a benchmark of 210K images spanning 62 visual forms and 12 generative models with expert annotations. Across AGIQA, HPDv2/3, and VisForm, our token-based metrics achieve state-of-the-art correlation with human judgments. We will release all code and datasets to facilitate future research, with the code publicly available at https://github.com/zexiJia/1d-Distance.
Figures & tables
Figure 1 : From feature distributions to token statistics. Conventional metrics such as Fréchet Inception Distance (FID) operate on continuous semantic features and assume a Gaussian distribution in feature space (left), which makes them insensitive to appearance details (e.g., texture, style) and unreliable on non-Gaussian data such as artistic or medical images. Our approach (right) quantizes images into a discrete vocabulary of 1D tokens and compares empirical token statistics directly.
Figure 2 : Sensitivity of Token Distributions to Image Degradation. To demonstrate how our discrete token space captures perceptual degradations, we apply 10 levels of progressive distortion to a set of 1,000 images and analyze the resulting shifts in their token distributions. As the severity of distortions like Gaussian noise or block shuffling increases (left), a small subset of perceptually-sensitive tokens exhibits consistent and predictable shifts in their distribution (middle). Our Codebook Histogram Distance (CHD) effectively aggregates these subtle changes, showing a robust, monotonic increase with the degradation level across all distortion types (right).
Figure 3 : Code Mixture Model Degradation. CMMS is trained on token sequences obtained from natural images that are progressively corrupted via uniform token injection, semantic fragment swapping, and pixel-space distortions, without any human labels.
Methods
Reference
AGIQA
Spearman ↑
Kendall ↑
N-MSE ↓
AttnGAN
DALLE2
Glide
Midjourney
SD-1.5
SD-XL
Human ↑
–
0.986
2.624
1.092
3.007
2.752
3.298
–
–
–
FID ↓ [ 8 ]
NeurIPS’17
77.7
77.5
101.45
59.45
41.2
78.45
0.771
0.600
0.119
KID ↓ [ 3 ]
ICLR’18
0.031
0.024
0.076
0.033
0.025
0.036
0.486
0.333
0.236
IS ↑ [ 19 ]
NeurIPS’16
13.8
15.8
16.8
20.8
26.6
15.2
0.543
0.467
0.224
CLIP-FID ↓ [ 13 ]
ICLR’23
0.607
0.547
0.676
0.572
0.451
0.656
0.714
0.467
0.170
Table 1 : Evaluation of different generative models on AGIQA [ 14 ] .
Metric
Real
Kolors
Flux
Infinity
SD-XL
Hunyuan
SD-3
SD-2.0
SD-1.4
Glide
Spearman ↑
Kendall ↑
N-MSE ↓
Human ↑
11.48
10.55
10.43
10.26
8.20
8.19
5.31
-0.24
-3.27
-7.46
–
–
–
FID ↓ [ 8 ]
24.7
41.2
35.3
36.8
35.7
35.7
30.5
53.9
41.6
64.1
0.648
0.467
0.043
IS ↑ [ 19 ]
26.0
27.5
30.1
27.0
29.9
22.5
32.3
13.7
24.7
20.3
0.491
0.289
0.085
KID ↓ [ 3 ]
0.010
0.022
0.018
0.020
0.019
0.017
0.015
0.027
0.021
0.042
0.515
0.333
0.045
CLIP-FID ↓ [ 13 ]
0.264
0.328
0.276
0.299
0.306
0.297
0.253
0.385
0.328
0.447
0.491
0.378
0.043
DINO-FID ↓ [ 20 ]
171.0
216.4
196.9
160.3
328.1
282.3
268.9
544.7
290.0
527.7
0.782
0.556
0.045
Table 2 : Evaluation of different generative models on HPDv3 [ 15 ] .
Figure 4 : Metric–human correlation on VisForm across models and domains. All metrics are normalized to [0,1] , higher is better.
Preference Model
AGIQA
HPDv2
HPDv3
VisForm
CLIPScore [ 7 ]
63.4
65.1
48.6
58.2
MUSIQ [ 11 ]
51.3
52.8
39.4
47.2
CLIP-IQA [ 22 ]
63.8
65.5
48.9
58.6
QualiCLIP [ 1 ]
66.7
68.6
51.2
61.3
MDIQA [ 28 ]
66.3
70.1
51.1
64.5
DeQA-Score [ 29 ]
68.7
70.6
52.7
63.1
Table 3 : Preference prediction on human preference benchmarks.
Setting
AGIQA
HPDv2
HPDv3
VisForm
CHD / N-MSE ↓
Tokenizer Architecture
VQ-VAE (2D tokens) [ 21 ]
0.268
0.152
0.114
0.147
VQGAN (2D tokens) [ 4 ]
0.245
0.145
0.123
0.139
Instella-T2I (1D tokens) [ 23 ]
0.118
0.028
0.019
0.027
TiTok (1D tokens) [ 30 ]
0.112
0.030
0.017
0.024
Table 4 : Ablation study of CHD (N-MSE ↓ ) and CMMS (Acc ↑ ).
Figure 5 : Mean CHD and FID values versus sample size. CHD converges with roughly 1,000 images, while FID needs over 10,000 samples to stabilize.