Organizations: WeChat AI, Tencent Inc., China · School of Intelligence Science and Technology, Peking University · College of Computer Science and Artificial Intelligence, Fudan University
Most evaluations of generative models rely on feature-distribution metrics such as FID, which operate on continuous recognition features that are explicitly trained to be invariant to appearance variations, and thus discard cues critical for perceptual quality. We instead evaluate models in the space of discrete visual tokens, where modern 1D image tokenizers compactly encode both semantic and perceptual information and quality manifests as predictable token statistics. We introduce Codebook Histogram Distance (CHD), a training-free distribution metric in token space, and Code Mixture Model Score (CMMS), a no-reference quality metric learned from synthetic degradations of token sequences. To stress-test metrics under broad distribution shifts, we further propose VisForm, a benchmark of 210K images spanning 62 visual forms and 12 generative models with expert annotations. Across AGIQA, HPDv2/3, and VisForm, our token-based metrics achieve state-of-the-art correlation with human judgments. We will release all code and datasets to facilitate future research, with the code publicly available at https://github.com/zexiJia/1d-Distance.
Figures & tables
Figure 1 : From feature distributions to token statistics. Conventional metrics such as Fréchet Inception Distance (FID) operate on continuous semantic features and assume a Gaussian distribution in feature space (left), which makes them insensitive to appearance details (e.g., texture, style) and unreliable on non-Gaussian data such as artistic or medical images. Our approach (right) quantizes images into a discrete vocabulary of 1D tokens and compares empirical token statistics directly.
Figure 2 : Sensitivity of Token Distributions to Image Degradation. To demonstrate how our discrete token space captures perceptual degradations, we apply 10 levels of progressive distortion to a set of 1,000 images and analyze the resulting shifts in their token distributions. As the severity of distortions like Gaussian noise or block shuffling increases (left), a small subset of perceptually-sensitive tokens exhibits consistent and predictable shifts in their distribution (middle). Our Codebook Histogram Distance (CHD) effectively aggregates these subtle changes, showing a robust, monotonic increase with the degradation level across all distortion types (right).
Figure 3 : Code Mixture Model Degradation. CMMS is trained on token sequences obtained from natural images that are progressively corrupted via uniform token injection, semantic fragment swapping, and pixel-space distortions, without any human labels.
Methods
Reference
AGIQA
Spearman ↑
Kendall ↑
N-MSE ↓
AttnGAN
DALLE2
Glide
Midjourney
SD-1.5
SD-XL
Human ↑
–
0.986
2.624
1.092
3.007
2.752
3.298
–
–
–
FID ↓ [ 8 ]
NeurIPS’17
77.7
77.5
101.45
59.45
41.2
78.45
0.771
0.600
0.119
KID ↓ [ 3 ]
ICLR’18
0.031
0.024
0.076
0.033
0.025
0.036
0.486
0.333
0.236
IS ↑ [ 19 ]
NeurIPS’16
13.8
15.8
16.8
20.8
26.6
15.2
0.543
0.467
0.224
CLIP-FID ↓ [ 13 ]
ICLR’23
0.607
0.547
0.676
0.572
0.451
0.656
0.714
0.467
0.170
Table 1 : Evaluation of different generative models on AGIQA [ 14 ] .
Metric
Real
Kolors
Flux
Infinity
SD-XL
Hunyuan
SD-3
SD-2.0
SD-1.4
Glide
Spearman ↑
Kendall ↑
N-MSE ↓
Human ↑
11.48
10.55
10.43
10.26
8.20
8.19
5.31
-0.24
-3.27
-7.46
–
–
–
FID ↓ [ 8 ]
24.7
41.2
35.3
36.8
35.7
35.7
30.5
53.9
41.6
64.1
0.648
0.467
0.043
IS ↑ [ 19 ]
26.0
27.5
30.1
27.0
29.9
22.5
32.3
13.7
24.7
20.3
0.491
0.289
0.085
KID ↓ [ 3 ]
0.010
0.022
0.018
0.020
0.019
0.017
0.015
0.027
0.021
0.042
0.515
0.333
0.045
CLIP-FID ↓ [ 13 ]
0.264
0.328
0.276
0.299
0.306
0.297
0.253
0.385
0.328
0.447
0.491
0.378
0.043
DINO-FID ↓ [ 20 ]
171.0
216.4
196.9
160.3
328.1
282.3
268.9
544.7
290.0
527.7
0.782
0.556
0.045
Table 2 : Evaluation of different generative models on HPDv3 [ 15 ] .
Figure 4 : Metric–human correlation on VisForm across models and domains. All metrics are normalized to [0,1] , higher is better.
Preference Model
AGIQA
HPDv2
HPDv3
VisForm
CLIPScore [ 7 ]
63.4
65.1
48.6
58.2
MUSIQ [ 11 ]
51.3
52.8
39.4
47.2
CLIP-IQA [ 22 ]
63.8
65.5
48.9
58.6
QualiCLIP [ 1 ]
66.7
68.6
51.2
61.3
MDIQA [ 28 ]
66.3
70.1
51.1
64.5
DeQA-Score [ 29 ]
68.7
70.6
52.7
63.1
Table 3 : Preference prediction on human preference benchmarks.
Setting
AGIQA
HPDv2
HPDv3
VisForm
CHD / N-MSE ↓
Tokenizer Architecture
VQ-VAE (2D tokens) [ 21 ]
0.268
0.152
0.114
0.147
VQGAN (2D tokens) [ 4 ]
0.245
0.145
0.123
0.139
Instella-T2I (1D tokens) [ 23 ]
0.118
0.028
0.019
0.027
TiTok (1D tokens) [ 30 ]
0.112
0.030
0.017
0.024
Table 4 : Ablation study of CHD (N-MSE ↓ ) and CMMS (Acc ↑ ).
Figure 5 : Mean CHD and FID values versus sample size. CHD converges with roughly 1,000 images, while FID needs over 10,000 samples to stabilize.
Generative models can produce images nearly indistinguishable from real data, yet rigorous and interpretable evaluation remains challenging. Conventional metrics such as FID provide only scalar scores with limited diagnostic insight. Widely adopted CLIP-based metrics enable semantic evaluation beyond simple training class labels, but inherit limitations from CLIP's training paradigm that restrict attribute-wise analysis. We propose RA-CLIPScore, a novel metric that mitigates these issues and extends CLIP-based evaluation to spatial distribution alignment, measuring whether generated objects adhere to the positional priors found in the training data. RA-CLIPScore introduces dual prompts to decouple competing attributes and leverages local patch tokens to capture fine-grained regional semantics. We evaluate image generative models on their ability to match both attribute and spatial distributions of the training data. Extensive experiments show that RA-CLIPScore provides more robust and interpretable evaluations than prior methods, particularly under distribution misalignment or partially irrelevant textual attributes. We further demonstrate how it reveals spatial biases in generative models. User evaluations confirm that Regional Single Attribute Divergence based on our RA-CLIPScore aligns more closely with human perception of visual diversity than existing semantic metrics.
Yifan Lu, Taras Kucherenko, Hedvig Kjellström +1
KTH Royal Institute of Technology · National Library of Sweden · Most of the work was performed while Taras was at Electronic Arts (EA). +3
Diffusion and continuous flow-based language models have emerged as the leading non-autoregressive alternatives to language modeling. Progress in both paradigms is overwhelmingly tracked by generative perplexity (gen-PPL): the per-token negative log-likelihood of samples under a frozen autoregressive (AR) scorer such as gpt2-large, typically paired with an empirical-entropy guardrail to rule out low-entropy collapse. We argue that this metric is unsound. By construction, gen-PPL measures only predictability under the scoring AR, not grammaticality or semantic coherence -- and the set of predictable but still low-quality sequences is combinatorially large. To make this concrete, we construct a suite of zero-parameter, deliberately naive samplers that achieve state-of-the-art gen-PPL on LM1B and OpenWebText at non-degenerate entropy, surpassing recently published diffusion and continuous-flow models while producing text that is incoherent by construction. We recommend evaluation suites that directly quantify the distributional divergence between generated and reference text, and use such a suite to re-benchmark recent non-autoregressive models, recovering a more faithful picture of the current state of the art.
We propose the Monge Inception Distance (MIND), a metric for evaluating generative models that addresses key limitations of the widely adopted Fréchet Inception Distance (FID). The MIND metric leverages the sliced Wasserstein distance to compare distributions by averaging one-dimensional optimal transport distances, efficiently computed via sorting. This approach circumvents the estimation of high-dimensional means and covariance matrices, which underlie FID's poor sample complexity and vulnerability to adversarial attacks. We empirically demonstrate three primary advantages: (i) it is more sample-efficient by one order of magnitude, (ii) it is faster to compute by two orders of magnitude, (iii) it is more robust to adversarial attacks such as moment-matching. We show that MIND with 5k samples can replace the evaluation performance of FID with 50k samples, providing high correlation with this standard benchmark and superior discriminative performance. We further demonstrate that even smaller sample sizes (e.g., 1k or 2k) remain highly informative for rapid model iteration.