cs.CVMar 9, 2026

Evaluating Generative Models via One-Dimensional Code Distributions

Authors: Zexi Jia, Pengcheng Luo, Yijia Zhong, Jinchao Zhang, Jie Zhou

Organizations: WeChat AI, Tencent Inc., China · School of Intelligence Science and Technology, Peking University · College of Computer Science and Artificial Intelligence, Fudan University

Abstract

Most evaluations of generative models rely on feature-distribution metrics such as FID, which operate on continuous recognition features that are explicitly trained to be invariant to appearance variations, and thus discard cues critical for perceptual quality. We instead evaluate models in the space of discrete visual tokens, where modern 1D image tokenizers compactly encode both semantic and perceptual information and quality manifests as predictable token statistics. We introduce Codebook Histogram Distance (CHD), a training-free distribution metric in token space, and Code Mixture Model Score (CMMS), a no-reference quality metric learned from synthetic degradations of token sequences. To stress-test metrics under broad distribution shifts, we further propose VisForm, a benchmark of 210K images spanning 62 visual forms and 12 generative models with expert annotations. Across AGIQA, HPDv2/3, and VisForm, our token-based metrics achieve state-of-the-art correlation with human judgments. We will release all code and datasets to facilitate future research, with the code publicly available at https://github.com/zexiJia/1d-Distance.

Figures & tables

Explore similar work

CardsList
  1. RA-ClipScore: Making Generative Model Evaluation More Interpretable

    Aug 12, 2026Yifan Lu, Taras Kucherenko, Hedvig Kjellström +1Generative Models

  2. MIND: Monge Inception Distance for Generative Models Evaluation

    May 7, 2026Quentin Berthet, Yu-Han Wu, Clement Crepy +3Fréchet Inception DistanceGenerative Models