cs.CVOct 7, 2026

UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation

Authors: Deyuan Liu, Yihao Hu, Jingxuan Zhang, Xingying Li, Jun Xie, Jiacheng Liu, Jungang Li, Yu Huang, +8 more

Organizations: Westlake University · Ant Group · Zhejiang University · Shanghai Innovation Institute · MBZUAI · HKUST · CityU · Peking University · CASIA · Wechat AI

Abstract

Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: https://github.com/LINs-lab/UltraText_Bench.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

    Oct 1, 2026Yu Huang, Jungang Li, Zhiyuan Wang +8Agentic Video GenerationVideo Generation

  2. WeGenBench: A Multidimensional Diagnostic Benchmark towards Text-to-Image Model Optimization

    Jun 18, 2026Qian Liang, Xiaomin Li, Ying Zhang +6T2I Generation Evaluation