cs.CVMar 7, 2026

TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images

Authors: Kirill Koltsov, Aleksandr Gushchin, Anastasia Antsiferova, Dmitriy Vatolin

Organizations: Lomonosov Moscow State University Moscow, Russia · ISP RAS Research Center for Trusted Artificial Intelligence Lomonosov Moscow State University Moscow, Russia · ISP RAS Research Center for Trusted Artificial Intelligence MSU Institute for Artificial Intelligence Moscow, Russia · MSU Institute for Artificial Intelligence Lomonosov Moscow State University Moscow, Russia

Abstract

Recent text-to-image models have improved global realism, but text rendering remains a persistent failure mode: images may look convincing overall, yet local typography often contains malformed glyphs, broken strokes, irregular spacing, and other artifacts that humans heavily penalize. We formulate Text-in-Image Quality Assessment (TIQA), a no-reference task that estimates a human-aligned perceptual quality score for detected text regions while disentangling visual text quality from semantic correctness. To support this setting, we introduce two datasets. TIQA-Crops contains 120k text crops from 36k AI-generated images produced by 12 generators, with 10k mean-opinion-score (MOS) labels and 110k proxy labels for pretraining. TIQA-Images contains 1,500 text-heavy images from 10 recent generators, including proprietary systems, with paired overall-quality and text-quality subjective scores. We also propose ANTIQA, a lightweight predictor with text-specific inductive biases. Across crop-level and image-level evaluations, ANTIQA achieves the best alignment with human judgments, reaching PLCC/SROCC of 0.942/0.935 on TIQA-Crops and 0.842/0.837 for text-quality MOS on unseen generators in TIQA-Images. In best-of-5 AI-generated image ranking, ANTIQA improves the text quality of the selected image by 0.36 MOS (14%), demonstrating utility for benchmarking, filtering, and generation-time selection. Together, these findings establish perceptual text quality as a distinct evaluation target for modern text-to-image generation.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation

    May 27, 2026Niantong Li, Guangzheng Hu, Weixu Qiao +35CreativityFidelity

  2. Naïve PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation

    Mar 12, 2026Joong Ho Kim, Nicholas Thai, Souhardya Saha Dip +2Text-To-ImageGenerative Quality

  3. DynEval: Holistic Evaluations of T2I Generative Models in the Wild

    Jul 13, 2026Shyam Marjit, Dheeraj Baiju, Anuj Shikarkhane +3Modern Text-To-Image ModelsText-To-Image