TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images
Authors: Kirill Koltsov, Aleksandr Gushchin, Anastasia Antsiferova, Dmitriy Vatolin
Organizations: Lomonosov Moscow State University Moscow, Russia · ISP RAS Research Center for Trusted Artificial Intelligence Lomonosov Moscow State University Moscow, Russia · ISP RAS Research Center for Trusted Artificial Intelligence MSU Institute for Artificial Intelligence Moscow, Russia · MSU Institute for Artificial Intelligence Lomonosov Moscow State University Moscow, Russia
Recent text-to-image models have improved global realism, but text rendering remains a persistent failure mode: images may look convincing overall, yet local typography often contains malformed glyphs, broken strokes, irregular spacing, and other artifacts that humans heavily penalize. We formulate Text-in-Image Quality Assessment (TIQA), a no-reference task that estimates a human-aligned perceptual quality score for detected text regions while disentangling visual text quality from semantic correctness. To support this setting, we introduce two datasets. TIQA-Crops contains 120k text crops from 36k AI-generated images produced by 12 generators, with 10k mean-opinion-score (MOS) labels and 110k proxy labels for pretraining. TIQA-Images contains 1,500 text-heavy images from 10 recent generators, including proprietary systems, with paired overall-quality and text-quality subjective scores. We also propose ANTIQA, a lightweight predictor with text-specific inductive biases. Across crop-level and image-level evaluations, ANTIQA achieves the best alignment with human judgments, reaching PLCC/SROCC of 0.942/0.935 on TIQA-Crops and 0.842/0.837 for text-quality MOS on unseen generators in TIQA-Images. In best-of-5 AI-generated image ranking, ANTIQA improves the text quality of the selected image by 0.36 MOS (14%), demonstrating utility for benchmarking, filtering, and generation-time selection. Together, these findings establish perceptual text quality as a distinct evaluation target for modern text-to-image generation.
Figures & tables
Figure 1. ANTIQA in action. Left: text-containing fragments from images generated by SOTA generators. Right: ANTIQA’s output — a Perceptual Quality score predicted for each crop. TODO
Figure 2. Overview of Text-in-Image Quality Assessment (TIQA). Left: AI-generated images contain multiple text regions that are detected and cropped. Middle: a TIQA model predicts a scalar text-quality score for each crop, trained on mean opinion scores (MOS). Right: representative model families used as baselines (VLM judges, OCR confidence, generic IQA) and the proposed specialized TIQA model. Bottom: example applications of TIQA for measuring generator quality, filtering candidates in production pipelines (best-of-K), and optimizing generation via reranking or closed-loop control. TODO
Figure 3. ANTIQA architecture. Each text crop is converted to grayscale, concatenated with a Sobel edge map, and then processed by a lightweight multi-scale CNN with residual stages and downsampling. Features from multiple resolutions are pooled to fixed grids using adaptive average and max pooling, fused via an MLP head, and regressed to a single MOS prediction. TODO
Type
Model
TIQA-Crops
TIQA-Images (OQ-MOS)
TIQA-Images (TQ-MOS)
Params
Speed (FPS)
PLCC ↑
SROCC ↑
PLCC ↑
SROCC ↑
PLCC ↑
SROCC ↑
Generic supervised models
ResNet50
0.917
0.920
0.735
0.732
0.728
0.731
25.6 M
220.4
ViT
0.926
0.927
0.740
0.738
0.734
0.735
86.6 M
244.7
IQA
TOPIQ
0.401
0.414
0.615
0.568
0.493
0.470
45.2 M
66.7
TOPIQ ⋆
0.870
0.879
0.752
0.754
0.748
0.749
45.2 M
66.7
HyperIQA
0.622
0.668
0.607
0.592
0.501
0.497
27.4 M
97.4
Table 1. Performance on TIQA-Crops (crop-level) and TIQA-Images (image-level) measured by PLCC/SROCC with human MOS. “ ⋆ ” denotes finetuned IQA models. Speed is computed on 256×256 images on NVIDIA A100 GPU. For VLMs, “A x B” denotes x active parameters out of the MoE total.
Type
Model
Within-group correlation to MOS (mean (std)) ↑
Best-of-5 selection outcome (mean MOS / gain)
TQ-MOS
OQ-MOS
Selected MOS
Δ MOS vs Random (gap closed)
PLCC
SROCC
PLCC
SROCC
TQ
OQ
TQ
OQ
Reference
Random
—
—
—
—
2.57
3.01
+0.00 (0%)
+0.00 (0%)
Oracle
—
—
—
—
3.07
3.47
+0.50 (100%)
+0.46 (100%)
Generic supervised models
ResNet50
0.351 (0.112)
0.364 (0.140)
0.260 (0.127)
0.265 (0.130)
2.69
3.06
+0.12 (24%)
+0.05 (10.9%)
ViT
0.342 (0.128)
0.359 (0.119)
0.274 (0.113)
0.259 (0.109)
2.70
3.08
+0.13 (26%)
+0.07 (14.1%)
Table 2. Best-of- K ranking/selection on TIQA-Images. Within-group PLCC/SROCC to MOS are averaged over groups (mean ± std). Selection reports the mean MOS of the top-scored image per group and Δ over Random. Gap closed is the improvement over Random relative to Oracle. “ ⋆ ” denotes finetuned IQA models. “*” denotes text-only masked image. “**” denotes separate crops as input and averaging their scores. Random is averaged over 1,000 runs; Oracle is max value over each group.
Figure 4. Box-plot distributions of OQ-MOS and TQ-MOS for separate generators. The models are sorted by mean TQ-MOS. TODO
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Aggregation
TOPIQ
PaddleOCR
Qwen3
ANTIQA
Average
PLCC ↑
SROCC ↑
PLCC ↑
SROCC ↑
PLCC ↑
SROCC ↑
PLCC ↑
SROCC ↑
PLCC ↑
SROCC ↑
Simple area-weighted mean
0.493
0.470
0.761
0.787
0.489
0.510
0.842
0.837
0.646
0.651
Area α -weighted mean
0.512
0.483
0.780
0.792
0.503
0.527
0.844
0.841
0.660
0.661
Coverage-aware blend
0.241
0.204
0.476
0.455
0.213
0.221
0.511
0.517
0.360
0.349
Softmin (log-sum-exp)
0.485
0.466
0.749
0.771
0.463
0.490
0.819
0.813
0.629
0.635
Bottom- k mean
0.461
0.459
0.712
0.756
0.448
0.475
0.793
0.802
0.604
0.623
Appendix
Table 3. Effect of aggregation on image-level correlation. Correlations were computed on the whole TIQA-Images dataset agains TQ-MOS (text-only quality score). Best score is bolded. Higher is better ( ↑ ).
Figure 8Figure 9
Figure 5. An example of gradual crop distortion, from left to right.
Figure 6. Binned mean normalized Levenshtein similarity as a function of predicted by TIQA models score.
Correlation level
# points
SROCC
Pooled (all images)
GPK=1500
0.78
Between generators (means)
G=10
0.98
Within (prompt, generator)
S=300
0.51 ± 0.43
(median)
S=300
0.59
Appendix
Table 4. Decomposing the correlation between human overall quality (OQ-MOS) and text quality (TQ-MOS) on TIQA-Images. TIQA-Images contains P=30 prompts, G=10 generators, and K=5 seeds per (prompt, generator) pair.
Figure 7. Prompt and seed dependencies across text-to-image models ON TIQA-Images. Colour encodes the mean TQ-MOS per prompt, and marker size encodes the standard deviation across five seed generations, capturing within-prompt sampling variability.
Variant
PLCC
SROCC
Full ANTIQA (proposed)
0.942
0.935
w/o OCR-confidence pretraining
0.887
0.881
w/o neural optimal transport mapping (use raw OCR confidence)
0.933
0.927
w/o Sobel edge-map input (grayscale only)
0.920
0.923
w/o strip conv blocks (use standard conv)
0.890
0.876
w/o SE channel gating
0.905
0.908
Appendix
Table 5. ANTIQA ablations on TIQA-Crops. Full-model numbers are from Table 1 in the paper; the provided PDF does not include the ablation-result values.
TQ–MOS
OQ–MOS
Prompt
PLCC
SROCC
PLCC
SROCC
1
0.6785
0.6912
0.6628
0.6866
2
0.4342
0.4983
0.4772
0.5392
3
0.6749
0.6505
0.7043
0.7066
4
0.5683
0.6004
0.5706
0.5708
Appendix
Table 6. Correlation of Qwen3 with MOS scores using different prompts on 60 images.
#
TIQA-Crops
TIQA-Images
1
PixArt Alpha ( Chen et al., 2023b )
GPT Image 1.5 ( OpenAI, 2025 )
2
SD 3.5 Large Turbo ( AI, 2024c )
FLUX1.1 [pro] ( Labs, 2024 )
3
SD 2.1 ( AI, 2022b )
FLUX.2 [max] ( Labs, 2025 )
4
PixArt Sigma ( Chen et al., 2024c )
Seedream 4.5 ( Seedream, 2025 )
5
SD 3.5 Medium ( AI, 2024b )
Ideogram 3.0 Turbo ( Ideogram, 2025 )
6
Kandinsky 2 ( Razzhigaev et al., 2023 )
Imagen 4 Fast ( DeepMind, 2025 )
Appendix
Table 7. Generators used for TIQA datasets
Figure 8. TIQA-Crops examples
Figure 9. MOS distributions (1–5) for three target strings on the OCR-correct subset (exact transcript match). Similar distributions across a real word, an anagram, and a nonword indicate limited sensitivity of visual-quality ratings to lexical plausibility when rendering is correct.
Figure 10. Distibution plot for OQ-MOS and TQ-MOS for TIQA-Images dataset. Vertical dashed lines denote mean values across all images.
Figure 11. Representative visual examples of the rating categories shown to annotators during the TIQA-Crops labeling task. Each row corresponds to a different score on the 0 – 5 scale and serves as a reference anchor for the rater.
Figure 21Figure 22Figure 23
Figure 12. TIQA-Images examples for overall quality TODO
Figure 13. TIQA-Images examples for text quality TODO
Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no longer satisfy users' pressing demands for faithful real-world reconstruction and genuine creative expression. Existing benchmarks, however, remain anchored in these foundational criteria and do not yet capture the nuanced capabilities that matter in authentic artistic practice, making it difficult to reliably distinguish state-of-the-art T2I models. To address the gap, we introduce Qwen-Image-Bench, a creator-centric benchmark co-designed with professional artists and grounded in real-world creation scenarios. Qwen-Image-Bench enriches conventional evaluation with two application-driven dimensions: Real-world Fidelity and Creative Generation. Drawing on the staged reasoning inherent in professional artistic workflows, we organize these five pillars into a top-down hierarchical taxonomy that further decomposes into 23 second-level sub-capabilities and 56 third-level verifiable rubrics. To ensure broad coverage, we curate 1000 stratified prompts with each prompt jointly exercising more than four fine-grained facets across multiple pillars. We train a unified judge model Q-Judger based on Qwen3.6-27B, supervised by 80 professional annotators from global art academies under blind labeling and triple-review protocols, that scores every image across all 56 verifiable facets, producing fine-grained, rubric-grounded, and fully attributable diagnostics rather than a single opaque score. Empirically, Qwen-Image-Bench reliably distinguishes leading T2I models, achieving the greatest separation on the two application-driven dimensions of Real-world Fidelity and Creative Generation where existing benchmarks provide little insight, while also providing a trustworthy optimization signal for production-level T2I development.
Text-to-Image (T2I) generation is primarily driven by Diffusion Models (DM) which rely on random Gaussian noise. Thus, like playing the slots at a casino, a DM will produce different results given the same user-defined inputs. This imposes a gambler's burden: To perform multiple generation cycles to obtain a satisfactory result. However, even though DMs use stochastic sampling to seed generation, the distribution of generated content quality highly depends on the prompt and the generative ability of a DM with respect to it. To account for this, we propose Naïve PAINE for improving the generative quality of Diffusion Models by leveraging T2I preference benchmarks. We directly predict the numerical quality of an image from the initial noise and given prompt. Naïve PAINE then selects a handful of quality noises and forwards them to the DM for generation. Further, Naïve PAINE provides feedback on the DM generative quality given the prompt and is lightweight enough to seamlessly fit into existing DM pipelines. Experimental results demonstrate that Naïve PAINE outperforms existing approaches on several prompt corpus benchmarks.
Joong Ho Kim, Nicholas Thai, Souhardya Saha Dip +2
Recent advances in text-to-image (T2I) generation have led to models capable of producing highly realistic images. Yet, reliably evaluating their outputs remains challenging, especially at scale. Existing automatic evaluators, often relying on a static prompt set, struggle to capture subtle failure modes such as partial prompt misalignment, compositional errors, or visually plausible but semantically incorrect generations. In this work, we introduce DynEval, a Dynamic Evaluation framework designed to jointly assess text-to-image alignment and image quality of T2I models. To support scalable training beyond limited human-annotated data, we construct two large datasets. First, we build GenDB, a collection of 500K prompt-image pairs generated from human-written prompts drawn from DiffusionDB using a tiered prompt-model generation strategy. Second, building upon GenDB, we construct DynEvalInstruct, a 250K instruction dataset comprising prompt-image-response triplets distilled from a structured evaluation pipeline that decomposes evaluation into text-image alignment and visual quality reasoning. Using this dataset, we perform full fine-tuning of a compact evaluator through a curriculum learning strategy to effectively distill the superior evaluation capabilities of a larger teacher vision-language model, resulting in DynEval-2B and DynEval-4B. In extensive comparisons against existing evaluators across 11 benchmarks, our evaluator achieves a higher overall correlation with human judgments. Furthermore, it provides fine-grained analysis of the capabilities and failure modes of 36 T2I models across 42 subcategories and 9 semantic dimensions.
Shyam Marjit, Dheeraj Baiju, Anuj Shikarkhane +3
Indian Institute of Science · VCL, IISc · Hugging Face