TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images
Authors: Kirill Koltsov, Aleksandr Gushchin, Anastasia Antsiferova, Dmitriy Vatolin
Organizations: Lomonosov Moscow State University Moscow, Russia · ISP RAS Research Center for Trusted Artificial Intelligence Lomonosov Moscow State University Moscow, Russia · ISP RAS Research Center for Trusted Artificial Intelligence MSU Institute for Artificial Intelligence Moscow, Russia · MSU Institute for Artificial Intelligence Lomonosov Moscow State University Moscow, Russia
Recent text-to-image models have improved global realism, but text rendering remains a persistent failure mode: images may look convincing overall, yet local typography often contains malformed glyphs, broken strokes, irregular spacing, and other artifacts that humans heavily penalize. We formulate Text-in-Image Quality Assessment (TIQA), a no-reference task that estimates a human-aligned perceptual quality score for detected text regions while disentangling visual text quality from semantic correctness. To support this setting, we introduce two datasets. TIQA-Crops contains 120k text crops from 36k AI-generated images produced by 12 generators, with 10k mean-opinion-score (MOS) labels and 110k proxy labels for pretraining. TIQA-Images contains 1,500 text-heavy images from 10 recent generators, including proprietary systems, with paired overall-quality and text-quality subjective scores. We also propose ANTIQA, a lightweight predictor with text-specific inductive biases. Across crop-level and image-level evaluations, ANTIQA achieves the best alignment with human judgments, reaching PLCC/SROCC of 0.942/0.935 on TIQA-Crops and 0.842/0.837 for text-quality MOS on unseen generators in TIQA-Images. In best-of-5 AI-generated image ranking, ANTIQA improves the text quality of the selected image by 0.36 MOS (14%), demonstrating utility for benchmarking, filtering, and generation-time selection. Together, these findings establish perceptual text quality as a distinct evaluation target for modern text-to-image generation.
Figures & tables
Figure 1. ANTIQA in action. Left: text-containing fragments from images generated by SOTA generators. Right: ANTIQA’s output — a Perceptual Quality score predicted for each crop. TODO
Figure 2. Overview of Text-in-Image Quality Assessment (TIQA). Left: AI-generated images contain multiple text regions that are detected and cropped. Middle: a TIQA model predicts a scalar text-quality score for each crop, trained on mean opinion scores (MOS). Right: representative model families used as baselines (VLM judges, OCR confidence, generic IQA) and the proposed specialized TIQA model. Bottom: example applications of TIQA for measuring generator quality, filtering candidates in production pipelines (best-of-K), and optimizing generation via reranking or closed-loop control. TODO
Figure 3. ANTIQA architecture. Each text crop is converted to grayscale, concatenated with a Sobel edge map, and then processed by a lightweight multi-scale CNN with residual stages and downsampling. Features from multiple resolutions are pooled to fixed grids using adaptive average and max pooling, fused via an MLP head, and regressed to a single MOS prediction. TODO
Type
Model
TIQA-Crops
TIQA-Images (OQ-MOS)
TIQA-Images (TQ-MOS)
Params
Speed (FPS)
PLCC ↑
SROCC ↑
PLCC ↑
SROCC ↑
PLCC ↑
SROCC ↑
Generic supervised models
ResNet50
0.917
0.920
0.735
0.732
0.728
0.731
25.6 M
220.4
ViT
0.926
0.927
0.740
0.738
0.734
0.735
86.6 M
244.7
IQA
TOPIQ
0.401
0.414
0.615
0.568
0.493
0.470
45.2 M
66.7
TOPIQ ⋆
0.870
0.879
0.752
0.754
0.748
0.749
45.2 M
66.7
HyperIQA
0.622
0.668
0.607
0.592
0.501
0.497
27.4 M
97.4
Table 1. Performance on TIQA-Crops (crop-level) and TIQA-Images (image-level) measured by PLCC/SROCC with human MOS. “ ⋆ ” denotes finetuned IQA models. Speed is computed on 256×256 images on NVIDIA A100 GPU. For VLMs, “A x B” denotes x active parameters out of the MoE total.
Type
Model
Within-group correlation to MOS (mean (std)) ↑
Best-of-5 selection outcome (mean MOS / gain)
TQ-MOS
OQ-MOS
Selected MOS
Δ MOS vs Random (gap closed)
PLCC
SROCC
PLCC
SROCC
TQ
OQ
TQ
OQ
Reference
Random
—
—
—
—
2.57
3.01
+0.00 (0%)
+0.00 (0%)
Oracle
—
—
—
—
3.07
3.47
+0.50 (100%)
+0.46 (100%)
Generic supervised models
ResNet50
0.351 (0.112)
0.364 (0.140)
0.260 (0.127)
0.265 (0.130)
2.69
3.06
+0.12 (24%)
+0.05 (10.9%)
ViT
0.342 (0.128)
0.359 (0.119)
0.274 (0.113)
0.259 (0.109)
2.70
3.08
+0.13 (26%)
+0.07 (14.1%)
Table 2. Best-of- K ranking/selection on TIQA-Images. Within-group PLCC/SROCC to MOS are averaged over groups (mean ± std). Selection reports the mean MOS of the top-scored image per group and Δ over Random. Gap closed is the improvement over Random relative to Oracle. “ ⋆ ” denotes finetuned IQA models. “*” denotes text-only masked image. “**” denotes separate crops as input and averaging their scores. Random is averaged over 1,000 runs; Oracle is max value over each group.
Figure 4. Box-plot distributions of OQ-MOS and TQ-MOS for separate generators. The models are sorted by mean TQ-MOS. TODO
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Aggregation
TOPIQ
PaddleOCR
Qwen3
ANTIQA
Average
PLCC ↑
SROCC ↑
PLCC ↑
SROCC ↑
PLCC ↑
SROCC ↑
PLCC ↑
SROCC ↑
PLCC ↑
SROCC ↑
Simple area-weighted mean
0.493
0.470
0.761
0.787
0.489
0.510
0.842
0.837
0.646
0.651
Area α -weighted mean
0.512
0.483
0.780
0.792
0.503
0.527
0.844
0.841
0.660
0.661
Coverage-aware blend
0.241
0.204
0.476
0.455
0.213
0.221
0.511
0.517
0.360
0.349
Softmin (log-sum-exp)
0.485
0.466
0.749
0.771
0.463
0.490
0.819
0.813
0.629
0.635
Bottom- k mean
0.461
0.459
0.712
0.756
0.448
0.475
0.793
0.802
0.604
0.623
Appendix
Table 3. Effect of aggregation on image-level correlation. Correlations were computed on the whole TIQA-Images dataset agains TQ-MOS (text-only quality score). Best score is bolded. Higher is better ( ↑ ).
Figure 8Figure 9
Figure 5. An example of gradual crop distortion, from left to right.
Figure 6. Binned mean normalized Levenshtein similarity as a function of predicted by TIQA models score.
Correlation level
# points
SROCC
Pooled (all images)
GPK=1500
0.78
Between generators (means)
G=10
0.98
Within (prompt, generator)
S=300
0.51 ± 0.43
(median)
S=300
0.59
Appendix
Table 4. Decomposing the correlation between human overall quality (OQ-MOS) and text quality (TQ-MOS) on TIQA-Images. TIQA-Images contains P=30 prompts, G=10 generators, and K=5 seeds per (prompt, generator) pair.
Figure 7. Prompt and seed dependencies across text-to-image models ON TIQA-Images. Colour encodes the mean TQ-MOS per prompt, and marker size encodes the standard deviation across five seed generations, capturing within-prompt sampling variability.
Variant
PLCC
SROCC
Full ANTIQA (proposed)
0.942
0.935
w/o OCR-confidence pretraining
0.887
0.881
w/o neural optimal transport mapping (use raw OCR confidence)
0.933
0.927
w/o Sobel edge-map input (grayscale only)
0.920
0.923
w/o strip conv blocks (use standard conv)
0.890
0.876
w/o SE channel gating
0.905
0.908
Appendix
Table 5. ANTIQA ablations on TIQA-Crops. Full-model numbers are from Table 1 in the paper; the provided PDF does not include the ablation-result values.
TQ–MOS
OQ–MOS
Prompt
PLCC
SROCC
PLCC
SROCC
1
0.6785
0.6912
0.6628
0.6866
2
0.4342
0.4983
0.4772
0.5392
3
0.6749
0.6505
0.7043
0.7066
4
0.5683
0.6004
0.5706
0.5708
Appendix
Table 6. Correlation of Qwen3 with MOS scores using different prompts on 60 images.
#
TIQA-Crops
TIQA-Images
1
PixArt Alpha ( Chen et al., 2023b )
GPT Image 1.5 ( OpenAI, 2025 )
2
SD 3.5 Large Turbo ( AI, 2024c )
FLUX1.1 [pro] ( Labs, 2024 )
3
SD 2.1 ( AI, 2022b )
FLUX.2 [max] ( Labs, 2025 )
4
PixArt Sigma ( Chen et al., 2024c )
Seedream 4.5 ( Seedream, 2025 )
5
SD 3.5 Medium ( AI, 2024b )
Ideogram 3.0 Turbo ( Ideogram, 2025 )
6
Kandinsky 2 ( Razzhigaev et al., 2023 )
Imagen 4 Fast ( DeepMind, 2025 )
Appendix
Table 7. Generators used for TIQA datasets
Figure 8. TIQA-Crops examples
Figure 9. MOS distributions (1–5) for three target strings on the OCR-correct subset (exact transcript match). Similar distributions across a real word, an anagram, and a nonword indicate limited sensitivity of visual-quality ratings to lexical plausibility when rendering is correct.
Figure 10. Distibution plot for OQ-MOS and TQ-MOS for TIQA-Images dataset. Vertical dashed lines denote mean values across all images.
Figure 11. Representative visual examples of the rating categories shown to annotators during the TIQA-Crops labeling task. Each row corresponds to a different score on the 0 – 5 scale and serves as a reference anchor for the rater.
Figure 21Figure 22Figure 23
Figure 12. TIQA-Images examples for overall quality TODO
Figure 13. TIQA-Images examples for text quality TODO