UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
Authors: Deyuan Liu, Yihao Hu, Jingxuan Zhang, Xingying Li, Jun Xie, Jiacheng Liu, Jungang Li, Yu Huang, +8 more
Organizations: Westlake University · Ant Group · Zhejiang University · Shanghai Innovation Institute · MBZUAI · HKUST · CityU · Peking University · CASIA · Wechat AI
Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: https://github.com/LINs-lab/UltraText_Bench.
Figures & tables
Figure 1: Scene coverage at L3. One selected output per category in taxonomy order, from GPT Image 2 at API quality Low . L3 denotes Extreme workload; EN/ZH denote English/Chinese.
Figure 2: Local success can hide whole-scene failures. This illustrative six-region scene contrasts a correct shop name with errors elsewhere. Left. The correctly rendered shop name, cropped. Center. The full scene, with five further requested regions marked. Each is missing, wrong, unreadable, misplaced, or legible but poorly integrated with its intended surface. Right. Dense text load and whole-scene evaluation cover the requested text.
Benchmark
Language
Extra input
Target reference
Evaluation
MARIO-Eval
EN
None
Target strings
OCR + image metrics
AnyText-Bench
EN/ZH
Glyph/position
Text + positions
OCR on text crops
LeX-Bench
EN
None
Text + style/position
PNED + attribute VQA
CVTG-2K
EN
None
Text instances
OCR + matching
LongText-Bench
EN/ZH
None
Target strings
VLM text accuracy
TextInVision
EN
None
Target strings
OCR edit distance
Table 1: Text-rendering benchmarks: inputs, target references, and evaluation. The table summarizes the main evaluation components; details appear in Section 2.2 . Notes: EN: English; ZH: Chinese. Extra input denotes benchmark-supplied controls beyond the prompt, not internally predicted layouts. OCR: optical character recognition; VQA: visual question answering. PNED matches words by edit distance and penalizes unmatched items. † No image is supplied for T2I; source images are used for editing or image-to-image tasks. UltraText pairs per-region text and visual attributes with image-level VLM ratings.
Figure 3: Task scenes and assessment in three benchmarks. Left: CVTG-2K ( Du et al., 2025 ) matches OCR-recognized words to targets using Word Accuracy and normalized edit distance, and separately reports CLIPScore. Center: LongText-Bench ( Geng et al., 2025 ) uses a VLM to extract text for comparison with target strings. Right: UltraText Bench uses a structured reference (Ref.) to rate the whole image: text accuracy (TA), completeness (TC), readability (TR), position correctness (PC), layout quality (LQ), and scene integration (SI). The scenes, abbreviated strings, and line marks are illustrations, not benchmark outputs or measured results; they do not exhaust each benchmark’s scene types.
Figure 4: Task definition and evaluation protocol. The generator receives only prompt p . Q-Judger, from Qwen-Image-Bench ( Li et al., 2026a ) , applies the UltraText rubric to the image and complete reference R , returning six image-level scores that map to four reporting dimensions and a composite. Reference fields and scores are not paired one-to-one. Grid cells give coarse locations; regions may share a cell. Invalid responses remain unscored and reduce coverage. The bakery schematic simplifies the four-region record A1_L1_EN_001 , reproduced in Appendix B.1 .
ID
Category
Rendering challenge
Density
Typical layout
A. Signage & Labels
A1
Sign
Perspective and weathered surfaces
Low–Med
Scattered blocks
A2
Label
Small text on curved surfaces
Med
Compact blocks
A3
Poster
Font hierarchy and composition
Med–High
Stacked blocks
A4
Billboard
Scale and viewing angles
Low–Med
Large text blocks
B. Documents & Print
Table 2: UltraText Bench scene taxonomy. Six domains organize 24 categories by rendering challenge and typical layout. Notes: Density describes qualitative text packing, not the GT-character load in Table 3 . Layouts are typical examples, not mandatory templates.
Figure 5: Realized difficulty statistics for the 432 prompts. Each row contains 72 prompts. Bars show minimum–maximum ranges and dots show means from Table 3 . (a) Target-string character counts, including spaces and punctuation, on a log axis. Grey brackets show the language-specific design bands; bounds are inclusive, ZH bands overlap, and L3 bands have no upper bound (shown by open ends). Levels are assigned workloads, not disjoint character bins; all records satisfy their own band. (b) Region counts; no per-level region-count bands are specified. Both panels describe benchmark inputs, not model performance.
Split
Prompts
GT characters per prompt
Regions per prompt
Mean
Range
Mean
Range
EN L1
72
489.90
278–592
4.67
4–7
EN L2
72
1,023.44
674–1,195
7.10
6–10
EN L3
72
2,386.68
1,369–5,230
9.38
8–12
ZH L1
72
394.62
350–555
4.93
4–6
ZH L2
72
627.25
551–892
6.17
6–8
Table 3: Dataset statistics by language and difficulty. Notes: Means and ranges are per prompt. GT characters include spaces and punctuation and are summed over all requested regions. The 432 prompts contain 407,918 GT characters and 2,926 regions in total; all statistics are recomputed from the released region arrays.
Dimension
Formula
Assessment focus
Weight
Text fidelity
(TA+TC)/2
Text correctness and completeness
60%
Text clarity
TR
Sharpness, contrast, and legibility
30%
Spatial quality
(PC+LQ)/2
Position, alignment, and organization
5%
Scene quality
SI
Natural fit with the text carrier and scene
5%
Table 4: From six raw scores to four reporting dimensions. Scores are VLM ratings on a 0–100 scale, with higher values better. TA/TC/TR denote text accuracy/completeness/readability; PC/LQ/SI denote position correctness/layout quality/scene integration. Weights apply to reporting dimensions, abbreviated Fidelity, Clarity, Spatial, and Scene in the figures and leaderboard.
Model
Composite
Fidelity
Clarity
Spatial
Scene
L1
L2
L3
EN
ZH
EN
ZH
EN
ZH
Open-weight models
Boogu-Image-0.1-Base
77.89
75.77
78.31
88.04
90.53
84.16
87.75
76.42
78.11
57.67
83.23
Qwen-Image-2512
67.96
59.30
79.89
79.13
88.93
86.50
89.47
71.57
67.44
42.86
49.90
Boogu-Image-0.1-Turbo
67.79
66.27
64.86
84.37
86.90
73.13
81.30
62.81
72.61
37.24
79.63
Z-Image-Base
61.92
55.51
70.60
70.88
77.72
82.92
84.66
67.54
64.41
26.34
45.64
Table 5: UltraText Bench leaderboard. VLM ratings range from 0 to 100 (higher is better). Overall columns average EN/ZH prompt-macro means equally; level columns show language-specific composites. Composite weights are 60/30/5/5 for Fidelity/Clarity/Spatial/Scene ( Table 4 ). L1/L2/L3 denote Hard/Very Hard/Extreme; bracketed Low/High labels are API quality settings. Bold marks column bests within each group; shading marks the Composite leader. Execution failures affect coverage, not quality means ( Section 5.1 ).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: L1/L2/L3 newspaper examples. All outputs use GPT Image 2 [Low], sample 1. Notes: The reference grows from a masthead, a headline, a lede, and a footer into several articles surrounded by smaller supporting text. Character load and region count describe the input. These thumbnails show the page layouts; checking individual characters requires enlargement.
Figure A2: Chinese L3 newspaper: whole image and native-pixel crop. B2_L3_ZH_001 , GPT Image 2 [Low], sample 1; 627 characters, 10 regions. Notes: The outline marks where the crop was taken and is not a GT bounding box. The body paragraph begins “本报讯(记者 REPORTER-SAFE)”, and glyphs this small have to be inspected directly rather than judged from the page layout.
Stage
Action
Output / validation
1. Generate
Generate four samples per prompt under the model’s default sampling configuration.
One PNG per sample with its model, prompt, and sample identifiers.
2. Validate
Check PNG structure and the prompt, GT, image, and model bindings.
Invalid images are held back from the judge with an explicit status.
3. Strict Judge
Send the complete GT together with the image, and require exactly six integer score keys in return.
Overflow, refusal, API, and parse failures become judge_failed rows with null scores.
4. Report
Read .summary.json ; report planned and successful image counts, failure causes, image and prompt coverage, and language, level, and category aggregates.
Averages run over successful images; a prompt-macro mean averages within each prompt first. No implicit zero, 50, or completion threshold is applied.
Appendix
Table A1: Evaluation workflow. Notes: All planned rows stay in the coverage denominators, while only strict judge successes enter the quality means. Runtime depends on hardware, generator, and service throughput. Command examples follow below.
Text-side diagnostics
Dimension
Severe 0–19
Weak 20–39
Partial 40–59
Strong 60–79
Excellent 80–100
TA Accuracy
Little or no correct text
Isolated correct fragments
Mixed correct and wrong text
Minor character errors
Near-exact transcription
TC Completeness
Little required text
Many missing regions
Substantial omissions
Minor omissions
Near-complete coverage
TR Readability
Illegible strokes
Few readable fragments
Mixed legibility
Mostly clear
Consistently legible
Appendix
Table A2: Interpretive score bands. Notes: These qualitative descriptions are a reading aid rather than calibrated thresholds on character accuracy or region recall, and they form no part of the judge prompt.
Figure A3: Character errors on a plausible sign. A1_L3_EN_002 , sample 1; 1,546 GT characters, 10 regions. GT excerpt ( region_5 ): “GREENFIELD COMMUNITY NOTICE”. Observation: The heading of the notice is built from malformed letters, even though the poster and the street around it stay entirely recognizable. The highlighted error occurs in that heading. Dimensions: TA, with TR also relevant.
Figure A4: Missing title amid detailed packaging text. C4_L1_EN_002 , sample 2; 563 GT characters, 5 regions. GT excerpt ( region_0 ): “BOTANICA — Repair & Restore Shampoo”. Observation: The bottle opens straight into product claims, and the prominent product title the reference asks for is absent. Neither the surrounding text nor the detailed ingredient list can stand in for that region. Dimension: TC.
Figure A5: Dense Chinese glyphs require local inspection. C1_L3_ZH_002 , sample 4; 611 GT characters, 8 regions. GT excerpt ( region_3 ): “鸡腿葱串 两串炭火现烤 ¥32”. Observation: Seen at page scale the dish-and-price rows look well organized, but the glyphs are small enough that their strokes and exact wording become checkable only under enlargement. Tidy rows by themselves say nothing about text fidelity. Dimensions: TA and TR.
Figure A6: Correct surface type at the wrong grid location. A1_L2_ZH_001 , sample 1; 564 GT characters, 7 regions. GT ( region_4 ): a small door plate at bottom-right , beginning “周一至周六 6:30-19:00”. Observation: The hours plate turns up in the bottom-left of the full image instead. The enlargement is what identifies the region, but its location can only be judged from the whole image. Dimension: PC.
Figure A7: Readable chat content and a discrepant judge score. F2_L3_ZH_002 , sample 4; 631 GT characters, 8 regions. GT excerpt ( region_2 ): “成员D:我DATE-SAFE从LOCBLOCK-SAFE出发”. Observation: Related text is plainly visible down the central chat column, with further messages and stickers filling the flanks, and yet the judge assigns TA=0 and TC=0 while giving LQ=100. Manual inspection identifies a disagreement between the visible content and the automated fidelity scores. The original scores are retained to make this failure case inspectable; they should not be interpreted as a human assessment that all requested text is absent. Dimensions: TC and LQ.
Figure A8: Carrier integration as a separate property. A1_L1_EN_001 , sample 1; 423 GT characters, 4 regions. GT ( region_0 ): the “DAILY BREAD” title belongs on a wooden shop sign. Observation: The lettering follows the plane of the sign and picks up its lighting and surface texture. Integration is judged separately from whether the text itself is exact. Dimension: SI. The score calculation is given in Appendix D.1 .
Figure A9: Maximum ratings despite degraded small text. E4_L3_EN_003 , sample 2; 5,230 GT characters, 8 regions; original image 1024×1024 pixels. GT excerpt ( region_7 ): “func TestParsePage(t *testing.T)” followed by the test cases and assertions. Observation: The bottom-right test block contains visibly deformed and difficult-to-read characters, despite maximum accuracy and readability ratings. The image hash matches its saved judge record. This selected case illustrates a local limitation of the automatic scores; it does not estimate how often maximum ratings overlook errors. Dimensions: TA and TR.
Figure A10: Scene coverage at L1 (Hard). One selected output per category in taxonomy order, from GPT Image 2 at API quality Low . EN/ZH denote English/Chinese. These examples illustrate scene diversity.
Figure A11: Scene coverage at L2 (Very Hard). One selected output per category in taxonomy order, from GPT Image 2 at API quality Low . EN/ZH denote English/Chinese. These examples illustrate scene diversity.
Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at https://github.com/hardenyu21/VTR-Bench.
Yu Huang, Jungang Li, Zhiyuan Wang +8
Department of Data Science, City University of Hong Kong · The Hong Kong University of Science and Technology (Guangzhou) · Westlake University +3
Recent text-to-image generation models have demonstrated remarkable capabilities in synthesizing highly realistic images from text inputs alone. Although existing benchmarks can evaluate the generation capabilities of various models to some extent, they struggle to comprehensively and accurately measure performance across multiple dimensions, often failing to reveal the inherent deficiencies of models in specific categories. To address these limitations, we propose WeGenBench, a novel benchmark designed for the comprehensive, multi-perspective evaluation of text-to-image generation capabilities. Our benchmark comprises a total of 4,000 test prompts across two primary categories, meticulously balanced between Chinese and English to evaluate bilingual and cross-cultural generation capabilities. Beyond macroscopic scene classification, we annotate each prompt with multi-dimensional tags tailored to the distinct content and challenges of each language, thereby refining the generation tasks into more specific sub-categories. Through a cross-dimensional evaluation mechanism leveraging both scene classifications and multi-dimensional tags, WeGenBench can precisely pinpoint model shortcomings in specific generation categories. Furthermore, to measure generation quality more accurately, we design and validate several novel evaluation metrics by integrating Vision-Language Models (VLMs), which assess model performance on domain-specific tasks from three core aspects. Crucially, our approach yields both the assessment outcomes and the detailed reasoning trajectories, facilitating a rigorous verification of the accuracy and soundness of the evaluation results. Finally, we conduct systematic benchmarking on current state-of-the-art methods and provide an in-depth analysis of the limitations present in existing models.
Qian Liang, Xiaomin Li, Ying Zhang +6
University of Electronic Science and Technology of China · Dalian University of Technology · Weixin, Tencent
Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no longer satisfy users' pressing demands for faithful real-world reconstruction and genuine creative expression. Existing benchmarks, however, remain anchored in these foundational criteria and do not yet capture the nuanced capabilities that matter in authentic artistic practice, making it difficult to reliably distinguish state-of-the-art T2I models. To address the gap, we introduce Qwen-Image-Bench, a creator-centric benchmark co-designed with professional artists and grounded in real-world creation scenarios. Qwen-Image-Bench enriches conventional evaluation with two application-driven dimensions: Real-world Fidelity and Creative Generation. Drawing on the staged reasoning inherent in professional artistic workflows, we organize these five pillars into a top-down hierarchical taxonomy that further decomposes into 23 second-level sub-capabilities and 56 third-level verifiable rubrics. To ensure broad coverage, we curate 1000 stratified prompts with each prompt jointly exercising more than four fine-grained facets across multiple pillars. We train a unified judge model Q-Judger based on Qwen3.6-27B, supervised by 80 professional annotators from global art academies under blind labeling and triple-review protocols, that scores every image across all 56 verifiable facets, producing fine-grained, rubric-grounded, and fully attributable diagnostics rather than a single opaque score. Empirically, Qwen-Image-Bench reliably distinguishes leading T2I models, achieving the greatest separation on the two application-driven dimensions of Real-world Fidelity and Creative Generation where existing benchmarks provide little insight, while also providing a trustworthy optimization signal for production-level T2I development.