VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
Authors: Yu Huang, Jungang Li, Zhiyuan Wang, Yonghua Hei, Song Dai, Jiayu Yang, Deyuan Liu, Xiang Zheng, +3 more
Organizations: Department of Data Science, City University of Hong Kong · The Hong Kong University of Science and Technology (Guangzhou) · Westlake University · The Hong Kong Institute of AI for Science, City University of Hong Kong · University of Electronic Science and Technology of China · The Hong Kong University of Science and Technology
Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at https://github.com/hardenyu21/VTR-Bench.
Figures & tables
Figure 1: Visual text rendering remains challenging for current video generation models, even in visually convincing videos. Top: a video generated by Seedance2.5. Video Score measures the proportion of satisfied requirements in a 20-question checklist, and WER measures word error rate against the reference text. Bottom left: four text rendering issues in the video. Bottom right: overall WER of 11 video generation models on VTR-Bench.
Figure 2: Overview of VTR-Bench. Left: five expert-designed scenario categories guide scene seed generation and human filtering. The retained seeds support prompt construction, followed by iterative refinement through reference image generation and VLM review and a final human review. Right: generated videos are evaluated along two separate dimensions. Video Score measures adherence to scene and video requirements using a 20-question checklist, while WER measures visual text fidelity by comparing carrier-specific VLM transcriptions with reference text.
Figure 3: Dataset statistics of VTR-Bench.
Figure 4: Keyframe-guided agentic generation. A Director agent uses visual feedback to refine image prompts, edit or regenerate first frames, and revise motion plans. The selected first frame, original prompt, and motion plan condition video generation.
Model
Resolution
WER ( ↓ )
Video Score ( ↑ )
AD
Sci
UI
Cul
DL
All
AD
Sci
UI
Cul
DL
All
Open-Sourced Models
Wan2.2-5B
1280
×
704
0.994
0.993
0.999
0.998
0.994
0.996
0.523
0.475
0.386
0.357
0.438
0.436
HunyuanVideo-1.5
1280
×
720
0.972
0.952
0.974
0.985
0.964
0.969
0.606
0.628
0.653
0.557
0.562
0.601
LTX-2.3
768
×
512
0.997
0.992
0.998
0.998
0.993
0.996
0.682
0.632
0.569
0.560
0.590
0.607
Lingbot-Video-Dense
832
×
480
0.992
0.974
0.993
0.996
0.988
0.988
0.573
0.571
0.556
0.469
0.503
0.534
Table 1: Main results on VTextBench. The top three results in each metric are highlighted in blue, with darker shades indicating better performance.
Figure 5: Effects of generation resolution and duration. Left: Video Score ( ↑ ). Right: WER ( ↓ ).
Figure 6: Transcription outcomes reported by VLM evaluator. Bars show the proportion of text targets in each state.
Figure 7: Comparison of direct generation, I2V, and the agentic framework on Minimax H3. Bold values indicate the best result within each scenario or overall.
Evaluation
Agreement ↑
Pearson ↑
Spearman ↑
Lin’s CCC ↑
MAE ↓
CoQ judgments
92.15%
—
—
—
—
Text-block WER
—
0.9541
0.9370
0.9536
0.0545
Video-level WER
—
0.9842
0.9731
0.9835
0.0407
Table 2: Alignment of Qwen3.8-27B.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
WER ( ↓ )
Video Score ( ↑ )
WER Rank ( ↓ )
Video Rank ( ↓ )
AD
Sci
UI
Cul
DL
All
AD
Sci
UI
Cul
DL
All
Qwen3.6
Qwen3.8
Qwen3.6
Qwen3.8
Open-Sourced Models
Wan2.2-5B
0.996
0.990
0.997
0.997
0.986
0.993
0.563
0.577
0.477
0.421
0.495
0.507
11
11
11
11
HunyuanVideo-1.5
0.966
0.944
0.966
0.982
0.964
0.964
0.668
0.718
0.708
0.618
0.602
0.663
7
7
8
8
LTX-2.3
0.993
0.987
0.996
0.997
0.986
0.992
0.728
0.733
0.696
0.639
0.637
0.687
10
10
7
7
Lingbot-Video-Dense
0.991
0.964
0.986
0.995
0.981
0.983
0.614
0.655
0.647
0.540
0.548
0.601
9
9
10
10
Appendix
Table 3: Results on VTR-Bench evaluated by Qwen3.6-27B. Darker blue highlights indicate better results among the top three models in each score column. Ranks use overall WER and Video Score before rounding.
Metric
Qwen3.6-27B
Qwen3.8-27B
CoQ judgments: 2,000 queries
Agreement ↑
91.60%
92.15%
Text-block WER: 404 blocks
Pearson ↑
0.9536
0.9541
Spearman ↑
0.9365
0.9370
Lin’s CCC ↑
0.9527
0.9536
Appendix
Table 4: Human alignments of the two evaluators. Better results are shown in bold.
Model
WER ( ↓ )
α=0
α=1
α=2
α=3
α=4
α=5
α=6
α=7
α=8
α=9
α=10
Qwen3.8-27B
Wan2.2-5B
0.993
0.994
0.995
0.995
0.996
0.996
0.996
0.996
0.996
0.996
0.996
HunyuanVideo-1.5
0.955
0.959
0.963
0.966
0.968
0.969
0.971
0.972
0.973
0.973
0.974
LTX-2.3
0.991
0.993
0.994
0.995
0.995
0.996
0.996
0.996
0.996
0.996
0.996
Lingbot-Video-Dense
0.971
0.978
0.982
0.986
0.987
0.988
0.989
0.990
0.990
0.991
0.991
Appendix
Table 5: Overall WER for α=0 – 10 under both evaluators. The shaded column marks the default α=5 ; bold values indicate the lowest WER in each column within each evaluator.
Model
Score ( ↑ )
Scene Attributes
Motion Adherence
Spatial Relationship
Entity Presence
Temporal Consistency
All
Open-Sourced Models
Wan2.2-5B
0.6308
0.1326
0.5277
0.4813
0.0974
0.4358
HunyuanVideo-1.5
0.7427
0.3431
0.6519
0.7511
0.3467
0.6013
LTX-2.3
0.7279
0.3873
0.6667
0.7076
0.3954
0.6068
Lingbot-Video-Dense
0.7068
0.2578
0.6039
0.5487
0.3496
0.5343
Appendix
Table 6: Scores across five video evaluation dimensions, evaluated by Qwen3.8-27B. Bold values indicate the best result in each column.
Figure 8: Word frequencies in transcription diagnostics pooled from Qwen3.8-27B and Qwen3.6-27B. (a) 6,433 descriptions for unreadable targets. (b) 6,549 descriptions accompanying transcriptions. Counts show the 15 most frequent terms within each selected diagnostic vocabulary.
Figure 9: Garbled text in ViduQ3. The problem heading and answer card contain pseudo-words despite their clear appearance.
Figure 12: Malformed Chinese characters in Lingbot-Video-Dense. Printed text blocks contain abnormal stroke combinations.
Figure 15: Incorrect parameter in Wan3.0. The required β=8/3 is rendered as β=4/3 .