VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
Authors: Yu Huang, Jungang Li, Zhiyuan Wang, Yonghua Hei, Song Dai, Jiayu Yang, Deyuan Liu, Xiang Zheng, +3 more
Organizations: Department of Data Science, City University of Hong Kong · The Hong Kong University of Science and Technology (Guangzhou) · Westlake University · The Hong Kong Institute of AI for Science, City University of Hong Kong · University of Electronic Science and Technology of China · The Hong Kong University of Science and Technology
Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at https://github.com/hardenyu21/VTR-Bench.
Figures & tables
Figure 1: Visual text rendering remains challenging for current video generation models, even in visually convincing videos. Top: a video generated by Seedance2.5. Video Score measures the proportion of satisfied requirements in a 20-question checklist, and WER measures word error rate against the reference text. Bottom left: four text rendering issues in the video. Bottom right: overall WER of 11 video generation models on VTR-Bench.
Figure 2: Overview of VTR-Bench. Left: five expert-designed scenario categories guide scene seed generation and human filtering. The retained seeds support prompt construction, followed by iterative refinement through reference image generation and VLM review and a final human review. Right: generated videos are evaluated along two separate dimensions. Video Score measures adherence to scene and video requirements using a 20-question checklist, while WER measures visual text fidelity by comparing carrier-specific VLM transcriptions with reference text.
Figure 3: Dataset statistics of VTR-Bench.
Figure 4: Keyframe-guided agentic generation. A Director agent uses visual feedback to refine image prompts, edit or regenerate first frames, and revise motion plans. The selected first frame, original prompt, and motion plan condition video generation.
Model
Resolution
WER ( ↓ )
Video Score ( ↑ )
AD
Sci
UI
Cul
DL
All
AD
Sci
UI
Cul
DL
All
Open-Sourced Models
Wan2.2-5B
1280
×
704
0.994
0.993
0.999
0.998
0.994
0.996
0.523
0.475
0.386
0.357
0.438
0.436
HunyuanVideo-1.5
1280
×
720
0.972
0.952
0.974
0.985
0.964
0.969
0.606
0.628
0.653
0.557
0.562
0.601
LTX-2.3
768
×
512
0.997
0.992
0.998
0.998
0.993
0.996
0.682
0.632
0.569
0.560
0.590
0.607
Lingbot-Video-Dense
832
×
480
0.992
0.974
0.993
0.996
0.988
0.988
0.573
0.571
0.556
0.469
0.503
0.534
Table 1: Main results on VTextBench. The top three results in each metric are highlighted in blue, with darker shades indicating better performance.
Figure 5: Effects of generation resolution and duration. Left: Video Score ( ↑ ). Right: WER ( ↓ ).
Figure 6: Transcription outcomes reported by VLM evaluator. Bars show the proportion of text targets in each state.
Figure 7: Comparison of direct generation, I2V, and the agentic framework on Minimax H3. Bold values indicate the best result within each scenario or overall.
Evaluation
Agreement ↑
Pearson ↑
Spearman ↑
Lin’s CCC ↑
MAE ↓
CoQ judgments
92.15%
—
—
—
—
Text-block WER
—
0.9541
0.9370
0.9536
0.0545
Video-level WER
—
0.9842
0.9731
0.9835
0.0407
Table 2: Alignment of Qwen3.8-27B.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
WER ( ↓ )
Video Score ( ↑ )
WER Rank ( ↓ )
Video Rank ( ↓ )
AD
Sci
UI
Cul
DL
All
AD
Sci
UI
Cul
DL
All
Qwen3.6
Qwen3.8
Qwen3.6
Qwen3.8
Open-Sourced Models
Wan2.2-5B
0.996
0.990
0.997
0.997
0.986
0.993
0.563
0.577
0.477
0.421
0.495
0.507
11
11
11
11
HunyuanVideo-1.5
0.966
0.944
0.966
0.982
0.964
0.964
0.668
0.718
0.708
0.618
0.602
0.663
7
7
8
8
LTX-2.3
0.993
0.987
0.996
0.997
0.986
0.992
0.728
0.733
0.696
0.639
0.637
0.687
10
10
7
7
Lingbot-Video-Dense
0.991
0.964
0.986
0.995
0.981
0.983
0.614
0.655
0.647
0.540
0.548
0.601
9
9
10
10
Appendix
Table 3: Results on VTR-Bench evaluated by Qwen3.6-27B. Darker blue highlights indicate better results among the top three models in each score column. Ranks use overall WER and Video Score before rounding.
Metric
Qwen3.6-27B
Qwen3.8-27B
CoQ judgments: 2,000 queries
Agreement ↑
91.60%
92.15%
Text-block WER: 404 blocks
Pearson ↑
0.9536
0.9541
Spearman ↑
0.9365
0.9370
Lin’s CCC ↑
0.9527
0.9536
Appendix
Table 4: Human alignments of the two evaluators. Better results are shown in bold.
Model
WER ( ↓ )
α=0
α=1
α=2
α=3
α=4
α=5
α=6
α=7
α=8
α=9
α=10
Qwen3.8-27B
Wan2.2-5B
0.993
0.994
0.995
0.995
0.996
0.996
0.996
0.996
0.996
0.996
0.996
HunyuanVideo-1.5
0.955
0.959
0.963
0.966
0.968
0.969
0.971
0.972
0.973
0.973
0.974
LTX-2.3
0.991
0.993
0.994
0.995
0.995
0.996
0.996
0.996
0.996
0.996
0.996
Lingbot-Video-Dense
0.971
0.978
0.982
0.986
0.987
0.988
0.989
0.990
0.990
0.991
0.991
Appendix
Table 5: Overall WER for α=0 – 10 under both evaluators. The shaded column marks the default α=5 ; bold values indicate the lowest WER in each column within each evaluator.
Model
Score ( ↑ )
Scene Attributes
Motion Adherence
Spatial Relationship
Entity Presence
Temporal Consistency
All
Open-Sourced Models
Wan2.2-5B
0.6308
0.1326
0.5277
0.4813
0.0974
0.4358
HunyuanVideo-1.5
0.7427
0.3431
0.6519
0.7511
0.3467
0.6013
LTX-2.3
0.7279
0.3873
0.6667
0.7076
0.3954
0.6068
Lingbot-Video-Dense
0.7068
0.2578
0.6039
0.5487
0.3496
0.5343
Appendix
Table 6: Scores across five video evaluation dimensions, evaluated by Qwen3.8-27B. Bold values indicate the best result in each column.
Figure 8: Word frequencies in transcription diagnostics pooled from Qwen3.8-27B and Qwen3.6-27B. (a) 6,433 descriptions for unreadable targets. (b) 6,549 descriptions accompanying transcriptions. Counts show the 15 most frequent terms within each selected diagnostic vocabulary.
Figure 9: Garbled text in ViduQ3. The problem heading and answer card contain pseudo-words despite their clear appearance.
Figure 12: Malformed Chinese characters in Lingbot-Video-Dense. Printed text blocks contain abnormal stroke combinations.
Figure 15: Incorrect parameter in Wan3.0. The required β=8/3 is rendered as β=4/3 .
A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is unforgiving in video generation: minor stroke corruption, temporal instability, or editing errors instantly break legibility and realism. Existing benchmarks overlook this challenge by treating text as incidental or using static OCR metrics that ignore temporal dynamics. We introduce VidScribe, a unified diagnostic benchmark spanning four generation regimes: writing from language (T2V), transferring text identity from a reference (R2V), sustaining text under dynamics (I2V), and localized text editing (V2V). VidScribe contains 803 human-verified samples across a 12-axis conditionally orthogonal factor space covering Intrinsic Text Properties, Physical Imaging Conditions, and Temporal Behavior. For reliable evaluation, we build a track-grounded, gated suite with 11 shared metrics and 2 task-specific probes under strict measurability conditions. Benchmarking 11 commercial and open-source systems shows that video text capability is non-monolithic, with content recognition decoupled from stroke-level glyph correctness. Performance is highly task-asymmetric: I2V sustains text most reliably, whereas V2V editing is the primary bottleneck. Counter-intuitively, degradation concentrates on a small subset of text-centric structural and temporal factors rather than adverse imaging conditions. Further probes show that visual references improve glyph and typographic fidelity rather than content accuracy, while localized editing fails to isolate target text without corrupting undeclared source text. Beyond evaluation, VidScribe also provides an actionable training signal, where benchmark-aligned preference optimization measurably improves visual text generation. https://huggingface.co/datasets/Vicky0720/VidScribe.
Ziying Zhang, Litao Li, Junchao Liao +4
Alibaba Group · Shanghai Jiao Tong University · Fudan University
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.
Shengyi Wang, Niantong Li, Guangzheng Hu +27
Alibaba Group · Moku Lab, Hujing Digital Media & Entertainment Group · Beijing Film Academy
Video-to-video (V2V) generation is difficult to evaluate because outputs must both follow editing instructions and preserve frame-level correspondence with the source video, which existing T2V and I2V metrics do not capture. We introduce V2V-Bench, a 11-dimension benchmark organized into five categories: temporal alignment, structural fidelity, transformation quality, video quality, and semantic alignment. V2V-Bench pairs diverse source videos with challenging editing tasks and evaluates two commercial models, Grok Imagine and Gemini Veo3, and one open-source model, Open Sora 2. Results show complementary model strengths: Grok performs better on editing fidelity, while Veo3 achieves stronger visual quality. On six V2V-specific dimensions, V2V-Bench reaches a Spearman correlation of 0.905 with human judgments.