cs.CVOct 1, 2026

VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

Authors: Yu Huang, Jungang Li, Zhiyuan Wang, Yonghua Hei, Song Dai, Jiayu Yang, Deyuan Liu, Xiang Zheng, +3 more

Organizations: Department of Data Science, City University of Hong Kong · The Hong Kong University of Science and Technology (Guangzhou) · Westlake University · The Hong Kong Institute of AI for Science, City University of Hong Kong · University of Electronic Science and Technology of China · The Hong Kong University of Science and Technology

Abstract

Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at https://github.com/hardenyu21/VTR-Bench.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation

    Sep 29, 2026Ziying Zhang, Litao Li, Junchao Liao +4Text-To-Video Generation ModelFew-Shot Font Generation

  2. FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

    Jul 27, 2026Shengyi Wang, Niantong Li, Guangzheng Hu +27Cinematic IdealLong-Video Benchmarks

  3. V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation

    Jun 4, 2026Tao Liu, Leela Krishna, Gouti Pavan Kumar +2Video QualityLong Videos