ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
Organizations: Texas A&M University
Abstract
Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits that must preserve the original scene dynamics. Video scene text editing replaces text on scene surfaces, such as storefront signs, whiteboards, and product labels, while preserving the surrounding content, motion, and camera dynamics. Although scene text editing is well studied for images, video scene text editing that achieves high visual quality, temporal consistency, and edit locality remains underexplored. Existing resources offer limited paired real-video data, and general video-editing metrics do not directly measure whether the requested text remains correct over time. We introduce ViTeX-Bench, a benchmark suite comprising ViTeX-Dataset and a three-axis evaluation protocol. The dataset contains 387 real-world 720p videos with text-region masks and editing instructions: 230 provide reviewed, pipeline-generated paired edits for training, and 157 form a frozen evaluation split. The protocol evaluates text correctness, visual and temporal quality, and edit locality through 13 metrics, with one primary metric per axis and a Pareto comparison of their trade-offs. OCR calibration, human evaluation, and annotation-sensitivity analyses support the interpretation of these scores. Across eight baselines from four editing families, accurate text, temporal stability, and scene preservation remain difficult to achieve together. We also release ViTeX-Edit-14B, an open-source reference editor fine-tuned on the paired training split with motion-aligned glyph-video conditioning. It achieves CharAcc 0.688, the highest mean among the evaluated video-native editors, and the lowest comparable text-crop Warp among raw editor outputs. ViTeX-Bench provides a reproducible foundation for studying these trade-offs in video scene text editing.
Figures & tables
| Videos | String length | Font style (%) | |||||
| Train | Eval | Scripts | Source | Target | Printed | Handwritten | Artistic |
| 230 | 157 | 4 | 23 | 44 | 33 | ||
| Text correctness | Visual quality | Edit locality | ||||||||||||
| Method | Fam. | SeqAcc | CharAcc | TTS | Flicker | Flicker | Warp | Warp | MUSIQ | MUSIQ | PSNR | SSIM | LPIPS | DreamSim |
| Source video | — | 0.000 | 0.317 | 0.760 | 3.72 | 3.68 | 1.46 | 1.27 | 70.33 | 45.12 | 1.000 | 0.000 | 0.000 | |
| AnyText2 [ 9 ] | A | 0.280 | 0.633 | 0.382 | 3.34 | 4.95 | 2.04 | 3.95 | 66.68 | 41.65 | 25.56 | 0.905 | 0.091 | 0.043 |
| TextCtrl [ 10 ] | A | 0.475 | 0.734 | 0.511 | 3.80 | 4.29 | 1.59 | 2.09 | 70.32 | 42.77 | 41.14 | 0.994 | 0.008 | 0.004 |
| FLUX-Text [ 11 ] | A | 0.528 | 0.737 | 0.326 | 5.11 | 14.81 | 3.03 | 13.01 | 70.26 | 43.85 | 31.49 | 0.975 | 0.029 | 0.012 |
| RS-STE [ 12 ] | A | 0.354 | 0.626 | 0.534 | 3.73 | 3.66 | 1.61 | 1.81 | 69.57 | 34.26 | 37.00 | 0.983 | 0.024 | 0.007 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | SeqAcc | CharAcc | TTS |
|---|---|---|---|
| Source video | 0.000 [0.000, 0.000] | 0.317 [0.285, 0.350] | 0.760 [0.721, 0.800] |
| AnyText2 | 0.280 [0.229, 0.333] | 0.633 [0.591, 0.676] | 0.382 [0.336, 0.427] |
| TextCtrl | 0.475 [0.409, 0.539] | 0.734 [0.684, 0.780] | 0.511 [0.459, 0.559] |
| FLUX-Text | 0.528 [0.483, 0.578] | 0.737 [0.696, 0.778] | 0.326 [0.286, 0.369] |
| RS-STE | 0.354 [0.288, 0.412] | 0.626 [0.574, 0.677] | 0.534 [0.484, 0.588] |
| TextCtrl + AnyV2V | 0.057 [0.031, 0.088] | 0.308 [0.274, 0.346] | 0.257 [0.222, 0.297] |
| Method | Flicker | Flicker | Warp | Warp | MUSIQ | MUSIQ |
|---|---|---|---|---|---|---|
| Source video | 3.72 [3.14, 4.41] | 3.68 [2.91, 4.63] | 1.46 [1.33, 1.62] | 1.27 [1.09, 1.46] | 70.33 [69.55, 71.14] | 45.12 [43.23, 47.09] |
| AnyText2 | 3.34 [2.89, 3.88] | 4.95 [4.36, 5.64] | 2.04 [1.85, 2.28] | 3.95 [3.56, 4.41] | 66.68 [65.82, 67.57] | 41.65 [39.96, 43.45] |
| TextCtrl | 3.80 [3.22, 4.49] | 4.29 [3.53, 5.21] | 1.59 [1.44, 1.76] | 2.09 [1.86, 2.35] | 70.32 [69.52, 71.14] | 42.77 [40.96, 44.74] |
| FLUX-Text | 5.11 [4.50, 5.80] | 14.81 [13.56, 16.01] | 3.03 [2.80, 3.29] | 13.01 [11.71, 14.21] | 70.26 [69.45, 71.07] | 43.85 [42.09, 45.74] |
| RS-STE | 3.73 [3.15, 4.40] | 3.66 [2.97, 4.53] | 1.61 [1.46, 1.77] | 1.81 [1.60, 2.06] | 69.57 [68.74, 70.43] | 34.26 [32.67, 36.01] |
| TextCtrl + AnyV2V | 4.98 [4.42, 5.67] | 4.98 [4.17, 5.89] | 4.11 [3.68, 4.59] | 3.97 [3.46, 4.56] | 69.41 [68.41, 70.46] | 33.85 [32.37, 35.35] |
| Method | PSNR | SSIM | LPIPS | DreamSim |
|---|---|---|---|---|
| Source video | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | 0.000 [0.000, 0.000] | |
| AnyText2 | 25.56 [25.11, 26.06] | 0.905 [0.894, 0.915] | 0.091 [0.087, 0.096] | 0.043 [0.040, 0.046] |
| TextCtrl | 41.14 [40.68, 41.61] | 0.994 [0.994, 0.995] | 0.008 [0.007, 0.009] | 0.004 [0.004, 0.005] |
| FLUX-Text | 31.49 [31.13, 31.85] | 0.975 [0.973, 0.976] | 0.029 [0.027, 0.030] | 0.012 [0.011, 0.013] |
| RS-STE | 37.00 [36.71, 37.30] | 0.983 [0.982, 0.984] | 0.024 [0.022, 0.025] | 0.007 [0.007, 0.008] |
| TextCtrl + AnyV2V | 21.08 [20.67, 21.51] | 0.785 [0.769, 0.801] | 0.225 [0.212, 0.239] | 0.073 [0.066, 0.083] |
| Editor | Scope | Working resolution | Per-frame pipeline |
|---|---|---|---|
| AnyText2 | full | edit Lanczos | |
| FLUX-Text | full | FLUX native | FLUX native edit resample |
| TextCtrl | bounding box | bounding-box crop edit Lanczos alpha-composite back | |
| RS-STE | bounding box | bounding-box crop edit Lanczos alpha-composite back |
| Flicker | PSNR-loc | ||||
|---|---|---|---|---|---|
| Method | Raw | +Composite | Raw | +Composite | Raw SeqAcc |
| AnyText2 | 3.34 | 3.85 | 25.6 | 42.9 | 0.280 |
| TextCtrl | 3.80 | 3.78 | 41.1 | 43.0 | 0.475 |
| FLUX-Text | 5.11 | 4.47 | 31.5 | 42.9 | 0.528 |
| RS-STE | 3.73 | 3.71 | 37.0 | 43.0 | 0.354 |
| TextCtrl + AnyV2V | 4.98 | 3.77 | 21.1 | 43.0 | 0.057 |
| Resample size | Mean Kendall | Top method retained |
|---|---|---|
| 40 | 0.892 | 0.79 |
| 80 | 0.917 | 0.89 |
| 120 | 0.932 | 0.92 |
| 152 | 0.936 | 0.95 |
| Quantity | Observed value |
|---|---|
| Source exact match within | |
| Source CharAcc within | |
| Source decoded-string TTS | |
| Source detectability | Median ; 25th percentile ; mean |
| Clips with | of evaluation clips |
| Pipeline-rendered | Exact match ; CharAcc ; TTS |
| Axis | Ordinal | Automatic metric | Spearman |
|---|---|---|---|
| Text correctness | 0.87 | SeqAcc | |
| Temporal quality | 0.80 | Warp | |
| Edit locality | 0.37 | DreamSim-loc |
| Method | On front | SeqAcc | Warp | DreamSim-loc |
|---|---|---|---|---|
| FLUX-Text | Yes | 0.528 | 13.010 | 0.012 |
| TextCtrl | Yes | 0.475 | 2.088 | 0.004 |
| RS-STE | Yes | 0.354 | 1.815 | 0.007 |
| ViTeX-Edit-14B | Yes | 0.341 | 1.530 | 0.024 |
| Wan2.1-VACE-14B | Yes | 0.000 | 1.561 | 0.007 |
| AnyText2 | No | 0.280 | 3.952 | 0.043 |
| Method | BG-Warp | DINOv2-drift | ArcFace-id |
|---|---|---|---|
| Source | 1.67 | 0.0036 | 1.00 |
| TextCtrl | 1.70 | 0.0038 | 0.99 |
| ViTeX-Edit-14B | 1.74 | 0.0037 | 0.89 |
| RS-STE | 1.76 | 0.0038 | 0.97 |
| Wan2.1-VACE-14B | 1.88 | 0.0039 | 0.96 |
| AnyText2 | 2.04 | 0.0046 | 0.82 |