Towards Unified Evaluation of Prompt Enhancers for Video Generation
Organizations: University of Science and Technology of China
Abstract
Modern video generators can realize increasingly complex visual narratives, positioning the prompt enhancer (PE) as a critical bridge from concise user instructions and multimodal references to structured cinematic plans. However, existing PE evaluation relies on rendered videos, imposing substantial computational and human costs, slowing PE training and iteration, and conflating PE quality with downstream generator behavior. To address this gap, we introduce PEBench, the first unified benchmark for direct PE evaluation across text-to-video, image-to-video, and reference-to-video prompt enhancement. It comprises 1,100 expert-verified cases and 1,005 visual assets, spanning 35 fine-grained tasks with diverse temporal, cinematic, audiovisual, and multi-reference requirements. In addition, we develop PEBench evaluation, an evidence-grounded framework that combines modality-aware fact extraction with rubric-based assessment across 24 criteria. Our systematic evaluation of representative open- and closed-source PE methods reveals an emerging shift from fine-grained descriptive expansion toward intent-preserving cinematic planning, while the caption-reconstruction and forward-refinement methods show complementary strengths in cinematic coverage and semantic fidelity or internal coherence, respectively. Human validation shows that PEBench scores align closely with expert judgments of enhanced prompts and downstream videos from Wan3.0 and MiniMax-H3, indicating that prompt-level evaluation reliably reflects downstream utility.
Figures & tables
| Benchmark | T2V | I2V | R2V | Cine. | Scripted | Multi- shot | Audio Req. | Multi- lang. | Video Ref. | Direct PE | # Metrics | # Prompts | # Ref. Assets |
| VBench [ 18 ] | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 16 | 1,600 | – |
| EvalCrafter [ 19 ] | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 17 | 700 | – |
| Video-Bench [ 20 ] | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 9 | 419 | – |
| OpenS2V-Nexus [ 21 ] | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 6 | 180 | 180 |
| UniVBench [ 22 ] | ✓ | ✗ | ✓ | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | 21 | 200 | 864 |
| MSAVBench [ 23 ] | ✓ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ | ✗ | 20 | 286 | 165 |
| Method | Semantic Fidelity | Internal Consistency | |||||||||||||||
| T2V-PE | |||||||||||||||||
| WanPE-397B | 89.20 | 70.00 | 80.80 | 94.20 | 62.00 | 57.40 | 80.40 | 80.60 | 94.80 | 74.89 | 91.20 | 87.40 | 94.60 | 82.40 | 92.80 | 94.40 | 91.51 |
| WanPE-35B | 86.40 | 71.80 | 82.60 | 88.00 | 59.60 | 54.40 | 73.20 | 81.20 | 97.00 | 72.83 | 91.60 | 89.20 | 92.40 | 81.60 | 91.60 | 90.20 | 91.40 |
| H3-Context-IR | 81.40 | 66.40 | 70.40 | 93.00 | 58.40 | 49.20 | 76.80 | 81.20 | 95.40 | 67.34 | 97.20 | 95.40 | 96.80 | 98.00 | 91.80 | 90.80 | 94.50 |
| SCMaPR | 94.60 | 87.60 | 85.60 | 80.20 | 82.00 | 87.20 | 90.60 | 92.60 | 95.00 | 87.92 | 94.60 | 92.60 | 96.80 | 91.40 | 93.60 | 95.60 | 95.06 |
| T2V-PE | I2V-PE | R2V-PE | |||||||
| Evaluation | SF | IC | DC | SF | IC | DC | SF | IC | DC |
| Enhanced-prompt alignment | |||||||||
| Direct VLM Scoring | 0.618 | 0.464 | 0.527 | 1.000 | 0.600 | 0.400 | 1.000 | 0.500 | 0.500 |
| PEBench Evaluation | 0.955 | 0.818 | 0.909 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| Downstream-video alignment | |||||||||
| PEBench Evaluation (Wan3.0) | 0.845 | 0.827 | 0.945 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| High-level category | # | Fine-grained task types |
| Action | 4 | Daily Activity, Combat, Sports, Dance |
| Animation | 2 | 2D Animation, 3D Animation |
| Visual Effects | 5 | VFX Character, VFX Object, Transformation, Light/Weather Effects, Sci-Fi Creature |
| General | 5 | Physical Laws, On-Screen Text, Portrait, Nature, Intellectual Property (IP) |
| Narrative | 3 | Modern Story, Historical Story, Fantasy Story |
| Cinematography | 3 | Camera Motion, Time-Lapse, Lighting and Color |
| Setting | Min. | Q1 | Median | Mean | Q3 | Max. |
| T2V-PE | 6 | 85.3 | 237.5 | 425.8 | 597.3 | 3,431 |
| I2V-PE | 8 | 91.8 | 221.5 | 448.9 | 642.8 | 4,120 |
| R2V-PE | 13 | 89.8 | 210.5 | 363.8 | 549.0 | 2,412 |
| All | 6 | 88.8 | 224.0 | 415.2 | 596.8 | 4,120 |
| Task | Metric | ||||
| T2V-PE | Avg. tokens | 394.7 | 490.9 | 584.9 | 861.1 |
| DC | 79.32 | 78.60 | 77.35 | 74.13 | |
| I2V-PE | Avg. tokens | 819.1 | 945.7 | 1066.6 | 1382.0 |
| DC | 75.00 | 75.20 | 75.99 | 67.20 | |
| R2V-PE | Avg. tokens | 809.9 | 1015.7 | 1164.4 | 1479.3 |
| DC | 79.37 | 73.75 | 72.07 | 69.28 |