Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by approximately 11K blind pairwise assessments. When powering Wan3.0's video generator, WanPE-397B boosts human preference over raw user prompts by 10.66-18.84 points at 5-15 seconds and by a dramatic 50.86 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.
Figures & tables
Figure 1 : Left: Earlier models use optional prompt expansions, while modern video generators can process longer contexts and follow complex instructions. Right: t-SNE shows WanPE outputs align closely with video-grounded captions, while forward-based enhancement exhibits a distribution gap.
Figure 2 : WanPE training pipeline. Video-grounded reverse SFT and semantic-consistency GRPO.
5 seconds
10 seconds
15 seconds
Overall
Method
S
BT
S
BT
S
BT
S
BT
LTX-2.5
18.69
31.23
18.47
34.74
16.10
32.20
17.30
33.02
Kling 3.0
31.11
43.02
23.47
39.84
16.10
34.07
20.80
37.40
HappyHorse 1.1
20.28
32.82
24.00
40.72
26.80
42.39
24.89
40.46
MiniMax-H3
29.09
41.24
36.46
49.59
38.30
51.13
36.40
49.27
Seedance 2.0
46.20
54.20
41.40
52.16
44.31
56.90
43.42
54.70
Table 1 : Expert preference score and Bradley–Terry score on the 5 – 15 seconds subset of WanPEval.
Method
Action
Anim.
Speech
Ad.
Sing & Dance
Drama
Know.
Overall
Exp1: 30-second subset
Seedance 2.5
68.18
46.43
55.00
81.25
60.00
56.00
60.71
59.76
downstream generator: Wan3.0’s video generator
Original request
4.17
9.38
15.79
5.00
22.50
4.00
3.57
9.38
+WanPE-397B
50.00
81.25
73.68
75.00
47.50
53.85
57.14
60.24
Exp2: Reverse vs. forward enhancement
Table 2 : Preference scores S on WanPEval. (i) 30-second subset against Seedance 2.5, under Wan3.0. (ii) Reverse-constructed vs. forward-based enhancement, under Wan3.0. (iii) Format-adapted WanPE vs. native enhancers on LTX-2.5 and MiniMax-H3 in two separate battles.
(i) Semantic consistency
Method
Overall ↑
5 seconds ↑
10 seconds ↑
15 seconds ↑
30 seconds ↑
Perfect ↑
Failure ↓
Fwd. Rewriting
90.7
93.5
91.8
90.3
89.4
55.0
10.0
LTX-2.5-PE
77.3
76.1
81.6
74.9
77.0
33.7
38.6
H3-Context-IR
88.3†
89.4
88.8
87.6
–
54.2†
17.3†
Ours
WanPE-4B-SFT
66.5
67.0
68.4
66.7
64.2
20.1
56.6
Table 3 : (i) Semantic consistency evaluated by gemini-3.1-pro-preview. † denotes results computed on the 5–15-second subset. (ii) Expert preference S for 397B SFT vs. WanPE-397B under Wan3.0.
Achieving high-quality, cinematic results in text-to-video generation remains challenging for non-experts, whose prompts often lack professional narrative and creative design. We propose SkillPE, a prompt engineering (PE) framework that evolves reusable cinematic skills from expert-authored seeds. SkillPE represents shot logic, composition, lighting, sound design, and other filmmaking cues in a fine-grained format, and retrieves movie references categorized as resonators (good matches), dissonants (weak matches), and divergents (creatively useful near-misses). The first two refine when and how a skill should be applied, while divergents inspire alternative cinematic realizations at different degrees of modification while preserving the user intent. Candidate skills are assessed through generated videos along prompt fidelity, cinematic quality, narrative appeal, and creativity to construct the final skill libraries. Experiments on StoryEval and VBench show improvements of up to 1.40 points over the strongest external baseline and 0.51 points over seed skills on 7-point four-dimensional evaluation, while remaining competitive on benchmark-native metrics. Overall, SkillPE offers a practical approach to balancing fidelity and creativity in cinematic text-to-video generation. Code is available at https://github.com/Ais0n/SkillPE .
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot planning. To address this challenge, we propose ShotPlan, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model. Our method introduces learnable planning tokens that capture shot-level transition cues and can be seamlessly integrated with the original video generation tokens to control transition timestamps. Unlike standard video generation tokens, the proposed planning tokens are equipped with Fractional Temporal Rotary Position Embedding (FRoPE), enabling shot transitions to be modeled at the frame level. Experiments demonstrate that ShotPlan significantly outperforms existing cinematic video generation methods, offering more flexible shot management and stronger inter-shot consistency.
Su Guo, Guangce Liu, Haosen Yang +7
Institute of Artificial Intelligence (TeleAI), China Telecom · Harbin Institute of Technology
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.
Shengyi Wang, Niantong Li, Guangzheng Hu +27
Alibaba Group · Moku Lab, Hujing Digital Media & Entertainment Group · Beijing Film Academy