cs.CVSep 28, 2026

SkillPE: Creativity-Oriented Cinematic Skill Evolution for Text-to-Video Prompt Engineering

Authors: Yanwei Huang, Mingxuan Zhu, Shujie Li, Shiyuan Liu, Yuanxing Zhang, Arpit Narechania

Organizations: HKUST · KlingAI · Shanghai Jiaotong University · The University of Hong Kong

Abstract

Achieving high-quality, cinematic results in text-to-video generation remains challenging for non-experts, whose prompts often lack professional narrative and creative design. We propose SkillPE, a prompt engineering (PE) framework that evolves reusable cinematic skills from expert-authored seeds. SkillPE represents shot logic, composition, lighting, sound design, and other filmmaking cues in a fine-grained format, and retrieves movie references categorized as resonators (good matches), dissonants (weak matches), and divergents (creatively useful near-misses). The first two refine when and how a skill should be applied, while divergents inspire alternative cinematic realizations at different degrees of modification while preserving the user intent. Candidate skills are assessed through generated videos along prompt fidelity, cinematic quality, narrative appeal, and creativity to construct the final skill libraries. Experiments on StoryEval and VBench show improvements of up to 1.40 points over the strongest external baseline and 0.51 points over seed skills on 7-point four-dimensional evaluation, while remaining competitive on benchmark-native metrics. Overall, SkillPE offers a practical approach to balancing fidelity and creativity in cinematic text-to-video generation. Code is available at https://github.com/Ais0n/SkillPE .

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

    Sep 24, 2026Yubo Zhu, Yawen Shao, Ziyun Dai +27Video Generation

  2. FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

    Jul 27, 2026Shengyi Wang, Niantong Li, Guangzheng Hu +27Cinematic IdealLong-Video Benchmarks

  3. VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

    Oct 1, 2026Yu Huang, Jungang Li, Zhiyuan Wang +8Text-To-Video Generation ModelVideo Generation