cs.CVSep 27, 2026

StoryEngine: A State-Grounded Agentic Framework for Video Storytelling

Authors: Yingrui Wang, Zeqing Wang, Yeying Jin

Organizations: Tencent · National University of Singapore

Abstract

Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, undermining both narrative coherence and visual consistency. To address these challenges, we propose StoryEngine, a state-grounded agentic framework for video storytelling. StoryEngine establishes a separation between authoritative semantic plans and unreliable visual observations. Specifically, StoryEngine maintains a structured representation of entity placement and story-relevant states, and propagates event-induced changes to define the intended start and end states of each shot. To visually realize these states, StoryEngine constructs canonical references for recurring entities and environments, and compiles state and visual constraints into executable render plans. Meanwhile, to realize these states correctly, a bounded evaluation-guided repair loop further corrects local state inconsistencies. Together, these mechanisms preserve causal story progression and prevent local visual errors from propagating across shots. To comprehensively evaluate long-form storytelling, we construct a benchmark across diverse scenarios and visual styles, with metrics assessing storytelling quality, narrative coherence, and visual consistency. Experimental results demonstrate that StoryEngine consistently outperforms state-of-the-art methods across all evaluation dimensions, validating its effectiveness for coherent and consistent video storytelling.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Closed-Loop Triplet Synergistic Generation for Long-Form Video

    Jun 15, 2026Xinlei Yin, Xiulian Peng, Xiao Li +2Interactive Video GenerationVisual Memory

  2. DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior

    Apr 19, 2026Junjia Huang, Binbin Yang, Pengxiang Yan +6Interactive Video GenerationVideo Diffusion Models