Video interactive storytelling enables viewers to actively steer how a video unfolds. However, once we allow viewers to intervene during generation, a new challenge arises: The viewer's request can have latent dependencies on both the grounding source and the current rendered video state. These dependencies may not be explicitly stated in any individual input, but emerge only when the source, rendered history, and new viewer intent are considered jointly. Existing interactive video generation systems primarily emphasize following viewer instructions, while source-grounded video generation methods focus on aligning generated content with an external narrative or knowledge source. This leaves a fundamental question underexplored: What context should a generation model ground on during interactive continuation, and how can heterogeneous, unstructured inputs be transformed into such grounding context? In this work, we formulate contextual grounding as the process of transforming heterogeneous input context into an executable constraint model for video generation. To address this challenge, we introduce VIS-Ground, which performs Structured Context Abstraction to recover grounded states and cross-context dependencies, Generation Constraints Induction to project relevant dependencies into candidate-specific constraints, and Constrained Video Generation to enforce these constraints through planning, verification, revision, and rendering. Across three video generation backbones, VIS-Ground consistently achieves the highest overall composite score, reaching an average absolute improvement of 10.3 points over the strongest per-backbone baselines. Detailed analysis further shows gains across both narrative and knowledge grounding, and reveals remaining challenges in dependency extraction, and faithful realization during video rendering.
Figures & tables
Figure 1: Formative Study Results. The study enrolled nine volunteers. Left: Learning gains for the seven participants who passed the attention test. Right: Experience ratings from four respondents.
Figure 2: Overview of V I S - G r o u n d . Structured Context Abstraction aligns the video, story, and source states into M . Generation Constraints Induction derives dependencies R from M and activates candidate-specific constraints C . Constrained Video Generation uses script and video verification to repair the continuation vcont from the prefix vprefix and the viewer request.
Figure 3: Dataset Construction and Dataset Statistics
Figure 4: Cross-backbone comparison. Left: overall composite-score variation. Right: performance across the six evaluation dimensions, averaged over settings for each backbone.
Figure 5: V I S - G r o u n d Performance gain over the baselines. Left: performance across six metrics for three video generators under equal grounding weights. Right: our method’s JCS gain over the strongest baseline in narrative and knowledge settings (percentage points).
Figure 7: The comparator is the IF-leading baseline, held fixed across all four metrics in each panel.
Figure 8: Score distribution in the 72-output analysis sample.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Interaction Type
Source Faithfulness
Story Continuity
Interaction Fulfillment
Video storytelling methods
MM-StoryAgent ( Xu et al., 2025 )
—
—
✓
—
MovieAgent ( Wu et al., 2025b )
—
✓
✓
—
VideoGen-of-Thought ( Zheng et al., 2024 )
—
—
✓
—
Co-Director ( Song et al., 2026 )
—
△a
✓
—
ViMax ( Huang et al., 2026 )
—
✓b
✓
—
Appendix
Table 2: Comparison with related methods and benchmarks. ✓ : explicitly addressed; △ : partial coverage; —: not explicitly established. a only asset or script alignment; b only screenplay adaptation; c action-conditioned control; d only temporal coherence; e only world-setting adherence; f live video interaction; g image-sequence evaluation.
Role
Model identifier
Planner
gemini-3.1-pro-preview ( Google, 2026a )
Reference-image generator
gemini-2.5-flash-image ( Google, 2025 )
Gemini evaluator
gemini-3.8-flash ( Google, 2026b )
Claude evaluator
claude-opus-5 ( Anthropic, 2026 )
Omni renderer
gemini-omni-1.1-flash-preview ( Google, 2026c )
Veo renderer
veo-3.1-generate-001 ( Google DeepMind, 2025b )
Appendix
Table 3: Shared model configuration for continuation generation and evaluation.
Ordered level
Interpretation
0
Absent or contradictory
1
Mostly unsuccessful
2
Partially fulfilled
3
Mostly fulfilled
4
Fully fulfilled
Appendix
Table 4: Semantic rating anchors. Unknown evidence is recorded separately.
Figure 9: Pre-test interface. The pre-test screen asks participants to answer ten questions using their current knowledge, without correctness feedback. Each visible question includes an “I don’t know” option.
Figure 10: Noninteractive viewing interface. The noninteractive screen presents the video, subtitles, scene title, and viewing progress. The screenshot shows an Apollo 13 example.
Figure 11: Interactive viewing interface. The interactive screen provides a question field and an “Add to Queue” control below the video. The displayed instructions ask participants to submit at least three questions and state that playback pauses while they type.
Field
Study details
Agreement statistic and scale
Krippendorff’s α computed over the same ordinal rating scale used by the automatic evaluators.
Output sampling and coverage
50 generated outputs sampled across narrative and knowledge grounding, multiple generation methods and backbones, and different automatic-score ranges to obtain broad coverage of output quality.
Rater recruitment and training
Three human annotators familiar with multimodal content evaluation. Annotators first reviewed the metric definitions, rubric, and several worked examples before independently completing the validation set.
Ratings per output
Each output received three independent human ratings for SF, SC, IF, VC, and VQ.
Blinding and presentation order
Annotators were not shown the generation method, model identity, or automatic evaluator scores. Outputs were presented in independently randomized order.
Disagreement handling
Ratings were collected independently without discussion. For human–automatic agreement, the median human rating across the three annotators was used as the reference judgment.
Appendix
Table 5: Protocol used for human validation of the automatic evaluation.
While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, and objects throughout a story. Existing training-free approaches typically condition target shots using retrieved historical visuals. However, these references often suffer from severe informational mismatch, either introducing irrelevant contextual redundancy or failing to provide the full combination of required elements for the target shot. To resolve this, we present Complementary Retrieval-Augmented Prompting, an agentic framework that strategically aggregates a compact set of mutually supportive historical references to achieve complete and targeted conditioning for long-form video generation without retraining or modifying the underlying generator. Specifically, our framework explicitly models the visual elements required by each target shot by parsing the narrative script into a text-grounded visual element registry that tracks characters, objects, scenes, and their shot-level states. A VLM-annotated keyframe library further maps these elements to past visual observations. Guided by the required elements, our agent retrieves complementary references that maximize target-element coverage while minimizing historical noise. Finally, the retrieved references, structured element states, and grounding instructions are assembled into a unified prompt for the frozen video generator. This element-aware process provides comprehensive conditioning while remaining fully interpretable. Quantitative and qualitative evaluations on multi-shot story generation demonstrate that our method consistently outperforms recent-frame, memory-based, and entity-level retrieval baselines in cross-shot consistency and text-controllability.
Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, undermining both narrative coherence and visual consistency. To address these challenges, we propose StoryEngine, a state-grounded agentic framework for video storytelling. StoryEngine establishes a separation between authoritative semantic plans and unreliable visual observations. Specifically, StoryEngine maintains a structured representation of entity placement and story-relevant states, and propagates event-induced changes to define the intended start and end states of each shot. To visually realize these states, StoryEngine constructs canonical references for recurring entities and environments, and compiles state and visual constraints into executable render plans. Meanwhile, to realize these states correctly, a bounded evaluation-guided repair loop further corrects local state inconsistencies. Together, these mechanisms preserve causal story progression and prevent local visual errors from propagating across shots. To comprehensively evaluate long-form storytelling, we construct a benchmark across diverse scenarios and visual styles, with metrics assessing storytelling quality, narrative coherence, and visual consistency. Experimental results demonstrate that StoryEngine consistently outperforms state-of-the-art methods across all evaluation dimensions, validating its effectiveness for coherent and consistent video storytelling.
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow a reference, or edit source footage while maintaining shared history. We formalize this setting as interactive multi-shot video creation (IMVC) and introduce ContextMaster, a unified model with a role-aware context representation for these operations. An interactive model must retain access to an expanding history without allowing the context read cost at each denoising step to grow. ContextMaster combines reusable clean context states with fixed budget sparse context routing and uses ConstraintSink to keep task constraints visible. To address the dual challenges of sparse context access and inference with few denoising steps, we propose a two-stage privileged context distillation framework, which transfers full context behavior from a dense teacher through consistency distillation and then refines deployment rollouts with distribution matching. Experiments on the three primitive tasks demonstrate improved task fulfillment and consistency across shots over specialized baselines. User studies further validate flexibly composed workflows, while the model reaches 16 FPS on a single GPU.
Xu Guo, Zhengxuan Wei, Xinghui Li +11
1Tsinghua University · 2Nanjing University · 3Kling Team, Kuaishou Technology +2