Video interactive storytelling enables viewers to actively steer how a video unfolds. However, once we allow viewers to intervene during generation, a new challenge arises: The viewer's request can have latent dependencies on both the grounding source and the current rendered video state. These dependencies may not be explicitly stated in any individual input, but emerge only when the source, rendered history, and new viewer intent are considered jointly. Existing interactive video generation systems primarily emphasize following viewer instructions, while source-grounded video generation methods focus on aligning generated content with an external narrative or knowledge source. This leaves a fundamental question underexplored: What context should a generation model ground on during interactive continuation, and how can heterogeneous, unstructured inputs be transformed into such grounding context? In this work, we formulate contextual grounding as the process of transforming heterogeneous input context into an executable constraint model for video generation. To address this challenge, we introduce VIS-Ground, which performs Structured Context Abstraction to recover grounded states and cross-context dependencies, Generation Constraints Induction to project relevant dependencies into candidate-specific constraints, and Constrained Video Generation to enforce these constraints through planning, verification, revision, and rendering. Across three video generation backbones, VIS-Ground consistently achieves the highest overall composite score, reaching an average absolute improvement of 10.3 points over the strongest per-backbone baselines. Detailed analysis further shows gains across both narrative and knowledge grounding, and reveals remaining challenges in dependency extraction, and faithful realization during video rendering.
Figures & tables
Figure 1: Formative Study Results. The study enrolled nine volunteers. Left: Learning gains for the seven participants who passed the attention test. Right: Experience ratings from four respondents.
Figure 2: Overview of V I S - G r o u n d . Structured Context Abstraction aligns the video, story, and source states into M . Generation Constraints Induction derives dependencies R from M and activates candidate-specific constraints C . Constrained Video Generation uses script and video verification to repair the continuation vcont from the prefix vprefix and the viewer request.
Figure 3: Dataset Construction and Dataset Statistics
Figure 4: Cross-backbone comparison. Left: overall composite-score variation. Right: performance across the six evaluation dimensions, averaged over settings for each backbone.
Figure 5: V I S - G r o u n d Performance gain over the baselines. Left: performance across six metrics for three video generators under equal grounding weights. Right: our method’s JCS gain over the strongest baseline in narrative and knowledge settings (percentage points).
Figure 7: The comparator is the IF-leading baseline, held fixed across all four metrics in each panel.
Figure 8: Score distribution in the 72-output analysis sample.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Interaction Type
Source Faithfulness
Story Continuity
Interaction Fulfillment
Video storytelling methods
MM-StoryAgent ( Xu et al., 2025 )
—
—
✓
—
MovieAgent ( Wu et al., 2025b )
—
✓
✓
—
VideoGen-of-Thought ( Zheng et al., 2024 )
—
—
✓
—
Co-Director ( Song et al., 2026 )
—
△a
✓
—
ViMax ( Huang et al., 2026 )
—
✓b
✓
—
Appendix
Table 2: Comparison with related methods and benchmarks. ✓ : explicitly addressed; △ : partial coverage; —: not explicitly established. a only asset or script alignment; b only screenplay adaptation; c action-conditioned control; d only temporal coherence; e only world-setting adherence; f live video interaction; g image-sequence evaluation.
Role
Model identifier
Planner
gemini-3.1-pro-preview ( Google, 2026a )
Reference-image generator
gemini-2.5-flash-image ( Google, 2025 )
Gemini evaluator
gemini-3.8-flash ( Google, 2026b )
Claude evaluator
claude-opus-5 ( Anthropic, 2026 )
Omni renderer
gemini-omni-1.1-flash-preview ( Google, 2026c )
Veo renderer
veo-3.1-generate-001 ( Google DeepMind, 2025b )
Appendix
Table 3: Shared model configuration for continuation generation and evaluation.
Ordered level
Interpretation
0
Absent or contradictory
1
Mostly unsuccessful
2
Partially fulfilled
3
Mostly fulfilled
4
Fully fulfilled
Appendix
Table 4: Semantic rating anchors. Unknown evidence is recorded separately.
Figure 9: Pre-test interface. The pre-test screen asks participants to answer ten questions using their current knowledge, without correctness feedback. Each visible question includes an “I don’t know” option.
Figure 10: Noninteractive viewing interface. The noninteractive screen presents the video, subtitles, scene title, and viewing progress. The screenshot shows an Apollo 13 example.
Figure 11: Interactive viewing interface. The interactive screen provides a question field and an “Add to Queue” control below the video. The displayed instructions ask participants to submit at least three questions and state that playback pauses while they type.
Field
Study details
Agreement statistic and scale
Krippendorff’s α computed over the same ordinal rating scale used by the automatic evaluators.
Output sampling and coverage
50 generated outputs sampled across narrative and knowledge grounding, multiple generation methods and backbones, and different automatic-score ranges to obtain broad coverage of output quality.
Rater recruitment and training
Three human annotators familiar with multimodal content evaluation. Annotators first reviewed the metric definitions, rubric, and several worked examples before independently completing the validation set.
Ratings per output
Each output received three independent human ratings for SF, SC, IF, VC, and VQ.
Blinding and presentation order
Annotators were not shown the generation method, model identity, or automatic evaluator scores. Outputs were presented in independently randomized order.
Disagreement handling
Ratings were collected independently without discussion. For human–automatic agreement, the median human rating across the three annotators was used as the reference judgment.
Appendix
Table 5: Protocol used for human validation of the automatic evaluation.