cs.CVSep 29, 2026

Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generation

Authors: Xianghan Wei, Xiaoda Yang, Zhi Wang, An Pan, Daoan Zhang, Huayi Zhang, Yan Zhang, Wei Xu, +2 more

Abstract

While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, and objects throughout a story. Existing training-free approaches typically condition target shots using retrieved historical visuals. However, these references often suffer from severe informational mismatch, either introducing irrelevant contextual redundancy or failing to provide the full combination of required elements for the target shot. To resolve this, we present Complementary Retrieval-Augmented Prompting, an agentic framework that strategically aggregates a compact set of mutually supportive historical references to achieve complete and targeted conditioning for long-form video generation without retraining or modifying the underlying generator. Specifically, our framework explicitly models the visual elements required by each target shot by parsing the narrative script into a text-grounded visual element registry that tracks characters, objects, scenes, and their shot-level states. A VLM-annotated keyframe library further maps these elements to past visual observations. Guided by the required elements, our agent retrieves complementary references that maximize target-element coverage while minimizing historical noise. Finally, the retrieved references, structured element states, and grounding instructions are assembled into a unified prompt for the frozen video generator. This element-aware process provides comprehensive conditioning while remaining fully interpretable. Quantitative and qualitative evaluations on multi-shot story generation demonstrate that our method consistently outperforms recent-frame, memory-based, and entity-level retrieval baselines in cross-shot consistency and text-controllability.

Figures & tables

Explore similar work

CardsList
  1. Closed-Loop Triplet Synergistic Generation for Long-Form Video

    Jun 15, 2026Xinlei Yin, Xiulian Peng, Xiao Li +2Interactive Video GenerationVisual Memory

  2. Memento: Reconstruct to Remember for Consistent Long Video Generation

    Jun 12, 2026Xuan Wei, Longbin Ji, Guan Wang +5Visual MemoryLong-Term Memory

  3. StoryEngine: A State-Grounded Agentic Framework for Video Storytelling

    Sep 27, 2026Yingrui Wang, Zeqing Wang, Yeying JinVideo StorytellingNarratives