Memory-Guided B-Roll Generation from User Video Collections
Authors: Cusuh Ham, Fabian Caba Heilbron, Josef Sivic, Bryan Russell
Organizations: Adobe Research, Cambridge, USA · Adobe Research, San Jose, USA · Adobe Research, USA · Czech Institute of Informatics, Robotics and Cybernetics, Czech Technical University, Czech Republic · Adobe Research, San Francisco, USA
We introduce an approach for collection-grounded B-roll sequence generation. Given a user's video collection, a directive given in natural language, and a target duration, the goal is to produce a multi-shot sequence that complements the user's primary footage (A-roll) while preserving the collection's characters, settings, objects, and style. This task is challenging as one must choose the visual evidence from hours of captured footage that should guide the generation of each shot in the sequence. We address this challenge with MemComposer, a three-stage system that turns raw footage into a structured memory with visual references (characters, settings, objects, and style) and uses it to plan, retrieve, and generate grounded B-roll sequences. First, in a one-time offline stage, MemComposer constructs an entity-centric memory from raw video. Second, it uses the memory and user directive to plan a grounded sequence and retrieve conditioning frames for each shot. Third, it iteratively generates and critiques the sequence to enforce identity, setting, and sequence-level consistency. We evaluate MemComposer in a user preference study along two dimensions: prompt adherence and visual alignment to the user's collection. Against an ungrounded text-to-video planner, MemComposer wins 60.0% of prompt-adherence and 92.8% of visual-alignment comparisons, showing the grounding benefit of collection memory and reference retrieval. Against retrieval-only sequences assembled from captured footage, MemComposer wins 94.5% of prompt-adherence comparisons, showing the value of generating missing shots, while retrieval-only sequences are preferred for visual alignment in 58.2% of comparisons.
Figures & tables
Figure 1 . Given a user’s video collection comprising over 10 hours of captured footage, we introduce an approach for generating a B-roll sequence that adheres to a user request given in natural language while preserving the characters, settings, and style of the captured footage. Key to our approach is the ability to construct a memory representation of the video collection for retrieving target characters and settings. Please view these results in the companion video. Collection thumbnails are copyrighted and belong to EditStock.
Figure 2 . Overview of MemComposer. Given a captured user video collection, user request, and target duration, MemComposer outputs a B-roll sequence that adheres to the inputs. MemComposer comprises three stages: memory construction (performed offline and once for the video collection; Section 3.1 ), planning and reference retrieval (Section 3.2 ), and finally B-roll generation (Section 3.3 ). See the appendix for illustrations of the different stages. Collection thumbnails are copyrighted and belong to EditStock.
Table 1 . Comparison against baselines from A/B user studies. The green bar and number indicates the win rate for our approach, and the red bar indicates the win rate for the baseline. * indicates a p-value <0.02 .
Table 2 . Ablation study results from A/B user studies. * indicates a p-value <0.02 .
Figure 3 . Qualitative comparison. We compare MemComposer to the baselines on two collections. MemComposer better matches the setting from the input collection than Text2Video+Planner and better adheres to the user request than EditDuet. Collection thumbnails are copyrighted and belong to EditStock.
Figure 4 . Qualitative result of swapping characters. For the Built by Life collection, we swap in different characters (spanning different projects). Notice how each row maintains the setting from Built by Life while correctly depicting the target character. Collection thumbnails are copyrighted and belong to EditStock.
Figure 5 . Judges and critic outputs. (a) We illustrate the outputs from judges Jchar and Jset over the generated keyframe candidate, with scores out of 10. While the generated keyframe passes the setting judge Jset , it does not pass the character judge Jchar . (b) We illustrate the critic C over two generation iterations. In the first iteration, the critic deems that the second keyframe fails due to visual consistency. In the second iteration, both keyframes are deemed acceptable by the critic. The retrieved reference images are copyrighted and belong to EditStock.
Figure 6 . Example failures. Example failures include (top) implausible motion due to complex physics and (bottom) fine-grained scene characteristics are not enforced to be consistent across shots.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 . Stage 1: Memory construction. The memory construction stage comprises steps for collection analysis, frame selection, and computing character sheets. The bottom row illustrates example outputs of each of the steps for this stage. All thumbnails are copyrighted and belong to EditStock.
Figure 8 . Stage 2: Planning and reference retrieval. The planning and reference retrieval stage comprises steps for narrative planning, assembly planning, and reference retrieval. The bottom row illustrates example outputs of each of the steps for this stage. All thumbnails are copyrighted and belong to EditStock.
Figure 9 . Stage 3: B-roll generation. The B-roll generation stage comprises steps for keyframe generation and animation & assembly. A set of judges and a critic gives feedback to the keyframe generation step for regenerating keyframes for improved adherence to the inputs. The right side illustrates example outputs of each of the steps for this stage.
Stage
Hyperparameter
Value
Description
Memory
SAMPLE_INTERVAL
2
Base frame sampling interval (sec)
ρ
15
Target seconds-per-frame for frame budgeting
MIN_CANDIDATES
6
Min. number of candidate frames to score before filtering
MIN_FACE_QUALITY
5
Min. face quality to keep frame in memory
MIN_SETTING_QUALITY
5
Min. setting quality to keep frame in memory
NMS_WINDOW
5
Temporal NMS window (sec)
Appendix
Table 3 . Hyperparameters used in our pipeline.
Table 4 . Comparison against baselines from A/B user studies for longer output sequences ( > 30s). The green bar and number indicate the win rate for our approach, and the red bar indicates the win rate for the baseline. * indicates a p-value <0.02 .
Table 5 . Ablation study results from A/B user studies for longer output sequences ( > 30s). The green bar and number indicate the win rate for our full approach, and the red bar indicates the win rate for the row variant. * indicates a p-value <0.02 .
We present SubjectAnchor, a Subject-Aware Memory-to-Video paradigm for multi-shot storytelling in which the current shot is generated by conditioning on explicit visual memories extracted from previous shots. The objective is to preserve subject identity and scene consistency across cuts while retaining the controllability of shot-wise prompting. Built on Wan2.2-I2V-A14B, SubjectAnchor contains three key components: subject-related memory construction, subject-aware temporal rotary position encoding, and memory-aware attention partition. For each target shot, the method constructs a compact memory bank by tracing each required subject to its historical appearance and retrieving the most relevant precomputed keyframes. These memory frames are encoded into the model input as explicit visual conditions, while different subjects are assigned to separated negative temporal slots to reduce identity interference. In addition, memory-aware attention partition regulates the interaction between memory tokens and generated content within a shared backbone. This formulation preserves the appearance anchoring of explicit visual memory while remaining compatible with script-driven shot-by-shot generation. Experiments show that SubjectAnchor improves cross-shot identity consistency over representative memory-based and holistic baselines while maintaining competitive visual quality.
Xinyu Wang, Huafeng Shi, Zian Li +4
Shenzhen International Graduate School, Tsinghua University · Kling Team · Peking University +1
Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and locations across shots remains a challenge over long sequences. Existing evaluations typically use independently generated prompt sets with limited entity coverage and simple consistency metrics, making standardized comparison difficult. We introduce EntityBench, a benchmark of 140 episodes (2,491 shots) derived from real narrative media, with explicit per-shot entity schedules tracking characters, objects, and locations simultaneously across easy / medium / hard tiers of up to 50 shots, 13 cross-shot characters, 8 cross-shot locations, 22 cross-shot objects, and recurrence gaps spanning up to 48 shots. It is paired with a three-pillar evaluation suite that disentangles intra-shot quality, prompt-following alignment, and cross-shot consistency, with a fidelity gate that admits only accurate entity appearances into cross-shot scoring. As a baseline, we propose EntityMem, a memory-augmented generation system that stores verified per-entity visual references in a persistent memory bank before generation begins. Experiments show that cross-shot entity consistency degrades sharply with recurrence distance in existing methods, and that explicit per-entity memory yields the highest character fidelity (Cohen's d = +2.33) and presence among methods evaluated. Code and data are available at https://github.com/Catherine-R-He/EntityBench/.
While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, and objects throughout a story. Existing training-free approaches typically condition target shots using retrieved historical visuals. However, these references often suffer from severe informational mismatch, either introducing irrelevant contextual redundancy or failing to provide the full combination of required elements for the target shot. To resolve this, we present Complementary Retrieval-Augmented Prompting, an agentic framework that strategically aggregates a compact set of mutually supportive historical references to achieve complete and targeted conditioning for long-form video generation without retraining or modifying the underlying generator. Specifically, our framework explicitly models the visual elements required by each target shot by parsing the narrative script into a text-grounded visual element registry that tracks characters, objects, scenes, and their shot-level states. A VLM-annotated keyframe library further maps these elements to past visual observations. Guided by the required elements, our agent retrieves complementary references that maximize target-element coverage while minimizing historical noise. Finally, the retrieved references, structured element states, and grounding instructions are assembled into a unified prompt for the frozen video generator. This element-aware process provides comprehensive conditioning while remaining fully interpretable. Quantitative and qualitative evaluations on multi-shot story generation demonstrate that our method consistently outperforms recent-frame, memory-based, and entity-level retrieval baselines in cross-shot consistency and text-controllability.