SubjectAnchor: Subject-Aware Memory-to-Video for Multi-Shot Storytelling
Authors: Xinyu Wang, Huafeng Shi, Zian Li, Yan Zhou, Xiaoqiang Liu, Yue Ma, Pengfei Wan
Organizations: Shenzhen International Graduate School, Tsinghua University · Kling Team · Peking University · The Hong Kong University of Science and Technology
We present SubjectAnchor, a Subject-Aware Memory-to-Video paradigm for multi-shot storytelling in which the current shot is generated by conditioning on explicit visual memories extracted from previous shots. The objective is to preserve subject identity and scene consistency across cuts while retaining the controllability of shot-wise prompting. Built on Wan2.2-I2V-A14B, SubjectAnchor contains three key components: subject-related memory construction, subject-aware temporal rotary position encoding, and memory-aware attention partition. For each target shot, the method constructs a compact memory bank by tracing each required subject to its historical appearance and retrieving the most relevant precomputed keyframes. These memory frames are encoded into the model input as explicit visual conditions, while different subjects are assigned to separated negative temporal slots to reduce identity interference. In addition, memory-aware attention partition regulates the interaction between memory tokens and generated content within a shared backbone. This formulation preserves the appearance anchoring of explicit visual memory while remaining compatible with script-driven shot-by-shot generation. Experiments show that SubjectAnchor improves cross-shot identity consistency over representative memory-based and holistic baselines while maintaining competitive visual quality.
Figures & tables
Figure 1: Showcase of SubjectAnchor. Our framework decouples complex video narratives into global background/subject settings (“Where” & “Who”) and local shot-by-shot action descriptions (“What”). By explicitly defining subject features and enforcing shot-level inclusion/exclusion constraints, it achieves multi-subject identity consistency and controls complex action interactions across consecutive shots. The dashed bounding boxes in the generated sequences demonstrate the accurate identity mapping of the corresponding subjects.
Figure 2: Comparison of multi-shot video generation paradigms. (a) Holistic Modeling faces text embedding limitations and an O(N2S2) scalability challenge. (b) Previous shot-by-shot methods improve efficiency ( O(NS2) ) but suffer from a mismatch challenge caused by unstructured FIFO memory buffers. (c) Our SubjectAnchor overcomes these limitations by utilizing a structured, identity-aware memory bank, allowing for selective retrieval of historical frames guided by local inclusion/exclusion constraints.
Figure 3: Data curation pipeline for Subject-Aware Memory-to-Video. Raw videos are segmented into shots, filtered into valid multi-shot stories, annotated with hierarchical story-level and shot-level descriptions, and paired with subject-centric keyframes. The resulting structured records provide the annotations and memory references used by SubjectAnchor.
Figure 4: Overview of the proposed framework. A target shot first retrieves subject-related memories from previous shots, then encodes them with the VAE, constructs the memory/video mask tensor, injects the resulting representation into the backbone, and organizes memory frames with subject-aware RoPE before diffusion decoding.
Figure 5: Qualitative comparison of multi-shot video generation. SubjectAnchor effectively preserves multi-subject identity and strictly adheres to fine-grained inclusion/exclusion constraints across consecutive shots. Baseline methods struggle with identity inconsistency, subject confusion, and failure to follow local action or exclusion prompts.
Method
Global Align. ↑
Per-shot Align. ↑
Subject Align. ↑
Identity-First ↑
Identity-Prev ↑
Aesthetic ↑
StoryDiffusion Zhou et al. [2024] +Wan2.2 Wan et al. [2025a]
0.3011
0.2310
0.2353
0.8416
0.8168
7.0332
StoryMem Zhang et al. [2025a]
0.3098
0.2698
0.2548
0.7250
0.7228
6.2093
HoloCine Meng et al. [2025]
0.2818
0.2686
0.2284
0.4477
0.4433
4.6530
Ours
0.3243
0.2729
0.2814
0.8689
0.8681
6.5625
Table 1: Main comparison on multi-shot narrative generation. Red and Blue denote the best and second best results.
Figure 6: Visualization of keyframes under different selection strategies. We compare keyframes retrieved from the same video using three selection strategies. Under “W/o Matching Eval”, the subject is less salient. Under “W/o Aesthetic Eval”, the subject is prominent, but the face is blurred. Our method balances subject relevance and frame quality.
Global Align. ↑
Per-shot Align. ↑
Subject Align. ↑
Identity-First ↑
Identity-Prev ↑
Aesthetic ↑
W/o Matching Eval
0.3132
0.2514
0.2337
0.8288
0.8276
6.4506
W/o Aesthetic Eval
0.3117
0.2461
0.2383
0.8506
0.8513
6.4408
Ours
0.3243
0.2729
0.2814
0.8689
0.8681
6.5625
Table 2: Memory selection ablation. Red and Blue denote the best and second best results.
Global Align. ↑
Per-shot Align. ↑
Subject Align. ↑
Identity-First ↑
Identity-Prev ↑
Aesthetic ↑
W/o Subject-Aware RoPE
0.3196
0.2537
0.2556
0.8476
0.8553
6.4696
W/o Memory-Aware Attention Partition
0.3150
0.2533
0.2507
0.8670
0.8600
6.4181
W/o Both
0.3165
0.2564
0.2574
0.8097
0.8176
6.4587
Ours
0.3243
0.2729
0.2814
0.8689
0.8681
6.5625
Table 3: Memory injection ablation. Red and Blue denote the best and second best results.
Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions, and scene transitions. Existing temporal decomposition methods improve scalability by generating videos shot by shot. However, they mainly focus on optimizing plausible next-shot continuations without verifying whether the historical memory preserves identity-critical subject evidence. Consequently, as generation proceeds, recurring subjects may be diluted, overwritten, or forgotten. In this paper, we propose Memento, a subject-reconstruction-guided framework that treats subject preservation as an explicit identity grounding problem, based on the premise that a memory bank faithfully preserving a subject should support reconstructing that subject from memory alone. Specifically, Memento jointly trains autoregressive next-shot generation with memory-based subject reconstruction, recovering target appearances using historical memory and global story captions. To disentangle long-range subject evidence from short-range cues, Memento introduces a dual-query memory mechanism, where one query retrieves identity-relevant memory and the other selects short-context keyframes for coherent continuation. Additionally, a subject-aware cinematic data pipeline provides precise reconstruction supervision via consistent, pronoun-free subject descriptions. Experiments demonstrate that Memento achieves state-of-the-art performance in long-term subject consistency, cross-shot coherence, and visual quality.
Multi-shot video generation requires maintaining a consistent appearance of recurring entities across shots while remaining faithful to shot-specific text prompts. Recent autoregressive methods reuse previously generated frames as memory. However, full-frame storage entangles persistent entity information with transient scene context, leading to irrelevant information leakage and high computational cost. We propose an entity-centric memory in the form of an entity-indexed bank of latent patches. We introduce sparse token conditioning compatible with pretrained models, restricting self-attention to entity-relevant tokens and reducing computational cost. To support this, we introduce a structured multi-shot script format. We additionally propose a budgeted memory update strategy to maintain a compact, evolving memory. Finally, we equip the entity representation with a noise-injection mechanism that enables fine-grained appearance control, preventing leakage of irrelevant information. Our method improves prompt adherence and efficiency while preserving subject consistency.
Jente Vandersanden, Matheus Gadelha, Chun-Hao P. Huang +2
Max Planck Institute for Informatics, Germany and Adobe Research, UK · Adobe Research, UK · Adobe Research, USA
Multi-shot long-form video generation remains challenging due to identity drift and compounding inconsistencies across shots. While storyboard-driven pipelines improve controllability, they are often executed in a feed-forward manner, with limited mechanisms to incorporate generated visual evidence back into subsequent conditioning. We propose CoTriSyGen, an agentic framework that formulates multi-shot long video generation as a closed-loop visual-text-memory synergy process, where planned intent, persistent memory, and generated visuals are jointly leveraged for iterative correction and long-range coherence. A vision-language-model-based analyzer reasons over this triplet and produces updates to both prompts and memory along two pathways: (i) intra-shot refinement, which triggers targeted regeneration when semantic or compositional violations are detected and refines image-to-video prompt for coherent motions; and (ii) inter-shot refinement, which rewrites subsequent-shot prompts to propagate newly manifested entities or attributes and improve prompt quality (e.g., compositional grounding and cinematic fluency) based on generated evidence. The loop is grounded in an entity-centric memory modeled as a mutable visual state that evolves as the story progresses, which is continuously updated by both the generator and the analyzer by adding new and evolved entities to reflect appearance changes, accumulated multi-view evidence, and multi-entity compositions. Experiments on our curated StoryBench benchmark demonstrate substantial improvements in cross-shot consistency, prompt adherence, and cinematic continuity over representative methods.
Xinlei Yin, Xiulian Peng, Xiao Li +2
Microsoft Research Asia · University of Science and Technology of China