cs.CVSep 28, 2026

SubjectAnchor: Subject-Aware Memory-to-Video for Multi-Shot Storytelling

Authors: Xinyu Wang, Huafeng Shi, Zian Li, Yan Zhou, Xiaoqiang Liu, Yue Ma, Pengfei Wan

Organizations: Shenzhen International Graduate School, Tsinghua University · Kling Team · Peking University · The Hong Kong University of Science and Technology

Abstract

We present SubjectAnchor, a Subject-Aware Memory-to-Video paradigm for multi-shot storytelling in which the current shot is generated by conditioning on explicit visual memories extracted from previous shots. The objective is to preserve subject identity and scene consistency across cuts while retaining the controllability of shot-wise prompting. Built on Wan2.2-I2V-A14B, SubjectAnchor contains three key components: subject-related memory construction, subject-aware temporal rotary position encoding, and memory-aware attention partition. For each target shot, the method constructs a compact memory bank by tracing each required subject to its historical appearance and retrieving the most relevant precomputed keyframes. These memory frames are encoded into the model input as explicit visual conditions, while different subjects are assigned to separated negative temporal slots to reduce identity interference. In addition, memory-aware attention partition regulates the interaction between memory tokens and generated content within a shared backbone. This formulation preserves the appearance anchoring of explicit visual memory while remaining compatible with script-driven shot-by-shot generation. Experiments show that SubjectAnchor improves cross-shot identity consistency over representative memory-based and holistic baselines while maintaining competitive visual quality.

Figures & tables

Explore similar work

CardsList
  1. Memento: Reconstruct to Remember for Consistent Long Video Generation

    Jun 12, 2026Xuan Wei, Longbin Ji, Guan Wang +5Visual MemoryLong-Term Memory

  2. EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation

    May 22, 2026Jente Vandersanden, Matheus Gadelha, Chun-Hao P. Huang +2Interactive Video GenerationTraining-Free

  3. Closed-Loop Triplet Synergistic Generation for Long-Form Video

    Jun 15, 2026Xinlei Yin, Xiulian Peng, Xiao Li +2Interactive Video GenerationVisual Memory