cs.CVJun 1, 2025

MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing

Authors: Tong Zhang, Victor Escorcia, Juan C Leon Alcazar, Bernard Ghanem

Organizations: Department of Computer Science King Abdullah University of Science and Technology Thuwal, Saudi Arabia · Independent Researcher

Abstract

Unlike traditional video editing or inpainting, video semantic mixing fuses a reference concept with a moving target entity to produce a hybrid while preserving the source video's motion and layout. We propose MoCA-Video, a training-free framework that steers a frozen video-diffusion denoising trajectory through concept-localized reference injection. At selected low-noise steps, MoCA-Video uses concept attention to localize the target object and injects the reference latent into the localized region, where object structure has formed but appearance remains editable. A momentum-based correction carries the injected prediction across frames to encourage coherent concept integration through the sequence. We further introduce CASS, a CLIP-based metric that measures the output's directional alignment shift toward the reference and away from the source prompt. Using the denoiser's internal attention avoids an external localization model; in our A100 FP16 setup, MoCA-Video takes 3.2 seconds per output frame, excluding preprocessing. Across the evaluated baselines, MoCA-Video achieves the highest CASS, rel-CASS, and ImageReward, while LPIPS-T and FVD expose separate temporal-coherence and video-quality trade-offs.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Unified Long Video Inpainting and Outpainting via Overlapping High-Order Co-Denoising

    Nov 5, 2025Shuangquan Lyu, Jian Mao, Yue MaText-To-Video Diffusion ModelsLong Video Generation

  2. MiVE: Multiscale Vision-language features for reference-guided video Editing

    May 14, 2026Tong Wang, Meng Zou, Chengjing Wu +4Video EditingVision Encoders

  3. One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing

    Sep 3, 2026Adheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen +6Video EditingEdit