MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
Organizations: Department of Computer Science King Abdullah University of Science and Technology Thuwal, Saudi Arabia · Independent Researcher
Abstract
Unlike traditional video editing or inpainting, video semantic mixing fuses a reference concept with a moving target entity to produce a hybrid while preserving the source video's motion and layout. We propose MoCA-Video, a training-free framework that steers a frozen video-diffusion denoising trajectory through concept-localized reference injection. At selected low-noise steps, MoCA-Video uses concept attention to localize the target object and injects the reference latent into the localized region, where object structure has formed but appearance remains editable. A momentum-based correction carries the injected prediction across frames to encourage coherent concept integration through the sequence. We further introduce CASS, a CLIP-based metric that measures the output's directional alignment shift toward the reference and away from the source prompt. Using the denoiser's internal attention avoids an external localization model; in our A100 FP16 setup, MoCA-Video takes 3.2 seconds per output frame, excluding preprocessing. Across the evaluated baselines, MoCA-Video achieves the highest CASS, rel-CASS, and ImageReward, while LPIPS-T and FVD expose separate temporal-coherence and video-quality trade-offs.
Figures & tables
| Method | CASS | rel-CASS | LPIPS-T | FVD | ImageReward | |
| Directional alignment | Temporal | Video quality | ||||
| Pretrained | ||||||
| AnimateDiffV2V | 0.68 | -0.41 | 0.009 | 1266 | 0.246 | |
| TokenFlow PnP | 2.87 | 0.02 | 0.010 | 4580 | -2.087 | |
| TokenFlow SDEdit | 1.98 | 0.05 | 0.150 | 4417 | -1.431 | |
| Training-free | ||||||
| Variant | CASS | rel-CASS | LPIPS-T | FVD | ImageReward |
| MoCA-Video (ours) | 7.15 | 0.135 | 0.111 | 3520 | 0.263 |
| w/o injection | 4.74 | 0.090 | 0.127 | 3447 | 0.256 |
| w/o momentum | 5.98 | 0.108 | 0.119 | 3602 | 0.324 |
| w/o concept attention | 5.25 | 0.098 | 0.108 | 3670 | 0.324 |
| 0.5 | 0.8 ∗ | 1.0 | 1.5 | 2.0 | |
| CASS | 2.69 | 4.09 | 1.02 | 0.54 | 0.40 |
| 0.1 | 0.2 | 0.3 | 0.4 ∗ | 0.5 | |
| CASS | 0.89 | 1.60 | 1.56 | 4.09 | 2.37 |
| off | 0.5 | 1.0 | 2.0 ∗ | 5.0 | |
| CASS | 0.64 | 0.69 | 1.33 | 4.09 | 0.64 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | CLIP-BS | DINO-BS |
| FreeBlend | 6.65 | 0.27 |
| MoCA-Video | 4.00 | 0.12 |
| Method | Frames | Time (s/frame) |
| AnimateDiffV2V | 32 | 3.8 |
| FreeBlend + DynamiCrafter a | 16 | 3.3 |
| TokenFlow (PnP) | 16 | 9.4 |
| TokenFlow (SDEdit) | 16 | 1.9 |
| RAVE | 16 | 3.0 |
| MoCA-Video (ours) | 50 | 3.2 |
| Source prompt | Condition |
| A surfer riding a wave at sunset | panda |
| A superhero in a dynamic pose against a city skyline, cape flowing in the wind | deer |
| A majestic eagle soaring through the sky with its wings spread wide | corgi |
| A horse galloping across an open meadow | unicorn |
| A firefighter in full gear, standing in front of a burning building | bear |
| An astronaut floating in space against a starry background, wearing a detailed white spacesuit | cat |
| Source prompt | Condition |
| A goat walking on rocks | wolf |
| A man windsurfing with the sail | dog |
| A rolling soccer ball | airplane |
| A man wearing hockey skates | panda |
| A cow with a bell around its neck | corgi |