cs.CVOct 8, 2026

Stop My Dancing! Understanding, Detecting and Attributing Motion-Aware Deepfake Videos

Authors: Fazhong Liu, Yan Meng, Tian Dong, Guoxing Chen, Haojin Zhu

Organizations: Shanghai Jiao Tong University Shanghai, China

Abstract

Pose-guided diffusion models can now synthesize entire human figures in motion, spawning a new class of deepfakes: Motion Aware Deepfake (MAD) that have already reached hundreds of millions of viewers. To better understand this emerging threat, we construct the first MAD-specific benchmark and measurement framework, containing over 1.5 million frames that mix 1,363 real and 30,122 synthetic videos from six controllable generators, with realistic perturbations and open-world evaluation splits. Then, we dissect MAD and discover that, despite their global coherence, these videos betray faint yet reliable cues: because the model relies on limited input frames for motion synthesis, it must predict and simulate coherent movement at motion boundaries, thereby producing high-frequency artifacts along with model-specific spectral fingerprints. Based on the observations obtained from analysis on dataset, we propose MoDA, the first defense framework tailored to detect and attribute MAD videos. MoDA couples spatial semantics with steganalysis-rich frequency features via cross-domain alignment and multi-scale aggregation, achieving 94.8% in-distribution and 89.1% cross-dataset detection accuracy gains of 10% to 25% over prior work and 91.5% model attribution accuracy. MoDA achieves 81.94% accuracy on 200 clips produced by two unseen commercial MAD platforms, indicating promising zero-shot transfer, and 78.13% detection accuracy on 1,200 unseen MAD video clips (55k frames in total) collected from the open Internet. Under white-box, gray-box, and black-box adaptive attacks, MoDA maintains relatively stable detection and attribution performance while the accuracies of the baselines drop rapidly.

Figures & tables

Explore similar work

CardsList
  1. Attribution-Guided Multimodal Deepfake Detection via Cross-Modal Forensic Fingerprints

    Apr 29, 2026Wasim Ahmad, Wei Zhang, Xuerui MaoAudio Deepfake DetectionCross-Modal Representation Learning

  2. Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

    Aug 7, 2026Xuechao Zou, Shun Zhang, Kai Li +6Explainable Deepfake DetectionDeepfake Detection

  3. DF26: We Cannot Tell Fake From Real Anymore

    Sep 7, 2026Severyn Shykula, Andrii Yermakov, Ivan Samarskyi +3Text-to-Video GenerationOOD Generalization