Pose-guided diffusion models can now synthesize entire human figures in motion, spawning a new class of deepfakes: Motion Aware Deepfake (MAD) that have already reached hundreds of millions of viewers. To better understand this emerging threat, we construct the first MAD-specific benchmark and measurement framework, containing over 1.5 million frames that mix 1,363 real and 30,122 synthetic videos from six controllable generators, with realistic perturbations and open-world evaluation splits. Then, we dissect MAD and discover that, despite their global coherence, these videos betray faint yet reliable cues: because the model relies on limited input frames for motion synthesis, it must predict and simulate coherent movement at motion boundaries, thereby producing high-frequency artifacts along with model-specific spectral fingerprints. Based on the observations obtained from analysis on dataset, we propose MoDA, the first defense framework tailored to detect and attribute MAD videos. MoDA couples spatial semantics with steganalysis-rich frequency features via cross-domain alignment and multi-scale aggregation, achieving 94.8% in-distribution and 89.1% cross-dataset detection accuracy gains of 10% to 25% over prior work and 91.5% model attribution accuracy. MoDA achieves 81.94% accuracy on 200 clips produced by two unseen commercial MAD platforms, indicating promising zero-shot transfer, and 78.13% detection accuracy on 1,200 unseen MAD video clips (55k frames in total) collected from the open Internet. Under white-box, gray-box, and black-box adaptive attacks, MoDA maintains relatively stable detection and attribution performance while the accuracies of the baselines drop rapidly.
Figures & tables
Figure 1. Overview of MAD dataset.
Dataset Partition
Source (Model)
Input Format
Ref. Image Num
Motion Sources
Total Videos
Total Frames
Output Size
Real Videos
Real
/
/
1,363
1,363
65k
604 × 1080
Synthetic Videos
Mimicmotion ( Zhang et al., 2024 )
DWpose
40
340
9,350
448k
576 × 1024
MusePose ( Tong et al., 2024 )
DWpose
40
340
9,350
448k
384 × 768
Animate Anybody ( Hu, 2024 )
DWPose
40
340
9,350
448k
512 × 784
Disco ( Wang et al., 2023a )
OpenPose
16
42
994
48k
256 × 256
MagicAnimate ( Xu et al., 2024 )
DensePose
18
42
844
40k
512 × 512
Table 1. Details of MAD dataset. Total Videos denote 48-frame clips after filtering failed generations and low-quality outputs; counts therefore do not necessarily equal the Cartesian product of reference identities and motion sources.
Figure 2. Multi-view evidence: understanding of MAD.
Model
T (Period)
H (Entropy)
Elow/Ehigh (Ratio)
Animate Anyone
Small
Low
Low
MimicMotion
Large
High
High
MusePose
Small
Medium-High
Highest
Table 4
Figure 3. Model-specific fingerprints in frequency domain, demonstrating clear separability rooted in architecture.
Figure 4. System overview of MoDA
Figure 5. Visual examples of Key-Region Focus Module. MoDA assigns higher attention weights to regions with salient motion boundaries.
Module
Filter
Alignment
Attention
Aggregation
End-to-end
Time (s)
0.026
0.062
0.081
0.014
0.183
Table 2. Runtime breakdown of MoDA (per 24 frames).
Category
Method
Original MAD-Dataset
Robustness under Post-processing
ACC
AUC
FPR
FNR
F1-Score
FS
LIC
PN
VE
Avg
Deepfake Detection
RECCE ( Cao et al., 2022 )
0.917
0.873
0.102
0.079
0.910
0.858
0.738
0.824
0.741
0.816
DFGaze ( Xu et al., 2023 )
0.876
0.822
0.113
0.135
0.873
0.816
0.762
0.754
0.704
0.783
XCeption ( Yang et al., 2023b )
0.886
0.834
0.085
0.131
0.889
0.956
0.828
0.509
0.724
0.781
M2TR ( Wang et al., 2022 )
0.813
0.751
0.175
0.198
0.811
0.813
0.803
0.682
0.758
0.774
FTCN ( Tan et al., 2024 )
0.671
0.599
0.346
0.312
0.669
0.624
0.592
0.467
0.517
0.574
Table 3. In-dataset evaluation results on MAD.
Training set
Method
Mimic
Muse
Anim
Avg
Mimic
RECCE ( Cao et al., 2022 )
0.968
0.864
0.771
0.868
DFGaze ( Peng et al., 2024 )
0.897
0.731
0.256
0.628
XCeption ( Rossler et al., 2019 )
0.733
0.664
0.570
0.656
DeFake ( Sha et al., 2023 )
0.525
0.328
0.084
0.312
MoDA
0.962
0.892
0.775
0.876
MusePose
RECCE ( Cao et al., 2022 )
0.774
0.956
0.449
0.728
Table 4. Cross-dataset evaluation results on MAD (recall on the MAD class).
Method
In-distribution
Cross-dataset
ACC
AUC
ACC
AUC
Spatial
94.51
0.779
73.06
0.642
Frequency
91.63
0.688
84.53
0.752
Dual-Stream
95.80
0.756
82.16
0.724
Dual-Att-Stream
96.45
0.854
85.44
0.739
CDFA
94.91
0.764
83.32
0.758
Table 5. Ablation study of different model components.
Method
Mimic
MusePose
Animate
Avg
RECCE ( Cao et al., 2022 )
0.862
0.137
0.441
0.480
DFGaze ( Peng et al., 2024 )
0.896
0.707
0.920
0.841
XCeption ( Rossler et al., 2019 )
0.741
0.187
0.829
0.586
DeFake ( Sha et al., 2023 )
0.452
0.081
0.361
0.298
MoDA
0.997
0.858
0.890
0.915
Table 6. Model tracing evaluation results on MAD (ACC).
Variant
AUC
FPR
Memory
Latency
Full MoDA (platform)
0.969
0.023
1.6GB
7.6 ms/frame
Lightweight (mobile)
0.756
0.067
164MB
2.0 ms/frame
Table 7. Full MoDA vs. lightweight variant.
Figure 6. Examples from commercial MAD platforms.
Figure 7. Detection under adaptive attacks and perturbation.
Defense
White-box PGD
Gray-box MI-FGSM
None
0.394
0.554
Adversarial training
0.523 (+18.2%)
0.703 (+14.9%)
Input preprocessing
0.417 (+3.3%)
0.619 (+6.5%)
Table 8. Mitigation results against adaptive attacks (ACC).
Audio-visual deepfakes have reached a level of realism that makes perceptual detection unreliable, threatening media integrity and biometric security. While multimodal detection has shown promise, most approaches are binary classification tasks that often latch onto dataset-specific artifacts rather than genuine generative traces. We argue that a detector incapable of identifying how a video was forged is likely learning the wrong signal. Unlike binary detection, attribution-guided learning imposes a stronger geometric constraint on the shared embedding space, forcing the model to encode generator-specific forensic content rather than shortcuts. We propose the Attribution-Guided Multimodal Deepfake Detection (AMDD) framework, which jointly learns to detect and attribute manipulation. AMDD treats generator attribution as a structured regularization that constrains representation geometry toward forensically meaningful features. We introduce a Cross-Modal Forensic Fingerprint Consistency (CMFFC) loss to enforce alignment between generator-induced artifacts in visual and audio streams. This exploits the fact that coherent manipulation leaves correlated traces across modalities, grounded in the physical coupling between speech and facial articulation that synthetic pipelines routinely disrupt. Architecturally, we pair a ResNet50 with temporal attention for visual encoding against a pretrained ResNet18 for mel spectrograms, closing the encoder capacity gap found in prior models. On FakeAVCeleb, AMDD achieves 99.7% balanced accuracy and 99.8% AUC with 95.9% attribution accuracy. Cross-dataset evaluation on DeepfakeTIMIT, DFDM, and LAV-DF confirms that real video detection generalizes robustly, while fake detection on unseen generators remains an open challenge that we analyze in depth.
Wasim Ahmad, Wei Zhang, Xuerui Mao
School of Interdisciplinary Science, Beijing Institute of Technology, Beijing 100081, China
The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmarks provide limited coverage of recent synthesis methods and generally lack reliable fine-grained textual annotations. Meanwhile, conventional detectors and multimodal large language models (MLLMs), whether operating as a single model or relying on a single analytical perspective, often fail to capture subtle forgery artifacts, limiting their generalization to emerging AI-generated methods. To address these limitations, we introduce FaceVid-Forensics-100K, a large-scale deepfake video dataset comprising 100,000 videos and spanning 33 synthesis methods across face swapping, face reenactment, and entire-face synthesis, including recent generators such as Seedance 2.0. The dataset provides fine-grained textual annotations of visual observations and verdict-consistent forensic explanations, automatically synthesized through a multi-model aggregation and conflict-resolution pipeline powered by advanced MLLMs. Building on this benchmark, we propose a multi-agent forensic reasoning framework that employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives: texture, lighting, motion, and physics. A judge agent then reconciles their reports to produce a final prediction together with an explanation. Extensive evaluations on out-of-domain test sets show that, despite being composed entirely of small open-source MLLMs, our framework outperforms all methods including closed-source GPT and Gemini models and ranks first across all reported metrics on this benchmark. The project page is available at https://xavierjiezou.github.io/ARGUS/.
Xuechao Zou, Shun Zhang, Kai Li +6
Beijing Jiaotong University · Tsinghua University · Ant Group
We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by recent text-to-video and image-to-video models. The videos capture single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interviews - 271 real and 2,420 synthetic videos generated by seven modern video models. The study on DF26 shows that human performance in detecting AI-generated videos, as well as state-of-the-art deepfake detectors, is close to random chance. Our results highlight the limitations of current evaluation protocols and motivate the need for benchmarks that explicitly measure robustness to modern generative model distribution shifts.
Severyn Shykula, Andrii Yermakov, Ivan Samarskyi +3
Ukrainian Catholic University · Faculty of Electrical Engineering, Czech Technical University in Prague · Hover Inc., USA +1