Video deepfakes targeting a specific individual, the Person-of-Interest (POI), are the most harmful ones, and, since a public figure is abundantly recorded, a detector can be built from genuine footage of that individual. Such detectors commonly describe a subject through a 3D Morphable Model (3DMM) and adopt its coefficients as a whole, so which part of that description carries the signal has never been measured. We dissect it, holding the encoder, the training corpus and the enrollment protocol fixed and varying only what the encoder observes. The groups of coefficients prove largely redundant, since the shape block alone recovers almost all the accuracy of the full vector, and their temporal evolution contributes a real but bounded amount. We further show that the dense surface the same fit returns, which these detectors discard, carries identity information that the coefficients do not, and that it helps precisely where they are weakest. We assemble the best configuration into MOTIF, a visual-only detector trained on real videos only, with no manipulated video and no POI-specific data. It improves on both state-of-the-art POI detectors in every dataset and manipulation of our benchmark and at two quality levels. Our experimental code will be released at https://github.com/polimi-ispl/MOTIF.
Figures & tables
Fig. 1 : MOTIF training overview.
Category
Dataset
Generation method
Face-swap (FS)
DF-TIMIT
autoencoder-based swap
FakeAVCeleb
FaceSwap, FSGAN
DeepSpeak
FaceFusion, INSwapper, SimSwap
KoDF
FaceSwap, DeepFaceLab, FSGAN
Face-swap + lip-sync (FS+LS)
FakeAVCeleb
FaceSwap and FSGAN followed by Wav2Lip [ 18 ]
Reenactment (RE)
DeepSpeak
LivePortrait, HelloMeme, Memo
TABLE I : Generation methods of each evaluation dataset, grouped into the four canonical categories used in the per-category results.
FakeAVCeleb
DeepSpeak
KoDF
Overall
Input
FS
FS+LS
LS
All
FS
RE
LS
All
FS
RE
LS
All
Mean
Shape (S)
91.6/92.2
93.1/93.5
60.7/78.3
85.3/87.0
98.5/98.4
87.6/90.9
86.7/91.5
89.7/91.3
97.6/97.4
85.7/86.8
93.0/93.5
93.1/92.0
89.4/90.1
Expression (S)
89.2/90.4
91.1/91.7
57.0/75.7
82.9/85.1
98.3/98.5
88.2/91.4
87.2/91.3
90.0/91.2
97.8/97.7
85.8/87.0
94.0/94.2
93.5/92.3
88.8/89.6
Shape + expr. (S)
90.5/91.6
91.8/92.9
59.0/78.3
83.9/86.3
98.3/98.3
87.9/90.8
86.3/91.6
89.6/91.1
97.6/97.6
86.6/87.5
96.2/95.5
93.7/92.6
89.0/90.0
Shape (D)
92.8/93.4
93.3/94.1
59.2/77.7
85.6/87.6
99.7/99.6
90.7/92.6
89.3/92.5
92.2/92.5
96.9/96.8
86.5/87.4
95.5/95.4
93.4/92.4
90.4/90.8
Expression (D)
90.2/91.0
91.7/92.8
57.9/77.0
83.8/85.6
99.6/99.7
90.2/92.8
88.9/92.6
91.8/92.8
97.5/97.3
86.4/87.3
95.5/95.9
93.6/92.5
89.7/90.3
TABLE II : AUC / BA (%) per dataset and manipulation category, considering static (S) coefficients or dynamic (D) ones. Best result per column in bold.
DF-TIMIT
FakeAVCeleb
DeepSpeak
KoDF
Overall
Config
FS
FS
FS+LS
LS
All
FS
RE
LS
All
FS
RE
LS
All
Mean
G-only
100.0/100.0
92.6/93.7
95.0/95.7
58.3/77.1
86.2/87.9
99.5/99.6
90.5/92.9
88.0/91.9
91.6/92.3
97.5/97.4
86.8/86.7
96.4/95.8
94.0/92.6
92.9/93.2
M-only
92.1/92.8
76.9/82.2
82.4/85.6
68.2/79.0
77.6/81.5
88.3/91.5
81.0/84.9
84.3/88.3
84.0/86.3
85.0/86.5
81.7/84.6
90.8/91.4
85.2/85.6
84.7/86.6
Fusion
100.0/100.0
91.0/92.3
93.3/94.0
67.4/79.7
86.8/87.7
99.1/99.3
88.5/91.6
87.5/91.8
90.5/91.6
97.1/97.1
90.4/90.6
98.0/97.4
95.3/94.0
93.1/93.3
TABLE III : AUC / BA (%) per dataset and manipulation category, considering the global branch only (G-only), the mouth branch only (M-only) or their fusion. Best result per column is highlighted in bold.
DF-TIMIT
FakeAVCeleb
DeepSpeak
KoDF
Overall
Method
FS
FS
FS+LS
LS
All
FS
RE
LS
All
FS
RE
LS
All
Mean
High quality
RealForensics [ 19 ]
100.0/100.0
98.2/98.8
95.7/94.3
75.7/78.9
86.9/84.9
86.5/90.0
84.9/84.0
86.1/86.7
85.5/83.5
87.1/85.5
94.5/91.4
99.2/98.1
93.1/89.9
91.4/89.6
LipForensics [ 20 ]
97.9/98.1
96.5/96.2
96.1/94.4
95.7/94.3
96.3/93.8
84.1/89.5
93.7/93.2
97.8/97.4
93.6/92.6
88.7/88.8
91.2/88.5
99.6/99.1
93.1/90.8
95.2/93.8
FTCN [ 21 ]
100.0/100.0
80.5/84.9
86.8/87.4
80.8/84.8
82.6/82.7
88.7/92.2
73.4/80.5
83.2/87.0
79.4/83.3
90.3/89.8
92.0/88.5
98.2/97.5
93.5/90.8
88.9/89.2
Seferbekov [ 23 ]
90.0/92.5
94.3/93.5
98.0/97.1
81.1/88.1
89.2/89.0
57.2/74.8
59.3/69.6
92.9/92.9
70.0/75.7
92.9/94.4
81.7/81.0
98.4/97.9
90.9/89.6
85.0/86.7
TABLE IV : State-of-the-art comparison per manipulation ( AUC / BA , %). Bold : best POI -specific method; italic : best general deepfake detector.
Deepfakes targeting a high-profile individual, known as Person-of-Interest (POI), are a threat to modern democracies and societies. Current POI deepfake detection methods still struggle to combine robustness to post-processing, efficiency and interpretability, key aspects of modern deepfake detectors. In this paper we propose CUPID, a POI video deepfake detector that combines UV texture maps, a facial appearance representation derived from 3D face reconstructions, with the representation learning capabilities of the Masked Autoencoder (MAE). Our method does not require any deepfake videos in its training phase. Moreover, it does not even require including a specific POI in the training set: the combination of UV texture maps extracted from real video frames and the MAE context-guided reconstruction yields a latent space that captures rich and discriminative facial features even for identities unseen during training. In the testing phase, the embeddings extracted from a query video depicting the POI can be matched against pristine reference videos to assess the video authenticity. Furthermore, operating in the UV space naturally provides an additional layer of interpretability. Specifically, we can extract decoded residual maps that highlight which facial regions of a test video deviate most from the identity representation of the corresponding POI. Experiments on four deepfake datasets show that CUPID outperforms the current state of the art on most datasets and achieves the best overall robustness against strong downscaling and compression, while also providing substantially faster inference. Our experimental code will be released at https://github.com/polimi-ispl/CUPID.
Giovanni Affatato, Sara Mandelli, Edoardo Daniele Cannas +2
Dipartimento di Elettronica, Informazione e Bioingegneria (DEIB), Politecnico di Milano, 20133 Milan, Italy
Deepfakes no longer need to fake a whole video. Generators that read the transcript now alter only the few seconds in which a video's meaning turns, so a forgery hides in a small, unknown fraction of the video. Yet detectors still read every one-second window of both the audio and image streams, spending nearly all of their compute where nothing was altered. We observe that deciding where to look is far cheaper than looking. We present ModalFidelity, a lightweight router that previews each window and decides, before any forensic detector runs, which stream is worth reading, under a hard compute budget it can never exceed. On AV-Deepfake1M, reading at most a fifth of the windows, it is more accurate than gating after the detectors at 15.9x less compute, and retains over 96% of the accuracy of an oracle that knows where every forgery lies.
Oguzhan Baser, Kaan Kale, Sriram Vishwanath +1
The University of Texas at Austin, Austin, TX, USA · Georgia Institute of Technology, Atlanta, GA, USA
Deepfake detectors show large performance gaps across demographic groups. Existing fairness approaches require demographic labels, retraining, or sacrifice accuracy. We introduce Face-Fairness (FF), a plug-and-play framework for bias mitigation. Our primary contribution, Face-Feature Tuning (FFT), is the first demographic label-free fairness method demonstrated for deepfake detection: a lightweight calibrator that performs a logit remapping conditioned on frozen face embeddings. We complement FFT with two variants: FF-Max, which maximizes worst-group accuracy when demographics are available, and FF-Discover, which does the same with embedding-discovered groups. Across in-domain and cross-dataset test settings, FF consistently reduces FPR/TPR gaps and improves minimum group accuracy while maintaining (often improving) overall accuracy. The approach is detector-agnostic, adds negligible runtime overhead, and requires no access to identity attributes.