This paper presents a systematic study of scaling laws for the deepfake detection task. Specifically, we analyze the model performance against the number of real image domains, deepfake generation methods, and training images. Since no existing dataset meets the scale requirements for this research, we construct ScaleDF, the largest dataset to date in this field, which contains over 5.8 million real images from 51 different datasets (domains) and more than 8.8 million fake images generated by 102 deepfake methods. Using ScaleDF, we observe power-law scaling similar to that shown in large language models (LLMs). Specifically, the average detection error follows a predictable power-law decay as either the number of real domains or the number of deepfake methods increases. This key observation not only allows us to forecast the number of additional real domains or deepfake methods required to reach a target performance, but also inspires us to counter the evolving deepfake technology in a data-centric manner. Beyond this, we examine the role of pre-training and data augmentations in deepfake detection under scaling, as well as the limitations of scaling itself.The ScaleDF dataset is available at https://huggingface.co/datasets/WenhaoWang/ScaleDF.
Figures & tables
Figure 1: Illustration of ScaleDF, which is the largest deepfake detection dataset across four dimensions: number of real domains, number of deepfake methods, number of videos, and number of images. To ensure a fair comparison, for datasets that do not explicitly provide the number of images, we estimate image counts by sampling one frame per second from the videos. Using this dataset, we present a systematic study on scaling laws for deepfake detection and reveal several insights.
Figure 2: Statistics of heights and widths of all real faces in the ScaleDF dataset.
Real
Deepfake
Videos
Images
Dataset
Domains
Methods
Real
Fake
Real
Fake
DF-TIMIT ( Korshunov and Marcel, 2018 )
1
2
320
640
-
-
UADFV ( Yang et al., 2019 )
1
1
49
49
-
-
FaceForensics++ ( Rossler et al., 2019 )
1
4
1,000
4,000
-
-
DeepFakeDetection ( Google, 2019 )
1
5
363
3,068
-
-
Celeb-DF V2 ( Li et al., 2020 )
1
1
590
5,639
-
-
Table 1: Compared to existing datasets, ScaleDF includes more real domains, more deepfake methods, and a larger number of videos and images, enabling scaling law research in these dimensions.
Figure 3: Left: observed power-law scaling as the number of training real domains changes; Right: observed power-law scaling as the number of training deepfake methods changes. μ represents the mean computed over 10 repetitions and 7 test datasets, while σ denotes the variance across the 10 repetitions.
Figure 4: Left: observed double-saturating power-law scaling as the number of training images changes; μ represents the mean computed over 5 repetitions and 7 test datasets, while σ denotes the variance across the 5 repetitions. Right: performance changes with respect to model sizes.
AUC
DFD
CDF V2
Wild
Forgery.
DFF
DF40
ScaleDF
Mean
ImageNet
0.793
0.915
0.815
0.824
0.909
0.980
0.968
0.886
CLIP
0.795
0.894
0.802
0.848
0.954
0.979
0.981
0.893
SigLIP 2
0.750
0.890
0.785
0.855
0.951
0.986
0.983
0.886
Table 2: Comparison of using different pre-training models: similar performance observed.
AUC
DFD
CDF V2
Wild
Forgery.
DFF
DF40
ScaleDF
Mean
N/A
0.667
0.805
0.781
0.756
0.808
0.937
0.975
0.818
QC
0.774
0.914
0.802
0.770
0.932
0.978
0.976
0.878
QC + PT
0.793
0.915
0.815
0.824
0.909
0.980
0.968
0.886
Table 3: Effectiveness of image quality compression (QC) and perturbations (PT).
Training sets
AUC
DFD
CDF V2
Wild
Forgery.
DFF
DF40
ScaleDF
DFD
−
0.802
0.738
0.610
0.545
0.594
0.524
CDF V2
0.709
−
0.777
0.640
0.599
0.665
0.612
Wild
0.643
0.757
−
0.555
0.489
0.562
0.556
Forgery.
0.813
0.913
0.821
−
0.687
0.833
0.657
DFF
0.583
0.568
0.552
0.578
−
0.669
0.622
DF40
0.587
0.800
0.684
0.677
0.607
−
0.667
Table 4: Comparison in cross-benchmark setting: with scaling, we achieve the best performance.
AUC
DFD
CDF V2
Wild
Forgery.
DFF
DF40
ScaleDF
Mean
w/o FF++
0.758
0.868
0.796
0.807
0.905
0.978
0.969
0.869
w FF++
0.793
0.915
0.815
0.824
0.909
0.980
0.968
0.886
Table 5: Comparison of whether includes FaceForensics++ ( Rössler et al., 2019 ) in the ScaleDF.
No.
Dataset
Format
Vol.
Link
0
GRID ( Cooke et al., 2006 )
Video
16 K+
Link
1
MORPH - 2 ( Ricanek and Tesafaye, 2006 )
Image
49 K+
Link
2
LFW ( Huang et al., 2008 )
Image
13 K+
Link
3
Multi-PIE ( Gross et al., 2008 )
Image
0.1 M+
Link
4
GENKI - 4K ( MPLab, 2009 )
Image
3.8 K+
Link
5
YouTubeFaces ( Wolf et al., 2011 )
Image
0.2 M+
Link
Table 6: Real datasets included in the ScaleDF, with the testing ones highlighted in .
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Left Top: observed power-law scaling as the number of training real domains changes; Right Top: observed power-law scaling as the number of training deepfake methods changes; Left Bottom: observed double-saturating power-law scaling as the number of training images changes; Right Bottom: performance changes with respect to model sizes. μ represents the mean computed over repetitions and test datasets, while σ denotes the variance across repetitions.
Figure 6: Used original face.
EER
DFD
CDF V2
Wild
Forgery.
DFF
DF40
ScaleDF
Mean
ImageNet
0.281
0.162
0.268
0.260
0.161
0.063
0.093
0.184
CLIP
0.278
0.179
0.279
0.231
0.085
0.065
0.052
0.167
SigLIP 2
0.316
0.193
0.290
0.229
0.094
0.052
0.052
0.175
Appendix
Table 7: Comparison of using different pre-training models: similar performance observed.
Figure 7: Perceived demographic distribution of faces in ScaleDF.
No.
Method
Cat.
Arch.
Link
0
Faceswap ( Earl, 2015 )
FS
Affine
Code
1
FaceSwap ( Kowalski, 2016 )
FS
3D
Code
2
DeepFakes ( deepfakes, 2017 )
FS
VAE
Code
3
FSGAN ( Nirkin et al., 2019 )
FS
GAN
Code
4
SimSwap ( Chen et al., 2020 )
FS
GAN
Code
5
HifiFace ( Wang et al., 2021c )
FS
3D
Code
Appendix
Table 11: Deepfake methods included in the ScaleDF, with the testing ones highlighted in .
AUC
DFD
CDF V2
Wild
Forgery.
DFF
DF40
ScaleDF
Mean
46
0.793
0.915
0.815
0.824
0.909
0.980
0.968
0.886
32
0.752
0.845
0.716
0.755
0.881
0.927
0.966
0.835
32
0.764
0.917
0.747
0.818
0.907
0.985
0.968
0.872
32
0.779
0.904
0.727
0.817
0.920
0.981
0.967
0.871
32
0.765
0.897
0.817
0.817
0.865
0.977
0.960
0.871
32
0.760
0.904
0.732
0.811
0.863
0.980
0.960
0.859
Appendix
Table 12: Full experimental results measured by AUC for scaling law with varying numbers of real domains . The concluded scaling law is shown in Fig. 3 (Left).
AUC
DFD
CDF V2
Wild
Forgery.
DFF
DF40
ScaleDF
Mean
88
0.793
0.915
0.815
0.824
0.909
0.980
0.968
0.886
64
0.774
0.914
0.776
0.815
0.874
0.982
0.965
0.871
64
0.756
0.912
0.802
0.820
0.893
0.978
0.955
0.874
64
0.790
0.917
0.802
0.836
0.888
0.981
0.963
0.882
64
0.782
0.915
0.796
0.823
0.888
0.977
0.963
0.878
64
0.766
0.908
0.806
0.802
0.894
0.978
0.965
0.874
Appendix
Table 13: Full experimental results measured by AUC for scaling law with varying numbers of deepfake methods . The concluded scaling law is shown in Fig. 3 (Right).
Table 18
EER
DFD
CDF V2
Wild
Forgery.
DFF
DF40
ScaleDF
Mean
46
0.281
0.162
0.268
0.260
0.161
0.063
0.093
0.184
32
0.317
0.224
0.337
0.310
0.189
0.149
0.092
0.231
32
0.308
0.162
0.313
0.264
0.162
0.055
0.092
0.194
32
0.293
0.173
0.326
0.265
0.148
0.063
0.093
0.194
32
0.307
0.179
0.257
0.269
0.209
0.070
0.105
0.199
32
0.308
0.173
0.334
0.270
0.210
0.064
0.104
0.209
Appendix
Table 16: Full experimental results measured by EER for scaling law with varying numbers of real domains . The concluded scaling law is shown in Fig. 5 (Left Top).
EER
DFD
CDF V2
Wild
Forgery.
DFF
DF40
ScaleDF
Mean
88
0.281
0.162
0.268
0.260
0.161
0.063
0.093
0.184
64
0.299
0.165
0.301
0.268
0.200
0.061
0.096
0.199
64
0.312
0.167
0.279
0.264
0.175
0.066
0.112
0.196
64
0.283
0.162
0.277
0.249
0.186
0.064
0.099
0.189
64
0.293
0.162
0.272
0.261
0.183
0.067
0.100
0.191
64
0.304
0.172
0.278
0.278
0.178
0.068
0.096
0.196
Appendix
Table 17: Full experimental results measured by EER for scaling law with varying numbers of deepfake methods . The concluded scaling law is shown in Fig. 5 (Right Top).
Table 21
Table 20: Visualization of processed faces in ScaleDF.
Table 21: Visualization of processed faces in ScaleDF.
Table 22: Visualization of processed faces in ScaleDF.
Table 23: Visualization of processed faces in ScaleDF.
Table 24: Visualization of processed faces in ScaleDF.
Table 25: Visualization of processed faces in ScaleDF.
Table 26: Visualization of processed faces in ScaleDF.
Table 27: Visualization of processed faces in ScaleDF.
Table 28: Visualization of processed faces in ScaleDF.
Table 29: Visualization of processed faces in ScaleDF.
Table 30: Visualization of processed faces in ScaleDF.
Table 31: Visualization of processed faces in ScaleDF.
Table 32: Visualization of processed faces in ScaleDF.
Table 33: Visualization of processed faces in ScaleDF.
Table 34: Visualization of processed faces in ScaleDF.
Table 35: Visualization of processed faces in ScaleDF.
Table 36: Demonstration of the perturbations used when training models on the ScaleDF dataset.
Table 37: Demonstration of the perturbations used when training models on the ScaleDF dataset.
Table 38: Demonstration of the perturbations used when training models on the ScaleDF dataset.
Department of Computer Science, University of Bucharest, Romania · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE · Department of Computer Science, University of Central Florida, Orlando, US
da/sec – Biometrics and Security Research Group, Hochschule Darmstadt · Computer Science and Engineering Department, Michigan State University · STARS team, Inria Center at Université Côte d’Azur