Generative models can synthesize high-quality inauthentic multimedia content that is already being misused at scale. We evaluate twenty deepfake detectors against ten generators released in the last four years and find accuracy decreasing over time, from near-perfect 99.5% to 76%. Adversarial perturbations further reduce every baseline detector to below 2% accuracy, effectively inverting the detector's assigned label. We argue that this unreliability reflects a fundamental ambiguity: generators can reproduce authentic content exactly (e.g., through memorization), so content alone cannot reveal the true provenance label.For this reason, content produced by a generator must admit a faithful reconstruction by that same generator, and finding such a reconstruction makes synthetic provenance plausible and authenticity plausibly deniable.We therefore propose and evaluate a detection paradigm that outputs a calibrated prediction of whether authenticity is plausibly deniable: a faithful reconstruction by any known generator establishes plausible deniability, while calibration bounds how often content from known generators fails to be reproduced. Our evaluation shows that (i) our detector can be calibrated so that at most 1% of generated content is wrongly certified, an operating point at which most baseline detectors reach near-zero recall, including the strongest with 93% accuracy; (ii) calibrating a stricter security threshold on attacked samples preserves this bound against adaptive adversaries within the evaluated bounded-perturbation attack space, whose perturbations break every baseline, but does not cover arbitrary adversarial transformations; and (iii) post-hoc verifiability is eroding, as 1,116 of 3,000 Reddit images resist reproduction by a 2022 generator, but only 55 to 79 resist reproduction by 2024 generators.
Figures & tables
Fig. 1 : The detection game increasingly favors generators. Among detectors available at each generator’s release, the best accuracy falls from 99.5% on 2022 generators to 76% on 2026 generators, suggesting that generators are winning.
Fig. 2 : Conceptual illustration of our method. (A) Traditional post-hoc detectors separate real from fake using feature cues but struggle as generators improve. (B) Our method instead tests whether a generator can resynthesize the query image x . The similarity s(x,x~) between x and its resynthesis x~ is calibrated into an authenticity score, where high similarity implies plausible deniability and low similarity indicates likely authentic . In this and subsequent figures, A-index denotes the authenticity score.
Fig. 3 : Computing of the Authenticity Score. Given an input image x , a reconstruction-free inverter Ge−1 produces an inverted reconstruction x~ . We then compute complementary similarities between x and x~ : pixel fidelity (PSNR) , structural fidelity (SSIM) , perceptual distance ( 1− LPIPS) , and semantic consistency ( CLIP cosine). A calibrated weighted combiner (learned α1,α2,α3,α4 ) produces a scalar s ( x , x~ ), yielding the A-index in [0,1] . A safety threshold τ certifies content as Authentic when A-index≥τ , and otherwise labels it as Plausibly Deniable . We further analyze robustness by applying ℓ∞ -bounded perturbations δ (PGD-style) through Ge−1 to maximize or minimize the A-index.
Fig. 4 : Reconstruction fidelity separates real and fake images. Distribution of the A-index a(x,x~) obtained by inverting both real and fake images with SD3 Medium . Fake images are concentrated at lower scores because they are reconstructed more faithfully, whereas real images are shifted toward higher scores. This separation provides the signal used to calibrate the generator-specific authentication threshold.
Fig. 5 : Examples of easy- and hard-to-invert images. Panels (a), (c), (e), and (g) are real inputs, while panels (b), (d), (f), and (h) are their generated reconstructions. The top row shows easy-to-invert images, and the bottom row shows hard-to-invert images. Visual complexity generally makes an image harder to invert because fine objects, overlapping structures, and spatial relationships are difficult to reproduce, as shown in (e)–(h). However, the visually complex input in (c) is reconstructed faithfully in (d) because its content is well represented within the generator’s latent space. Inversion difficulty therefore depends on both image complexity and how well the generator’s latent space represents the image.
Fig. 6 : C2PClip score distributions overlap under zero-shot distribution shift. Predictions on a balanced test set of 1,000 real and 1,000 fake images from unseen generators. C2PClip correctly classifies most real images but assigns most fake images to the real class. The resulting overlap prevents the detector from retaining useful recall when its threshold is calibrated to limit false authenticity claims.
Before attack
After attack
Attack success (%)
Model
Correct fake
Correct real
Acc. (%)
Correct fake
Correct real
Acc. (%)
on fake
on real
UFD [ 13 ]
1
974
48.75
0
0
0.00
100.0
100.0
FreqNet [ 42 ]
53
995
52.40
0
0
0.00
100.0
100.0
NPR [ 41 ]
86
653
36.95
0
0
0.00
100.0
100.0
FatFormer [ 45 ]
1
987
49.40
0
0
0.00
100.0
100.0
AEROBLADE † [ 31 ]
778
778
77.80
0
30
1.50
100.0
96.1
TABLE I : Results before/after PGD ( ϵ=2558 ) on 2,000 Images (1,000 fake + 1,000 real, 512×512 ). Following standard practice, only samples classified correctly before the attack are perturbed, so accuracy can only drop. Attack success is the fraction of those correct samples flipped. † : threshold-degenerate detector (score has no natural 0.5 cut). “ − ”: no sample of that class was correct pre-attack.
Fig. 7 : PGD perturbations collapse binary detectors but preserve separated A-index distributions. Panels (a)–(c) show an image before and after an ℓ∞ -bounded perturbation with ϵ=8/255 , and the perturbation δ amplified 16× for visibility. Panel (d) shows that D3 declines from 83.90% to 1.75% accuracy, the highest post-attack accuracy among the twenty baselines. Panel (e) shows that PGD shifts both A-index distributions, while the distributions retain distinct peaks and motivate the security threshold.
Fig. 8 : Best-of- 100 search does not reach the calibrated thresholds. The attacker samples 100 candidates from one prompt and selects the highest A-index before PGD refinement. The selected candidate increases from 0.0148 to 0.0154 after PGD, remaining below the safety and security thresholds.
Fig. 9 : Calibrated thresholds retain different subsets of 3,000 unverified internet images. We apply five generator-specific thresholds to the same SD3 inversion-score distribution: 1,116 images exceed the SD2.1 threshold, compared with 55 - 79 for the four newer configurations. This comparison measures threshold sensitivity on one score distribution and does not compare generator-specific reconstructions of the corpus. Examples from the corpus are shown in Figure 15 .
Fig. 10 : Frame-aggregated video authentication. We sample eight frames at 30 -frame intervals, reconstruct and score each frame independently, and average their A-index values to obtain avideo . Lower values indicate that the sampled frames are easier to reconstruct, while higher values provide stronger authentication evidence.
Model
AUC
Prec.
Rec.
F1
GenConViT [ 47 ]
0.615
0.59
0.49
0.53
FTCN [ 46 ]
0.483
0.50
0.64
0.40
StyleFlow [ 48 ]
0.509
0.53
0.42
0.47
TABLE II : Video detectors transfer poorly to Deepfake-Eval-2024. Results on 50 authentic and 50 generated videos; no baseline exceeds 0.615 AUC or 0.59 precision.
Fig. 11 : Frame-aggregated A-index scores retain the expected ordering on video. Each score averages eight independently processed frames from one of 100 Deepfake-Eval-2024 videos. Authentic videos generally receive higher scores and generated videos generally receive lower scores, although the distributions overlap.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Description
Number of inference steps
28
Denoising timesteps
Guidance scale
3.5
Classifier-free guidance strength
Scheduler
FlowMatchEulerDiscrete
Sampling scheduler used during inversion
DDIM η base
0.95
Base eta for interpolated denoise
Eta trend
constant
Trend of eta values (e.g., constant, linear)
Eta start step
0
Step where eta schedule starts
Appendix
TABLE III : General inversion hyperparameters used across all experiments.
SD2.1
SD3
SD3.5
FLUX.1
FLUX.1
Boreal
FLUX.2
Boreal
HiDream
GPT
[ 16 ]
[ 1 ]
[ 17 ]
[ 2 ]
+LoRA [ 19 ]
FLUX.1 [ 20 ]
[ 18 ]
FLUX.2 [ 21 ]
O1 [ 23 ]
Image-2 [ 22 ]
Detector
Year
2022
2024
2024
2024
2024
2024
2025
2025
2026
2026
UFD [ 13 ]
2023
49.5
50.0
50.5
48.5
49.0
49.0
49.0
48.5
49.0
50.0
AEROBLADE [ 31 ]
2024
50.0
50.0
50.0
50.0
50.0
50.0
50.0
50.0
50.0
50.0
FatFormer [ 45 ]
2024
49.5
49.0
63.0
49.0
51.0
59.5
52.5
52.0
57.0
60.5
FreqNet [ 42 ]
2024
46.5
44.5
90.0
82.5
84.0
81.0
83.0
74.5
78.5
81.0
Appendix
TABLE IV : Balanced accuracy (%) of twenty detectors on ten generators. Each cell uses 100 authentic and 100 generated images at 512×512 . Bold marks the best detector per generator.
Fig. 12 : Zero-shot prediction score distributions for each evaluated detector. Each subplot shows the histogram of classifier outputs for real and fake images. Vertical dashed lines indicate classification thresholds. Most methods struggle to generalize to unseen generators, misclassifying a significant number of fake images as real.
Fig. 13 : Prediction score distributions before and after PGD attack for each evaluated detector (excluding D3, shown separately in Figure 7 ). Each plot shows how adversarial perturbation erodes separability between real and fake inputs, often collapsing confidence entirely.
Fig. 14 : Six image samples generated from the same prompt ( Appendix G ), with different random seeds. All were generated under the medium-resource attacker’s sampling budget. These highlight the diversity in candidate outputs that the attacker can choose from.
Metric
Fake mean
Real mean
Fake − Real
PSNR
23.6922
22.0028
+1.6894
SSIM
0.7207
0.6036
+0.1171
1−LPIPS
0.9240
0.9126
+0.0114
CLIP similarity
0.9666
0.9293
+0.0373
Appendix
TABLE V : Inversion metric statistics on the calibration split. Synthetic images reconstruct more faithfully on average, consistent with the hypothesis that generated images are easier to re-synthesize under inversion.
Metric (alone)
AUCeff↑
Overlap ↓
Cohen’s d↑
PSNR
0.6515
0.6942
0.4896
SSIM
0.7284
0.6040
0.8630
(1−LPIPS)
0.5967
0.8062
0.3378
CLIP similarity
0.8546
0.4401
1.2996
Appendix
TABLE VI : Single-metric separability under inversion (weights optimized independently for each metric). CLIP similarity provides the strongest standalone signal.
Removed metric
AUCeff↑
Overlap ↓
Cohen’s d↑
Δ Overlap / Δd
Remove PSNR
0.8841
0.3476
1.5690
+0.0040 / − 0.0180
Remove SSIM
0.8609
0.3936
1.3956
+0.0500 / − 0.1913
Remove (1−LPIPS)
0.8623
0.4035
1.4163
+0.0599 / − 0.1706
Remove CLIP
0.7693
0.5284
0.9619
+0.1848 / − 0.6251
None (All 4)
0.8852
0.3436
1.5870
baseline
Appendix
TABLE VII : Leave-one-out ablation of A-index metrics. Removing CLIP causes the largest degradation in separability, while removing PSNR has minimal impact.
# Metrics
Best subset (by overlap)
Overlap ↓
Cohen’s d↑
1
CLIP
0.4401
1.2996
2
SSIM + CLIP
0.4028
1.3596
3
SSIM + (1−LPIPS) + CLIP
0.3476
1.5690
4
PSNR + SSIM + (1−LPIPS) + CLIP
0.3436
1.5870
Appendix
TABLE VIII : Best-performing subset at each subset size from the exhaustive metric sweep. The full four-metric combination achieves the strongest separability.
Subset
AUCeff↑
Overlap ↓
Cohen’s d↑
Single metrics
PSNR
0.6515
0.6942
0.4896
SSIM
0.7284
0.6040
0.8630
(1−LPIPS)
0.5967
0.8062
0.3378
CLIP
0.8546
0.4401
1.2996
Two metrics
Appendix
TABLE IX : Performance of all 15 non-empty metric subsets with independently optimized weights. Combining complementary inversion signals consistently improves separability, with the full four-metric combination achieving the strongest performance.
Perturbation level
Overlap (mean ± std) ↓
Cohen’s d (mean ± std) ↑
AUCeff↑
0%
0.3436±0.0000
1.5870±0.0000
0.8852
10%
0.3781±0.0090
1.5834±0.0172
0.8843
20%
0.3830±0.0094
1.5756±0.0302
0.8821
30%
0.3917±0.0199
1.5480±0.0660
0.8772
50%
0.4380±0.0784
1.4083±0.2395
0.8495
70%
0.4932±0.1365
1.2415±0.4086
0.8129
Appendix
TABLE X : Sensitivity of separability to perturbations of the learned A-index weights. Performance degrades gradually under moderate perturbations, indicating robustness of the learned weighting scheme.
Method
AUCeff↑
Overlap ↓
Cohen’s d↑
#Params
Interp.
CV std (Overlap)
SVM-RBF (5-fold)
0.9125
0.3003
2.0610
O(n)
N
0.0323
Random forest (5-fold)
0.8948
0.3470
1.9023
> 1000
N
0.0504
Linear + sigmoid (ours)
0.8852
0.3436
1.5870
4
Y
–
Logistic reg. (5-fold)
0.8848
0.3629
1.7271
5
Y
0.0425
Product (log-space)
0.8777
0.3489
1.5437
4
Y
–
MLP (16,8) (5-fold)
0.8518
0.4537
1.5169
225
N
0.0742
Appendix
TABLE XI : Comparison to alternative aggregation strategies sorted by AUCeff . Our method ( Linear + sigmoid ) achieves competitive performance while remaining interpretable and using only four parameters. Supervised models are evaluated via 5-fold cross-validation.
Fig. 15 : Examples from the keyword-sampled Reddit corpus. The corpus contains photographs, screenshots, memes, and digital artwork.
Attack
SSIM
CLIP
A-index
Δτs
Horizontal flip
0.645
0.963
0.0045
− 87.7%
Crop 20%
0.638
0.916
0.0067
− 81.6%
Rotation 15 ∘
0.607
0.926
0.0074
− 79.7%
Crop 10%
0.671
0.948
0.0086
− 76.4%
Brightness ↓ (0.5 × )
0.690
0.942
0.0115
− 68.5%
Brightness ↑ (1.5 × )
0.810
0.946
0.0132
− 63.8%
Appendix
TABLE XII : A-index of generated SD3 images under semantic edits and transformations. The safety threshold is τsafety=0.0365 and the security threshold is τsecurity=0.038 (SD3 medium). Every edit remains below τsafety except Gaussian noise, which remains below τsecurity .
Fig. 16 : Qualitative examples of some transformations and their SD3 inversions. Each panel shows the edited query on the left and its SD3 Medium inverted reconstruction on the right. Overlays, crops, JPEG compression, downsampling, and flipping leave the reconstruction faithful, so the A-index stays well below τsafety , consistent with Table XII .
We survey deepfake generation and detection techniques, covering all deepfake media types: image, video, audio and multimodal content. We identify various kinds of deepfakes and construct taxonomies of deepfake generation and detection methods, illustrating the important groups of methods. Next, we gather datasets used for deepfake detection and provide updated rankings of the best performing detectors on the most popular datasets. In addition, we develop a novel multimodal benchmark to evaluate deepfake detectors on out-of-distribution content. The results indicate that state-of-the-art detectors fail to generalize to deepfakes generated by unseen generators. Our project page and new benchmark are available at https://github.com/CroitoruAlin/biodeep.
Department of Computer Science, University of Bucharest, Romania · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE · Department of Computer Science, University of Central Florida, Orlando, US
The rapid advancement of generative models has blurred the boundary between synthetic and real imagery, creating an urgent need for reliable deepfake detection. Yet most existing approaches rely on massive real--fake datasets, which are increasingly difficult to maintain as new generators continue to emerge. In this work, we investigate how much information about image authenticity is already encoded in modern multimodal vision representations. We find that frozen multimodal encoders naturally separate real and synthetic images in their embedding space, enabling a simple linear classifier to achieve strong performance without task-specific fine-tuning. Motivated by this observation, we develop a representation-aware data curation strategy that selects a compact set of representative generators for training. The resulting training set contains only 10K images, compared to 288K in AIGIBench and 4M in OpenFake, while improving robustness to unseen generators and distribution shifts. We additionally introduce RealWorldBench, a benchmark consisting of modern camera photographs, contemporary stock images, and outputs from recent commercial generators. Experiments across multiple benchmarks show that combining frozen multimodal representations with carefully curated training data provides a simple and effective approach to AI-generated image detection.
Modern deepfake detectors are rarely consumed as bare classifiers. In moderation, provenance, and verification pipelines their output probability is read as a degree of trust, so its calibration matters as much as raw accuracy. We reframe deepfake detection as a calibrated, self-auditing trust instrument, the Calibrated Deepfake Trust Score (CDTS), and identify what governs its trustworthiness. Our central finding is a competence-calibration coupling: the calibration of the trust score degrades as the detector's discriminative competence falls. We establish it across 32 configurations (pooled Pearson r = -0.81), demonstrate it within a single dataset, reinforce it by inducing low competence directly, and replicate it on a fourth held-out dataset the detectors never trained on. It holds across three architecturally distinct detectors, two convolutional networks and a CLIP vision transformer (r = -0.88, -0.83, -0.86). The result is also deployable: a single calibrator frozen on in-domain data fails on exactly the low-competence generators the coupling flags (its error tracks competence at r = -0.98), and competence is estimable without labels, so a label-free monitor flags calibration risk on unseen generators and routing source-batches on a reference-free competence estimate lowers overall AURC and improves the low-to-mid coverage operating region relative to confidence-based routing. The same competence factor also drives calibration inequity across demographic subgroups (distinct from accuracy inequity) and explanation faithfulness. We therefore argue that detector trustworthiness is organized by competence as a shared driver, that competence is the right quantity to estimate and condition on, and that trust scoring must be competence-aware. We offer the CDTS wrapper as the mechanism, and report openly where the unification is tight and where it is architecture-specific.
Md Anas Biswas
School of Computing, University of Portsmouth, Portsmouth, United Kingdom