Organizations: School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing, 100049, China · State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, 100190, China · University of Chinese Academy of Sciences, Beijing, 100049, China · Faculty of Information Science and Engineering, Ocean University of China, Qingdao, 266404, China
As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulations. To address this limitation, researchers have turned to large-scale foundation models, which can provide richer representations and better generalization. Nevertheless, relying on a single foundation model alone remains insufficient for effective forgery detection. While models like CLIP offer robust global semantic cues, they lack the capacity to capture detailed local facial features. In contrast, DINO excels at capturing local structural features of faces, but provides weaker global semantic context. To fully utilize the synergies among multiple foundation models, we propose a hierarchical multi-granular framework that integrates complementary pretrained representations. Specifically, a Global Context Branch (GCB) based on CLIP captures holistic semantic cues, while a Fine-grained Cue Branch (FCB) built on DINOv3 captures localized structural irregularities. In addition, we design a feature fusion module that enables parameter-efficient adaptation of the frozen foundation backbones by adaptively extracting and integrating complementary features from the two models. By jointly leveraging global context and fine-grained cues, our method learns more comprehensive forgery representations and achieves strong cross-manipulation performance. Extensive experiments on multiple benchmarks demonstrate the benefit of the proposed design, particularly under cross-dataset and cross-manipulation settings.
Figures & tables
Figure 1: Overview of the proposed framework. A Global Context Branch (GCB) and a Fine-grained Cue Branch (FCB) operate in parallel and are connected via adaptive cross-feature interaction modules, enabling effective collaboration between global context and fine-grained forgery cues.
Figure 2: Detailed structures of the core components in the proposed framework. (a) The Adaptive Feature Learner (AFL) extracts multi-scale spatial priors. (b) The Cross-Feature Interaction (CFI) block progressively exchanges complementary information between global context and fine-grained forgery cues. (c) The multi-scale decoder aggregates hierarchical interaction features for the final real/fake prediction.
Method
Venue
CDF
DFDC
DFDCP
Xception [ 37 ]
ICCV’19
73.65
70.77
73.74
FaceX-ray [ 24 ]
CVPR’20
67.86
63.26
69.42
RECCE [ 2 ]
CVPR’22
73.19
71.33
74.19
SBI [ 40 ]
CVPR’22
81.30
71.96
79.90
UIA-ViT [ 55 ]
ECCV’22
82.41
–
75.80
UCF [ 50 ]
ICCV’23
75.27
71.91
75.94
Table 1: Cross-dataset comparison using frame-level AUC. The results reported in the table are taken from [ 51 , 31 ] , or directly obtained from the corresponding original papers.
Method
Venue
CDF
DFDC
DFDCP
FaceX-ray [ 24 ]
CVPR’20
–
–
71.1
FTCN [ 54 ]
ICCV’21
86.9
67.6
74.0
SBI [ 40 ]
CVPR’22
92.8
71.9
85.5
UIA-ViT* [ 55 ]
ECCV’22
82.4
75.0
75.8
CFM* [ 30 ]
TIFS’23
85.3
75.0
80.2
AltFreezing* [ 46 ]
CVPR’23
85.1
71.7
79.3
Table 2: Comparison with SOTA methods using the video-level AUC. Results marked with * are obtained by using the authors’ released models or code, while the remaining results are directly taken from [ 51 ] or corresponding original papers.
Method
Venue
uniface
facedancer
fsgan
inswap
simswap
Avg.
RECCE [ 2 ]
CVPR’22
84.2
78.3
88.4
79.5
73.0
80.7
SBI [ 40 ]
CVPR’22
64.4
44.7
87.9
63.3
56.8
63.4
IID [ 18 ]
CVPR’23
79.5
79.0
86.4
74.4
64.0
76.7
UCF [ 50 ]
ICCV’23
78.7
80.0
88.1
76.8
64.9
77.7
LSDA [ 48 ]
CVPR’24
85.4
75.9
83.2
81.0
72.7
79.6
CDFA [ 26 ]
ECCV’24
76.5
75.4
84.8
72.0
76.1
77.0
Table 3: Cross-manipulation comparison on five representative face swapping forgery types in DF40 [ 49 ] using frame-level AUC (%)
Component Settings
Frame-level AUC
GCB
MG
FCB
FF++
Celeb-DF
DFDC
facedancer
inswap
✓
82.6 ± 0.1
75.6 ± 1.0
72.6 ± 0.6
77.4 ± 0.6
76.2 ± 0.9
✓
92.5 ± 0.3
82.8 ± 0.3
73.2 ± 0.2
75.0 ± 0.3
76.1 ± 0.6
✓
✓
88.8 ± 0.2
75.7 ± 0.7
73.2 ± 0.5
78.0 ± 0.5
79.4 ± 1.1
✓
✓
96.0 ± 0.1
80.7 ± 0.4
76.8 ± 0.9
88.1 ± 0.4
92.2 ± 1.4
✓
✓
95.9 ± 0.2
81.7 ± 0.9
75.3 ± 0.5
83.9 ± 0.6
92.3 ± 1.2
Table 4: Ablation study of different components on frame-level AUC across multiple datasets. GCB denotes the Global Context Branch for modeling global semantic consistency, FCB denotes the Fine-grained Cue Branch for capturing local manipulation artifacts, and MG denotes the multi-granular feature aggregation module for hierarchical feature fusion.
Variant
Avg. frame-level AUC ↑
Latency (per frame)
Params
GFLOPs
GCB only
76.83
30 ms
330M
248
FCB only
79.92
31 ms
330M
249
GCB + FCB (Naive concat)
79.03
33 ms
607M
367
DBCF (Ours)
88.06
56 ms
678M
544
Table 5: Effectiveness-efficiency comparison of representative model variants. Latency and computational metrics are measured under the same inference setting and reported on a per-frame basis.
Figure 3: Qualitative visualization of attention maps from the GCB and FCB branches on different forgery datasets.
Figure 4: Frame-level AUC reduction ratio under different degradation levels and perturbation types. "Average" score represents the mean across all levels for each type of perturbation.
Method
O&C&I → M HTER ↓ /AUC ↑
O&M&I → C HTER ↓ /AUC ↑
O&C&M → I HTER ↓ /AUC ↑
C&I&M → O HTER ↓ /AUC ↑
Avg HTER ↓
Avg AUC ↑
NAS-FAS
19.53 / 88.63
16.54 / 90.18
14.51 / 93.84
13.80 / 93.43
16.10
91.52
SSAN-R
6.67 / 98.75
10.00 / 96.67
8.88 / 96.79
13.72 / 93.63
9.82
96.46
PatchNet
7.10 / 98.46
11.33 / 94.58
13.40 / 95.67
11.82 / 95.07
10.91
95.95
SA-FAS
5.95 / 96.55
8.78 / 95.37
6.58 / 97.54
10.00 / 96.23
7.83
96.42
AG-FAS
5.71 / 98.03
5.44 / 98.55
6.71 / 98.23
9.43 / 96.62
6.82
97.86
Ours (GCB)
5.95 / 98.50
2.67 / 99.54
17.14 / 91.24
10.14 / 96.31
8.98
96.40
Table 6: Performance comparison of different face anti-spoofing methods under cross-dataset evaluation. All results for the compared methods are taken from [ 29 ] . The metrics reported include HTER and AUC for each source-target dataset combination, as well as the mean performance across all settings.
The rapid evolution of generative models has enabled the creation of hyper-realistic facial deepfakes, exposing a critical vulnerability in modern digital forensics: the inability of detectors to generalize to unseen manipulation techniques. Traditional networks suffer from representation collapse, overfitting to localized artifact fingerprints of specific training generators. This work investigates whether modern Vision Foundation Models can serve as generalizable, out-of-the-box feature extractors capable of tracking forensic anomalies across entirely unseen generative manifolds. We conduct a systematic cross-domain evaluation comparing three foundational learning paradigms: fully supervised macro-semantic features (RoPE-ViT), pure self-supervised geometric features (DINOv3), and multi-teacher agglomerative representations (NVIDIA C-RADIOv4-H). By deploying frozen backbones subjected to downstream linear probing, we map the performance limitations of these architectures on the challenging DF40 benchmark. Our empirical findings expose the intrinsic trade-offs between pre-training paradigms and parameter scale, proving that while foundation models retain high discriminative capabilities for entire face synthesis, localized face editing techniques expose fundamental boundaries in linear probe evaluation structures. Source code and model weights are available in http://github.com/mribrahim/deepfake
Ibrahim Delibasoglu
Department of Software Engineering, Faculty of Computer and Information Sciences, Sakarya University, Esentepe, Sakarya, 54050 Türkiye.
Existing image forgery detectors often suffer from generalization to unseen manipulation methods due to the limited ability to capture transferable forensic cues. Recent cross-reconstruction based methods attempt to improve generalization through semantic-artifact disentanglement, but typically align heterogeneous artifacts across generators and exclude artifact representations during reconstruction, which may overlook the inherent diversity and visual cues of manipulation artifacts. In this work, we revisit cross-reconstruction and introduce an artifact-oriented disentanglement framework for robust image forgery detection. We argue that \textbf{artifact diversity}, i.e., the intrinsic variations of manipulation artifacts introduced by different generation processes, contains complementary forensic cues rather than undesirable domain variations. Instead of enforcing explicit artifact alignment, our framework preserves diverse artifact characteristics through semantically aligned cross-generator reconstruction. Furthermore, we incorporate artifact representations into the reconstruction process and introduce a masked frequency-aware reconstruction strategy to emphasize manipulation-related residuals while reducing semantic interference. This design enables the model to learn transferable forensic representations from diverse artifacts. Extensive experiments on multiple benchmark datasets demonstrate improvements under both cross-dataset and cross-generator evaluation settings. Further analysis and ablation studies validate the effectiveness of artifact diversity preservation and artifact-aware cross-reconstruction.
Bingjian Yang, Shilei Zhao, Zheng Wang
National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University, China
The growing realism and accessibility of manipulated and generated faces threaten the trustworthiness of digital media. To detect such forgeries, deepfake detectors based on vision foundation models have shown promising performance, but they typically rely on a single pretrained representation and are prone to overfitting to particular training distributions. To improve generalization to unseen forgeries, we propose UCF-Net, an uncertainty-aware cascaded fusion network that harnesses CLIP's language-aligned semantic priors and DINO's self-supervised visual-structure priors. UCF-Net extracts hierarchical features across Transformer depths, uses layer-wise expert aggregation to adaptively combine each encoder's multi-level cues, and performs weighted fusion of the resulting representations based on entropy-derived uncertainty. We further consolidate public deepfake datasets into a unified benchmark of approximately 4M images and construct a separate cross-generator evaluation set with over 8K face images from eight recent generators. On the unified benchmark, UCF-Net achieves the best mean AUC among the evaluated methods in both in-domain and cross-domain evaluations. On the cross-generator set, it adapts effectively with limited target-domain data, although zero-shot transfer remains challenging.
Xuechao Zou, Yi Zhou, Kai Li +4
Beijing Jiaotong University · Tsinghua University · Ant Group