Organizations: School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing, 100049, China · State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, 100190, China · University of Chinese Academy of Sciences, Beijing, 100049, China · Faculty of Information Science and Engineering, Ocean University of China, Qingdao, 266404, China
As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulations. To address this limitation, researchers have turned to large-scale foundation models, which can provide richer representations and better generalization. Nevertheless, relying on a single foundation model alone remains insufficient for effective forgery detection. While models like CLIP offer robust global semantic cues, they lack the capacity to capture detailed local facial features. In contrast, DINO excels at capturing local structural features of faces, but provides weaker global semantic context. To fully utilize the synergies among multiple foundation models, we propose a hierarchical multi-granular framework that integrates complementary pretrained representations. Specifically, a Global Context Branch (GCB) based on CLIP captures holistic semantic cues, while a Fine-grained Cue Branch (FCB) built on DINOv3 captures localized structural irregularities. In addition, we design a feature fusion module that enables parameter-efficient adaptation of the frozen foundation backbones by adaptively extracting and integrating complementary features from the two models. By jointly leveraging global context and fine-grained cues, our method learns more comprehensive forgery representations and achieves strong cross-manipulation performance. Extensive experiments on multiple benchmarks demonstrate the benefit of the proposed design, particularly under cross-dataset and cross-manipulation settings.
Figures & tables
Figure 1: Overview of the proposed framework. A Global Context Branch (GCB) and a Fine-grained Cue Branch (FCB) operate in parallel and are connected via adaptive cross-feature interaction modules, enabling effective collaboration between global context and fine-grained forgery cues.
Figure 2: Detailed structures of the core components in the proposed framework. (a) The Adaptive Feature Learner (AFL) extracts multi-scale spatial priors. (b) The Cross-Feature Interaction (CFI) block progressively exchanges complementary information between global context and fine-grained forgery cues. (c) The multi-scale decoder aggregates hierarchical interaction features for the final real/fake prediction.
Method
Venue
CDF
DFDC
DFDCP
Xception [ 37 ]
ICCV’19
73.65
70.77
73.74
FaceX-ray [ 24 ]
CVPR’20
67.86
63.26
69.42
RECCE [ 2 ]
CVPR’22
73.19
71.33
74.19
SBI [ 40 ]
CVPR’22
81.30
71.96
79.90
UIA-ViT [ 55 ]
ECCV’22
82.41
–
75.80
UCF [ 50 ]
ICCV’23
75.27
71.91
75.94
Table 1: Cross-dataset comparison using frame-level AUC. The results reported in the table are taken from [ 51 , 31 ] , or directly obtained from the corresponding original papers.
Method
Venue
CDF
DFDC
DFDCP
FaceX-ray [ 24 ]
CVPR’20
–
–
71.1
FTCN [ 54 ]
ICCV’21
86.9
67.6
74.0
SBI [ 40 ]
CVPR’22
92.8
71.9
85.5
UIA-ViT* [ 55 ]
ECCV’22
82.4
75.0
75.8
CFM* [ 30 ]
TIFS’23
85.3
75.0
80.2
AltFreezing* [ 46 ]
CVPR’23
85.1
71.7
79.3
Table 2: Comparison with SOTA methods using the video-level AUC. Results marked with * are obtained by using the authors’ released models or code, while the remaining results are directly taken from [ 51 ] or corresponding original papers.
Method
Venue
uniface
facedancer
fsgan
inswap
simswap
Avg.
RECCE [ 2 ]
CVPR’22
84.2
78.3
88.4
79.5
73.0
80.7
SBI [ 40 ]
CVPR’22
64.4
44.7
87.9
63.3
56.8
63.4
IID [ 18 ]
CVPR’23
79.5
79.0
86.4
74.4
64.0
76.7
UCF [ 50 ]
ICCV’23
78.7
80.0
88.1
76.8
64.9
77.7
LSDA [ 48 ]
CVPR’24
85.4
75.9
83.2
81.0
72.7
79.6
CDFA [ 26 ]
ECCV’24
76.5
75.4
84.8
72.0
76.1
77.0
Table 3: Cross-manipulation comparison on five representative face swapping forgery types in DF40 [ 49 ] using frame-level AUC (%)
Component Settings
Frame-level AUC
GCB
MG
FCB
FF++
Celeb-DF
DFDC
facedancer
inswap
✓
82.6 ± 0.1
75.6 ± 1.0
72.6 ± 0.6
77.4 ± 0.6
76.2 ± 0.9
✓
92.5 ± 0.3
82.8 ± 0.3
73.2 ± 0.2
75.0 ± 0.3
76.1 ± 0.6
✓
✓
88.8 ± 0.2
75.7 ± 0.7
73.2 ± 0.5
78.0 ± 0.5
79.4 ± 1.1
✓
✓
96.0 ± 0.1
80.7 ± 0.4
76.8 ± 0.9
88.1 ± 0.4
92.2 ± 1.4
✓
✓
95.9 ± 0.2
81.7 ± 0.9
75.3 ± 0.5
83.9 ± 0.6
92.3 ± 1.2
Table 4: Ablation study of different components on frame-level AUC across multiple datasets. GCB denotes the Global Context Branch for modeling global semantic consistency, FCB denotes the Fine-grained Cue Branch for capturing local manipulation artifacts, and MG denotes the multi-granular feature aggregation module for hierarchical feature fusion.
Variant
Avg. frame-level AUC ↑
Latency (per frame)
Params
GFLOPs
GCB only
76.83
30 ms
330M
248
FCB only
79.92
31 ms
330M
249
GCB + FCB (Naive concat)
79.03
33 ms
607M
367
DBCF (Ours)
88.06
56 ms
678M
544
Table 5: Effectiveness-efficiency comparison of representative model variants. Latency and computational metrics are measured under the same inference setting and reported on a per-frame basis.
Figure 3: Qualitative visualization of attention maps from the GCB and FCB branches on different forgery datasets.
Figure 4: Frame-level AUC reduction ratio under different degradation levels and perturbation types. "Average" score represents the mean across all levels for each type of perturbation.
Method
O&C&I → M HTER ↓ /AUC ↑
O&M&I → C HTER ↓ /AUC ↑
O&C&M → I HTER ↓ /AUC ↑
C&I&M → O HTER ↓ /AUC ↑
Avg HTER ↓
Avg AUC ↑
NAS-FAS
19.53 / 88.63
16.54 / 90.18
14.51 / 93.84
13.80 / 93.43
16.10
91.52
SSAN-R
6.67 / 98.75
10.00 / 96.67
8.88 / 96.79
13.72 / 93.63
9.82
96.46
PatchNet
7.10 / 98.46
11.33 / 94.58
13.40 / 95.67
11.82 / 95.07
10.91
95.95
SA-FAS
5.95 / 96.55
8.78 / 95.37
6.58 / 97.54
10.00 / 96.23
7.83
96.42
AG-FAS
5.71 / 98.03
5.44 / 98.55
6.71 / 98.23
9.43 / 96.62
6.82
97.86
Ours (GCB)
5.95 / 98.50
2.67 / 99.54
17.14 / 91.24
10.14 / 96.31
8.98
96.40
Table 6: Performance comparison of different face anti-spoofing methods under cross-dataset evaluation. All results for the compared methods are taken from [ 29 ] . The metrics reported include HTER and AUC for each source-target dataset combination, as well as the mean performance across all settings.