cs.CVSep 28, 2026

DBCF: Dual-Branch Complementary Fusion of Foundation Models for Generalized Deepfake Detection

Authors: Fengming Gu, Mingjie He, Zonghui Guo, Jie Zhangb, Shiguang Shan

Organizations: School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing, 100049, China · State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, 100190, China · University of Chinese Academy of Sciences, Beijing, 100049, China · Faculty of Information Science and Engineering, Ocean University of China, Qingdao, 266404, China

Abstract

As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulations. To address this limitation, researchers have turned to large-scale foundation models, which can provide richer representations and better generalization. Nevertheless, relying on a single foundation model alone remains insufficient for effective forgery detection. While models like CLIP offer robust global semantic cues, they lack the capacity to capture detailed local facial features. In contrast, DINO excels at capturing local structural features of faces, but provides weaker global semantic context. To fully utilize the synergies among multiple foundation models, we propose a hierarchical multi-granular framework that integrates complementary pretrained representations. Specifically, a Global Context Branch (GCB) based on CLIP captures holistic semantic cues, while a Fine-grained Cue Branch (FCB) built on DINOv3 captures localized structural irregularities. In addition, we design a feature fusion module that enables parameter-efficient adaptation of the frozen foundation backbones by adaptively extracting and integrating complementary features from the two models. By jointly leveraging global context and fine-grained cues, our method learns more comprehensive forgery representations and achieves strong cross-manipulation performance. Extensive experiments on multiple benchmarks demonstrate the benefit of the proposed design, particularly under cross-dataset and cross-manipulation settings.

Figures & tables

Explore similar work

CardsList
  1. Cross-Domain Generalization Limits of Vision Foundation Models in Facial Deepfake Detection

    May 24, 2026Ibrahim DelibasogluDeepfake DetectionDigital Forensics

  2. Revisiting Cross-Reconstruction for Generalizable Deepfake Detection

    Oct 1, 2026Bingjian Yang, Shilei Zhao, Zheng WangFace Forgery DetectionDeepfake Detection

  3. Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection

    Sep 7, 2026Xuechao Zou, Yi Zhou, Kai Li +4Deepfake DetectionForgeries