cs.CVOct 1, 2026

Revisiting Cross-Reconstruction for Generalizable Deepfake Detection

Authors: Bingjian Yang, Shilei Zhao, Zheng Wang

Organizations: National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University, China

Abstract

Existing image forgery detectors often suffer from generalization to unseen manipulation methods due to the limited ability to capture transferable forensic cues. Recent cross-reconstruction based methods attempt to improve generalization through semantic-artifact disentanglement, but typically align heterogeneous artifacts across generators and exclude artifact representations during reconstruction, which may overlook the inherent diversity and visual cues of manipulation artifacts. In this work, we revisit cross-reconstruction and introduce an artifact-oriented disentanglement framework for robust image forgery detection. We argue that \textbf{artifact diversity}, i.e., the intrinsic variations of manipulation artifacts introduced by different generation processes, contains complementary forensic cues rather than undesirable domain variations. Instead of enforcing explicit artifact alignment, our framework preserves diverse artifact characteristics through semantically aligned cross-generator reconstruction. Furthermore, we incorporate artifact representations into the reconstruction process and introduce a masked frequency-aware reconstruction strategy to emphasize manipulation-related residuals while reducing semantic interference. This design enables the model to learn transferable forensic representations from diverse artifacts. Extensive experiments on multiple benchmark datasets demonstrate improvements under both cross-dataset and cross-generator evaluation settings. Further analysis and ablation studies validate the effectiveness of artifact diversity preservation and artifact-aware cross-reconstruction.

Figures & tables

Explore similar work

May 14, 2026cs.CV

Reduce the Artifacts Bias for More Generalizable AI-Generated Image Detection

As the misuse of AI-generated images grows, generalizable image detection techniques are urgently needed. Recent state-of-the-art (SOTA) methods adopt aligned training datasets to reduce content, size, and format biases, empowering models to capture robust forgery cues. A common strategy is to employ reconstruction techniques, e.g., VAE and DDIM, which show remarkable results in diffusion-based methods. However, such reconstruction-based approaches typically introduce limited and homogeneous artifacts, which cannot fully capture diverse generative patterns, such as GAN-based methods. To complement reconstruction-based fake images with aligned yet diverse artifact patterns, we propose a GAN-based upsampling approach that mimics GAN-generated fake patterns while preserving content, size, and format alignment. This naturally results in two aligned but distinct types of fake images. However, due to the domain shift between reconstruction-based and upsampling-based fake images, direct mixed training causes suboptimal results, where one domain disrupts feature learning of the other. Accordingly, we propose a Separate Expert Fusion (SEF) framework to extract complementary artifact information and reduce inter-domain interference. We first train domain-specific experts via LoRA adaptation on a frozen foundational model, then conduct decoupled fusion with a gating network to adaptively combine expert features while retaining their specialized knowledge. Rather than merely benefiting GAN-generated image detection, this design introduces diverse and complementary artifact patterns that enable SEF to learn a more robust decision boundary and improve generalization across broader generative methods. Extensive experiments demonstrate that our method yields strong results across 13 diverse benchmarks. Codes are released at: https://github.com/liyih/SEF_AIGC_detection.
Sep 28, 2026cs.CV

DBCF: Dual-Branch Complementary Fusion of Foundation Models for Generalized Deepfake Detection

As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulations. To address this limitation, researchers have turned to large-scale foundation models, which can provide richer representations and better generalization. Nevertheless, relying on a single foundation model alone remains insufficient for effective forgery detection. While models like CLIP offer robust global semantic cues, they lack the capacity to capture detailed local facial features. In contrast, DINO excels at capturing local structural features of faces, but provides weaker global semantic context. To fully utilize the synergies among multiple foundation models, we propose a hierarchical multi-granular framework that integrates complementary pretrained representations. Specifically, a Global Context Branch (GCB) based on CLIP captures holistic semantic cues, while a Fine-grained Cue Branch (FCB) built on DINOv3 captures localized structural irregularities. In addition, we design a feature fusion module that enables parameter-efficient adaptation of the frozen foundation backbones by adaptively extracting and integrating complementary features from the two models. By jointly leveraging global context and fine-grained cues, our method learns more comprehensive forgery representations and achieves strong cross-manipulation performance. Extensive experiments on multiple benchmarks demonstrate the benefit of the proposed design, particularly under cross-dataset and cross-manipulation settings.
Jun 12, 2026cs.CV

A Multi-Domain Feature Fusion Framework for Generalizable Deepfake Detection Across Different Generators

Deepfakes are artificially generated images, audio, or videos that threaten privacy, security, and information integrity. Detecting such content is crucial for countering disinformation, as the latest models generate highly realistic content. While spatial- or frequency-based approaches achieve good detection rates on Generative Adversarial Networks (GANs)-based generated deepfakes, they often struggle with recent diffusion model-generated images. In particular, existing approaches rarely exploit complementary multi-domain representations or systematically evaluate cross-generator robustness. To address these challenges, we propose a multi-domain deepfake detection framework called SGFF-Net (Spatial-Gradient-Frequency Fusion Network) that integrates spatial, gradient, and DWT (Discrete Wavelet Transform)-based frequency representations within a dual residual learning architecture. Experimental results show that the SGFF-Net achieves 98.95% accuracy in intra-dataset evaluation and improves performance in both cross-model (70.46%) and cross-paradigm (69.94%) settings. Incorporating multi-source training and data augmentation further enhances robustness, increasing accuracy from 70.46% to 79.80% in cross-model evaluation, from 69% to 78% in cross-paradigm evaluation, and from 61.50% to 75.80% on real-world data. Unlike single-domain detectors, the SGFF-Net learns complementary forensic cues across spatial, gradient, and wavelet-frequency domains, resulting in greater robustness under cross-generator and cross-paradigm evaluation. The results further show that combining multi-domain representations with data diversity and augmentation substantially improves generalization, providing practical insights for developing more reliable deepfake detection systems.