Organizations: Department of Networks and Digital Media, Kingston University London, UK · Centre for Research and Technology Hellas (CERTH), Thessaloniki, Greece · Department of Computer Science, Kingston University London, UK
Deepfake detection methods have become increasingly effective yet most provide limited insight into the evidence behind their predictions. However, in forensic settings users also need to know which manipulation cues support the decision and where they appear. Existing explainability methods only partially address this need since localization-based approaches lack semantic descriptions while language-based explanation methods are only weakly grounded in visual evidence. In this work, we propose DF-CBM, a region-aware concept bottleneck model for explainable deepfake detection. DF-CBM builds a compact vocabulary of manipulation-related concepts from textual artifact annotations and links each concept to plausible facial and boundary regions. It then predicts these concepts from visual features using a concept-specific masked attention mechanism guided by parsed facial masks and the final real/fake decision is made from the predicted concept bottleneck. Our experiments show that DF-CBM outperforms concept-based baselines in concept prediction and deepfake classification while remaining competitive with state-of-the-art black-box detectors. Finally, qualitative results and intervention analyses demonstrate that DF-CBM provides spatially grounded concept evidence and enables counterfactual explanations of how individual manipulation concepts influence the final prediction. Our code is available at: https://github.com/GeorgeTsoumplekas/DF-CBM.
Figures & tables
Figure 1 : Explanation capabilities of existing deepfake detection approaches and DF-CBM that provides both localized visual evidence and grounded manipulation concepts.
Figure 2 : Overview of the region-aware concept bottleneck model in DF-CBM.
Model
B-Acc.
F1
F-AUC
V-AUC
Joint-CBM [ 26 ]
0.687
0.310
0.739
0.753
BotCL [ 56 ]
0.651
0.300
0.689
0.693
DF-CBM
0.675
0.563
0.743
0.772
Table 1 : Macro-averaged concept prediction performance on the FaceForensics++ test set.
Methods
UniFace
BlendFace
MobSwap
e4s
FaceDan
FSGAN
InSwap
SimSwap
Avg.
SBI [ 44 ]
0.724
0.891
0.952
0.750
0.594
0.803
0.712
0.701
0.766
UCF [ 62 ]
0.831
0.827
0.950
0.731
0.862
0.937
0.809
0.647
0.824
IID [ 16 ]
0.839
0.789
0.888
0.766
0.844
0.927
0.789
0.644
0.811
LSDA [ 59 ]
0.872
0.875
0.930
0.694
0.721
0.939
0.855
0.793
0.835
ProDet [ 5 ]
0.908
0.929
0.975
0.771
0.747
0.928
0.837
0.844
0.867
CDFA [ 33 ]
0.762
0.756
0.823
0.631
0.803
0.942
0.772
0.757
0.781
Table 2 : Intra-dataset deepfake detection performance on FaceForensics++ using video-level AUC. DF-CBM is compared with black-box detectors (top-part) and concept-based baselines (bottom part) across eight manipulation methods. Best results among the concept-based methods are shown in bold .
Methods
CDF-v2
DFD
DFDC
DFDCP
UADFV
Avg.
SBI [ 44 ]
0.886
0.827
0.717
0.848
-
-
UCF [ 62 ]
0.837
0.867
0.742
0.770
-
-
IID [ 16 ]
0.838
0.939
0.700
0.689
-
-
LSDA [ 59 ]
0.875
0.881
0.701
0.812
-
-
ProDet [ 5 ]
0.926
0.901
0.707
0.828
-
-
CDFA [ 33 ]
0.938
0.954
0.830
0.881
-
-
Table 3 : Cross-dataset deepfake detection performance using video-level AUC. Models are trained on FaceForensics++ and evaluated on unseen datasets. Best results among the concept-based methods (bottom part) are shown in bold .
Methods
CDF-v2
DFD
DFDC
DFDCP
UADFV
Avg.
SBI [ 44 ]
0.813
0.774
-
0.799
-
-
UCF [ 62 ]
0.753
0.807
0.719
0.759
-
-
ED [ 1 ]
0.864
-
0.721
0.851
-
-
CFM [ 35 ]
0.828
0.915
-
0.758
-
-
FoCus [ 50 ]
0.720
-
0.669
0.778
-
-
LSDA [ 59 ]
0.830
0.880
0.736
0.815
-
-
Table 4 : Cross-dataset deepfake detection performance using frame-level AUC. Models are trained on FaceForensics++ and evaluated on unseen datasets. Best results among the concept-based methods (bottom part) are shown in bold .
Figure 3 : Qualitative examples of DF-CBM explanations. For each manipulated image, we show concept-specific attention maps, their overlays on the input image and the corresponding concept contribution scores.
Methods
UniFace
BlendFace
MobSwap
e4s
FaceDan
FSGAN
InSwap
SimSwap
Avg.
DF-CBM
0.912
0.878
0.919
0.968
0.845
0.928
0.835
0.891
0.897
w/o region prior
0.884
0.789
0.884
0.926
0.812
0.879
0.807
0.809
0.849
w/o concept-specific attention
0.830
0.810
0.866
0.936
0.811
0.896
0.768
0.843
0.845
w/o concept bottleneck
0.868
0.837
0.927
0.933
0.877
0.912
0.844
0.877
0.884
Table 5 : Architectural component ablation on FaceForensics++ using video-level AUC.
Methods
UniFace
BlendFace
MobSwap
e4s
FaceDan
FSGAN
InSwap
SimSwap
Avg.
DF-CBM
0.912
0.878
0.919
0.968
0.845
0.928
0.835
0.891
0.897
frozen text queries
0.911
0.867
0.919
0.967
0.851
0.932
0.829
0.887
0.895
random query initialization
0.827
0.811
0.875
0.955
0.805
0.907
0.784
0.825
0.849
Table 6 : Query design ablation for the masked cross-attention mechanism on FaceForensics++ using video-level AUC.
Figure 4 : Sequential concept intervention traces for two misclassified samples: (a) ground-truth real, (b) ground-truth fake. At each step, a single concept is intervened on while retaining all previous interventions. The left panels show the evolution of the raw class logits, whereas the right panels show the corresponding softmax probabilities.
Deepfake detection is moving beyond binary classification decisions toward systems that can also explain the visual evidence supporting those decisions. This transition is important for real-world verification settings, where diverse users need to understand not only whether an image is manipulated, but also why it is considered suspicious. The Explainable Deepfake Detection Challenge at ACM Multimedia 2026 is designed to benchmark this joint capability. Built on XPlainVerse, a million-scale benchmark for explainable deepfake detection, the challenge evaluates methods on image classification and grounded natural-language explanation generation. Participants submit a real/fake label together with two explanations for each image: a detailed complex explanation for technical users and a concise simple explanation for general users. The evaluation combines classification metrics with semantic similarity, simplicity, and intent-aware grounding metrics that assess whether explanations identify the relevant manipulated entities and supporting visual evidence. The methodologies developed through the challenge will contribute to the development of next-generation explainable deepfake detectors. Evaluation script, baseline models, and accompanying code are available on https://github.com/Abhijeet8901/XPlainVerse-ACMChallenge.
Abhijeet Narang, Kartik Kuckreja, Shreya Ghosh +4
American University of Sharjah Sharjah, United Arab Emirates
Multimodal large language models (MLLMs) can explain deepfake verdicts in natural language, but such explanations are not necessarily visually grounded in the visual evidence underlying the prediction. A model may describe plausible artifacts inferred from language priors rather than from image evidence. Existing grounding methods improve visual reliance through decoding or attention interventions, but they generally strengthen grounding over the entire image, making them ill-suited for forensic artifacts that are subtle, spatially localized, and image-dependent. We propose Look Before You Judge, a training-free framework that formulates explainable deepfake detection as a sequential evidence acquisition process. Instead of directly predicting image authenticity from holistic visual reasoning, our framework first identifies image-specific candidate evidence regions by contrasting the MLLM's decoder-to-visual attention between an original image and its Gaussian-blurred counterpart. The identified regions are then inspected individually, and the resulting local evidence is integrated with the global image context before reaching a final verdict. The framework operates without manipulation masks, external forensic models, or parameter updates, making it directly applicable to off-the-shelf MLLMs. Across five open-source MLLMs on TriDF and MMTD-Set, our framework improves detection accuracy by up to 12.8%, reduces CHAIR by up to 33.4% and hallucination rate by up to 21.3%, and outperforms representative training-free decoding and attention methods.
Chia-Ling Chen, Yu-Ting Ta, Jian-Yu Jiang-Lin +8
National Taiwan University · National Tsing Hua University · National Yang Ming Chiao Tung University
Deepfake (DF) technology poses a significant threat to information integrity, driving the need for robust detection methods. Most DF detectors only consider predicting a binary label for whether the input is real or fake, lacking the justification required for real-world applications like legal proceedings. Explainable DF Detection has emerged to address this limitation, but existing techniques frequently fall short by either relying on human annotations for precise artifact localization or generating superficially plausible textual explanations without grounding. This work investigates the use of post-hoc explainable AI (XAI) to analyze the decision-making process of state-of-the-art black-box DF detectors. Specifically, we employ Encoding-Decoding Direction Pairs (EDDP), a technique suitable for uncovering the concept space of DF detectors (their semantic vocabulary) as well as the mechanism for writing and reading concept information to and from internal representations. Our analysis reveals previously hidden real and fake features learned implicitly during detector training, offering nuanced explanations unattainable through conventional methods. This enables global model understanding, spatially aware concept localization, and counterfactual what-if analysis, all contributing to a deeper comprehension of DF detection strategies.