Recent advances in generative image models have made many manipulated images highly realistic, raising the need for detectors that are not only accurate but also able to provide visual evidence for their decisions. In this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues. Our method follows a modular detection-and-explanation design. For the real/fake decision, we build a multi-backbone detector that combines several DINOv3 models with Mesorch manipulation-localization features, bringing together pretrained visual representations, DCT-aware cues, and multi-scale forensic information. To inject explanation evidence into the detector, we use a Grounding-DINO-based pseudo-mask generation pipeline that converts local artifact descriptions from training explanations into weak patch- level supervision for an Artifact Evidence Map. We further introduce a local patch-level contrastive objective that separates artifact and authenticity evidence in the detector feature space without requiring paired images or pixel-level manipulation masks. For language output, we use class-conditional Qwen3-VL models to generate complex explanations for fake and real predictions, followed by a GRPO-optimized text simplification model. The proposed methods were trained and evaluated on the challenge subset of XPlainVerse. On the full test split, our submission achieves 0.9349 detection accuracy, 0.5571 explanation score, and a 0.7456 final challenge score.
Figures & tables
Figure 1. Overview of the proposed system, including offline pseudo-mask generation, multi-backbone forensic detection, and class-conditional explanation generation. A three-part pipeline: explanation-grounded pseudo-mask generation, a multi-backbone detector with image and patch losses, and class-conditional complex and simple explanation generation.
CF
UF
WF
CF
pull(CF, CF)
pull(sg(CF), UF)
–
UF
pull(UF, sg(CF))
pull(UF, UF)
–
WF
–
–
–
Table 1. Local contrastive rules for fake–fake patch pairs.
CR
UR
WR
CR
pull(CR, CR)
pull(sg(CR), UR)
pull(sg(CR), WR)
UR
pull(UR, sg(CR))
pull(UR, UR)
pull(sg(UR), WR)
WR
pull(WR, sg(CR))
pull(WR, sg(UR))
–
Table 2. Local contrastive rules for real–real patch pairs.
CR
UR
WR
CF
push(CF, CR)
push(sg(CF), UR)
push(sg(CF), WR)
UF
push(UF, sg(CR))
push(UF, UR)
pull(sg(UF), WR)
WF
–
pull(sg(WF), UR)
pull(sg(WF), WR)
Table 3. Local contrastive rules for fake–real patch pairs.
Figure 2. Internal-validation examples. From top to bottom: pseudo boxes, ground-truth simple explanations, Artifact Evidence Maps, and generated simple explanations. Eleven examples showing grounded pseudo boxes, reference simple explanations, detector evidence heatmaps, and generated simple explanations.
Variant
λimg
λart
λlcl
Val. AUC
Image-level detector only
15
0
0
0.8975
+ pseudo-mask artifact supervision
15
1
0
0.9139
+ LCL
15
1
1
0.9417
Table 4. Detector ablation on the public validation split.
Method
Bc
Fent
Fevid
Bs
Ls
Ss
Mexp
Zero-shot Qwen3-VL-253B-A22B-Instruct
0.6023
0.5310
0.3934
0.4495
0.3418
0.4172
0.4812
Organizer baseline 1 1 1 The organizer baseline model can be found on the official challenge website: https://explainable-deepfake-detection.github.io/
0.7109
0.5536
0.4620
0.4578
0.4264
0.4484
0.5365
Ours
0.7253
0.6235
0.5447
0.4991
0.9703
0.6405
0.6236
Table 5. Explanation generation evaluation on the internal held-out validation subset. Bc and Bs denote complex and simple BERTScore-F1; Fent and Fevid denote Entity and Evidence F1; Ls , Ss , and Mexp denote normalized simple SLE, Simple score, and Explanation score.
Metric
Score
Detection accuracy
0.9349
Detection fake F1
0.9418
Detection macro F1
0.9340
Detection real F1
0.9261
Complex BERTScore-F1
0.7004
Simple BERTScore-F1
0.6509
Table 6. Final submission results on the full final test split.
Deepfake detection is moving beyond binary classification decisions toward systems that can also explain the visual evidence supporting those decisions. This transition is important for real-world verification settings, where diverse users need to understand not only whether an image is manipulated, but also why it is considered suspicious. The Explainable Deepfake Detection Challenge at ACM Multimedia 2026 is designed to benchmark this joint capability. Built on XPlainVerse, a million-scale benchmark for explainable deepfake detection, the challenge evaluates methods on image classification and grounded natural-language explanation generation. Participants submit a real/fake label together with two explanations for each image: a detailed complex explanation for technical users and a concise simple explanation for general users. The evaluation combines classification metrics with semantic similarity, simplicity, and intent-aware grounding metrics that assess whether explanations identify the relevant manipulated entities and supporting visual evidence. The methodologies developed through the challenge will contribute to the development of next-generation explainable deepfake detectors. Evaluation script, baseline models, and accompanying code are available on https://github.com/Abhijeet8901/XPlainVerse-ACMChallenge.
Abhijeet Narang, Kartik Kuckreja, Shreya Ghosh +4
American University of Sharjah Sharjah, United Arab Emirates
As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limiting reliability and user trust. Existing benchmarks mainly evaluate classification accuracy, overlooking whether explanations reflect the actual manipulations. This gap hinders progress toward deployable, explainable deepfake detection systems. To this end, we introduce XPlainVerse, a large-scale benchmark designed for joint deepfake detection and human-centered explanation. XPlainVerse comprises one million real and manipulated images, pairing authentic images from five established sources with forgeries generated by twelve off-the-shelf image editing and synthesis models. We further propose a multi-stage filtering pipeline, Edit-Check, to verify if manipulations satisfy their intended edits, enabling reliable reasoning supervision at scale. Beyond dataset scale, XPlainVerse provides two complementary explanation styles: technical explanations for expert analysis and simplified explanations optimized for non-technical users. To evaluate explanation quality beyond surface similarity, we propose novel metrics, EntityScore and EvidenceScore, that measure reasoning fidelity by checking whether explanations correctly identify manipulated entities and visual evidence. Human annotations on 2,000 explanation pairs validate our dataset quality against human judgment. We believe XPlainVerse will establish grounded explanation quality as a measurable dimension of deepfake detection and support scalable research on trustworthy, interpretable models.
Abhijeet Narang, Kartik Kuckreja, Shreya Ghosh +3
Monash University · MBZUAI · The University of Queensland
Multimodal large language models (MLLMs) can explain deepfake verdicts in natural language, but such explanations are not necessarily visually grounded in the visual evidence underlying the prediction. A model may describe plausible artifacts inferred from language priors rather than from image evidence. Existing grounding methods improve visual reliance through decoding or attention interventions, but they generally strengthen grounding over the entire image, making them ill-suited for forensic artifacts that are subtle, spatially localized, and image-dependent. We propose Look Before You Judge, a training-free framework that formulates explainable deepfake detection as a sequential evidence acquisition process. Instead of directly predicting image authenticity from holistic visual reasoning, our framework first identifies image-specific candidate evidence regions by contrasting the MLLM's decoder-to-visual attention between an original image and its Gaussian-blurred counterpart. The identified regions are then inspected individually, and the resulting local evidence is integrated with the global image context before reaching a final verdict. The framework operates without manipulation masks, external forensic models, or parameter updates, making it directly applicable to off-the-shelf MLLMs. Across five open-source MLLMs on TriDF and MMTD-Set, our framework improves detection accuracy by up to 12.8%, reduces CHAIR by up to 33.4% and hallucination rate by up to 21.3%, and outperforms representative training-free decoding and attention methods.
Chia-Ling Chen, Yu-Ting Ta, Jian-Yu Jiang-Lin +8
National Taiwan University · National Tsing Hua University · National Yang Ming Chiao Tung University