Recent advances in generative image models have made many manipulated images highly realistic, raising the need for detectors that are not only accurate but also able to provide visual evidence for their decisions. In this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues. Our method follows a modular detection-and-explanation design. For the real/fake decision, we build a multi-backbone detector that combines several DINOv3 models with Mesorch manipulation-localization features, bringing together pretrained visual representations, DCT-aware cues, and multi-scale forensic information. To inject explanation evidence into the detector, we use a Grounding-DINO-based pseudo-mask generation pipeline that converts local artifact descriptions from training explanations into weak patch- level supervision for an Artifact Evidence Map. We further introduce a local patch-level contrastive objective that separates artifact and authenticity evidence in the detector feature space without requiring paired images or pixel-level manipulation masks. For language output, we use class-conditional Qwen3-VL models to generate complex explanations for fake and real predictions, followed by a GRPO-optimized text simplification model. The proposed methods were trained and evaluated on the challenge subset of XPlainVerse. On the full test split, our submission achieves 0.9349 detection accuracy, 0.5571 explanation score, and a 0.7456 final challenge score.
Figures & tables
Figure 1. Overview of the proposed system, including offline pseudo-mask generation, multi-backbone forensic detection, and class-conditional explanation generation. A three-part pipeline: explanation-grounded pseudo-mask generation, a multi-backbone detector with image and patch losses, and class-conditional complex and simple explanation generation.
CF
UF
WF
CF
pull(CF, CF)
pull(sg(CF), UF)
–
UF
pull(UF, sg(CF))
pull(UF, UF)
–
WF
–
–
–
Table 1. Local contrastive rules for fake–fake patch pairs.
CR
UR
WR
CR
pull(CR, CR)
pull(sg(CR), UR)
pull(sg(CR), WR)
UR
pull(UR, sg(CR))
pull(UR, UR)
pull(sg(UR), WR)
WR
pull(WR, sg(CR))
pull(WR, sg(UR))
–
Table 2. Local contrastive rules for real–real patch pairs.
CR
UR
WR
CF
push(CF, CR)
push(sg(CF), UR)
push(sg(CF), WR)
UF
push(UF, sg(CR))
push(UF, UR)
pull(sg(UF), WR)
WF
–
pull(sg(WF), UR)
pull(sg(WF), WR)
Table 3. Local contrastive rules for fake–real patch pairs.
Figure 2. Internal-validation examples. From top to bottom: pseudo boxes, ground-truth simple explanations, Artifact Evidence Maps, and generated simple explanations. Eleven examples showing grounded pseudo boxes, reference simple explanations, detector evidence heatmaps, and generated simple explanations.
Variant
λimg
λart
λlcl
Val. AUC
Image-level detector only
15
0
0
0.8975
+ pseudo-mask artifact supervision
15
1
0
0.9139
+ LCL
15
1
1
0.9417
Table 4. Detector ablation on the public validation split.
Method
Bc
Fent
Fevid
Bs
Ls
Ss
Mexp
Zero-shot Qwen3-VL-253B-A22B-Instruct
0.6023
0.5310
0.3934
0.4495
0.3418
0.4172
0.4812
Organizer baseline 1 1 1 The organizer baseline model can be found on the official challenge website: https://explainable-deepfake-detection.github.io/
0.7109
0.5536
0.4620
0.4578
0.4264
0.4484
0.5365
Ours
0.7253
0.6235
0.5447
0.4991
0.9703
0.6405
0.6236
Table 5. Explanation generation evaluation on the internal held-out validation subset. Bc and Bs denote complex and simple BERTScore-F1; Fent and Fevid denote Entity and Evidence F1; Ls , Ss , and Mexp denote normalized simple SLE, Simple score, and Explanation score.
Metric
Score
Detection accuracy
0.9349
Detection fake F1
0.9418
Detection macro F1
0.9340
Detection real F1
0.9261
Complex BERTScore-F1
0.7004
Simple BERTScore-F1
0.6509
Table 6. Final submission results on the full final test split.