Multimodal large language models (MLLMs) can explain deepfake verdicts in natural language, but such explanations are not necessarily visually grounded in the visual evidence underlying the prediction. A model may describe plausible artifacts inferred from language priors rather than from image evidence. Existing grounding methods improve visual reliance through decoding or attention interventions, but they generally strengthen grounding over the entire image, making them ill-suited for forensic artifacts that are subtle, spatially localized, and image-dependent. We propose Look Before You Judge, a training-free framework that formulates explainable deepfake detection as a sequential evidence acquisition process. Instead of directly predicting image authenticity from holistic visual reasoning, our framework first identifies image-specific candidate evidence regions by contrasting the MLLM's decoder-to-visual attention between an original image and its Gaussian-blurred counterpart. The identified regions are then inspected individually, and the resulting local evidence is integrated with the global image context before reaching a final verdict. The framework operates without manipulation masks, external forensic models, or parameter updates, making it directly applicable to off-the-shelf MLLMs. Across five open-source MLLMs on TriDF and MMTD-Set, our framework improves detection accuracy by up to 12.8%, reduces CHAIR by up to 33.4% and hallucination rate by up to 21.3%, and outperforms representative training-free decoding and attention methods.
Figures & tables
Figure 1: Look Before You Judge. A general MLLM produces an ungrounded explanation and an incorrect authentic verdict. Our framework instead inspects localized candidate regions and aggregates the resulting evidence to reach the correct manipulated verdict.
Figure 2: Overview of Look Before You Judge . Stage 1 derives a blur-sensitive visual prior by contrasting output-to-visual attention between the input image and its blurred counterpart. Stage 2 constructs candidate forensic regions from high-response tokens. Stage 3 performs region-wise inspection with attention steering and aggregates the filtered evidence to produce the final prediction and explanation. The entire framework is training-free and requires only frozen MLLMs.
Method
TriDF
MMTD-Set
ACC ↑
Cover ↑
CHAIR ↓
Hal ↓
F0.5↑
DeepFake
AIGC-Editing
ACC ↑
F1 ↑
ACC ↑
F1 ↑
Open-source models
InternVL-3.5-8B
0.4176
0.0270
0.9745
1.0000
0.0296
0.5235
0.4936
0.5200
0.4419
InternVL-3.5-8B + Ours
0.5458
0.2239
0.6407
0.7875
0.2564
0.5591
0.6278
0.5260
0.5917
InternVL-3.5-14B
0.4213
0.0395
0.9763
0.9980
0.0217
0.5300
0.5488
0.5105
0.4872
Table 1: Overall Quantitative Comparison. Results across five open-source MLLMs on TriDF and MMTD-Set, with commercial MLLMs reported as performance references.
Method
TriDF
ACC ↑
Cover ↑
CHAIR ↓
Hal ↓
F0.5↑
Vanilla
0.4176
0.0270
0.9745
1.0000
0.0296
+ VCD
0.4241
0.0372
0.9777
1.0000
0.0212
+ AttnReal
0.4206
0.0719
0.9750
0.9993
0.0246
+ Ours
0.5458
0.2239
0.6407
0.7875
0.2564
Table 2: Comparison with Training-Free Grounding. Results on TriDF using InternVL-3.5-8B.
Method
TriDF
ACC ↑
Cover ↑
CHAIR ↓
Hal ↓
F0.5↑
Vanilla
0.4176
0.0270
0.9745
1.0000
0.0296
Prompt-only
0.5113
0.1369
0.8545
0.9927
0.1340
w/o blur contrast
0.4867
0.1920
0.8600
0.9568
0.1132
w/o component steering
0.4884
0.1890
0.8574
0.9548
0.1182
Ours
0.5458
0.2239
0.6407
0.7875
0.2564
Table 3: Ablation Study. Effects of evidence-before-judgment prompting, blur-contrastive region mining, and component-wise steering on TriDF.
Figure 3: Qualitative Spatial Grounding. Candidate regions discovered by Stages 1 and 2 align with annotated tampered areas without localization supervision.
Figure 4: From Global Error to Local Evidence. Top: Vanilla inference overlooks the annotated blending and texture artifacts and incorrectly predicts the image as authentic. Bottom: Our framework discovers three candidate regions, inspects them independently, and aggregates the retained local evidence under a global view to produce the correct manipulated verdict.
Method
TriDF
ACC ↑
Cover ↑
CHAIR ↓
Hal ↓
F0.5↑
Vanilla
0.4176
0.0270
0.9745
1.0000
0.0296
Prompt-only
0.5113
0.1369
0.8545
0.9927
0.1340
Self-Proposed Region Inspection
0.5076
0.0919
0.8870
0.9495
0.0836
Ours
0.5458
0.2239
0.6407
0.7875
0.2564
Table S1: Comparison with Self-Proposed Region Inspection. Results on TriDF using InternVL-3.5-8B.
Method
TriDF
ACC ↑
Cover ↑
CHAIR ↓
Hal ↓
F0.5↑
Vanilla
0.4176
0.0270
0.9745
1.0000
0.0296
+ Facial Regions
0.5176
0.2096
0.6727
0.8014
0.2343
+ Ours
0.5458
0.2239
0.6407
0.7875
0.2564
Table S2: Comparison with Alternative Region Selection Method. Results on TriDF using InternVL-3.5-8B.
Figure S1: Qualitative result on a manipulated image generation example. Vanilla inference overlooks the localized artifacts and predicts the image as authentic. Our framework identifies blur and texture/color inconsistencies in the selected regions and produces the correct manipulated verdict.
Figure S2: Qualitative result on an authentic face. Vanilla inference hallucinates facial artifacts and produces a false positive. Our framework assigns only weak, uncertain, or negative evidence to the inspected regions and recovers the correct authentic verdict.
Figure S3: Qualitative comparison when both methods correctly predict manipulation. Vanilla inference provides generic and weakly supported claims, whereas our framework grounds its explanation in blending and texture irregularities around the hairline and forehead.
Figure S4: Representative failure case on a manipulated outdoor scene. Our framework identifies a relevant background-blur cue, but the cue is also compatible with natural focus variation and lacks sufficient complementary evidence. Both methods therefore predict the image as authentic.
Method
Inference Efficiency
Avg. Invoc. ↓
Time (s/img) ↓
Rel. Latency ↓
Peak GPU Mem. (GB) ↓
Vanilla
1.000
13.480
1.000 ×
16.129
Ours
6.970
19.951
1.480 ×
16.624
Table S3: Inference efficiency comparison. inference efficiency on a balanced 100-image subset of TriDF. Both methods use the same InternVL-3.5-8B backbone under identical single-GPU settings. A model invocation denotes either an autoregressive generation or an explicit attention forward evaluation.
Forensic deepfake analysis demands more than binary classification: investigators need region-grounded natural language explanations they can verify against the image. Multimodal large language models (MLLMs) are a natural fit, but pretrained MLLMs fail systematically, producing globally coherent text that misses the small localized cues defining manipulations. We argue this is an inductive bias problem rather than a capacity issue: the image-text contrastive objective training MLLM visual encoders optimizes for whole-image semantic summaries, not patch-level forensic detail. The same mismatch explains why prior deepfake reasoning methods target either face manipulation or fully AI-generated content, never both. We propose FORGE, which addresses the mismatch by routing a second visual stream into the language model from a Vision-Only Model (VOM) trained on dense patch prediction rather than image-text alignment. The MLLM's native encoder and the VOM operate on a shared patch grid, which lets us interleave their tokens with preserved spatial correspondence; we show this beats naive concatenation. A two-stage adapter training protocol (generic image-caption alignment, then joint task-specific optimization) prevents the localized stream from overfitting to training-domain manipulations. Across face-manipulated and fully synthetic content, FORGE produces region-referential explanations answering fine-grained attribute queries ("Does the eyes/nose/mouth look real or fake?") and substantially outperforms in-domain baselines on cross-domain evaluations; region-specific evaluation and human studies confirm explanation faithfulness.
Multimodal large language models (MLLMs) offer a promising path toward interpretable deepfake detection by generating textual explanations. However, the reasoning process of current MLLM-based methods combines evidence generation and manipulation localization into a unified step. This combination blurs the boundary between faithful observations and hallucinated explanations, leading to unreliable conclusions. Building on this, we present VIGIL, a part-centric structured forensic framework inspired by expert forensic practice through a plan-then-examine pipeline: the model first plans which facial parts warrant inspection based on global visual cues, then examines each part with independently sourced forensic evidence. A stage-gated injection mechanism delivers part-level forensic evidence only during examination, ensuring that part selection remains driven by the model's own perception rather than biased by external signals. We further propose a progressive three-stage training paradigm whose reinforcement learning stage employs part-aware rewards to enforce anatomical validity and evidence--conclusion coherence. To enable rigorous generalizability evaluation, we construct OmniFake, a hierarchical 5-Level benchmark where the model, trained on only three foundational generators, is progressively tested up to in-the-wild social-media data. Extensive experiments on OmniFake and cross-dataset evaluations demonstrate that VIGIL consistently outperforms both expert detectors and concurrent MLLM-based methods across all generalizability levels.
Xinghan Li, Junhao Xu, Jingjing Chen
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China.
Deepfake detection methods have become increasingly effective yet most provide limited insight into the evidence behind their predictions. However, in forensic settings users also need to know which manipulation cues support the decision and where they appear. Existing explainability methods only partially address this need since localization-based approaches lack semantic descriptions while language-based explanation methods are only weakly grounded in visual evidence. In this work, we propose DF-CBM, a region-aware concept bottleneck model for explainable deepfake detection. DF-CBM builds a compact vocabulary of manipulation-related concepts from textual artifact annotations and links each concept to plausible facial and boundary regions. It then predicts these concepts from visual features using a concept-specific masked attention mechanism guided by parsed facial masks and the final real/fake decision is made from the predicted concept bottleneck. Our experiments show that DF-CBM outperforms concept-based baselines in concept prediction and deepfake classification while remaining competitive with state-of-the-art black-box detectors. Finally, qualitative results and intervention analyses demonstrate that DF-CBM provides spatially grounded concept evidence and enables counterfactual explanations of how individual manipulation concepts influence the final prediction. Our code is available at: https://github.com/GeorgeTsoumplekas/DF-CBM.
Department of Networks and Digital Media, Kingston University London, UK · Centre for Research and Technology Hellas (CERTH), Thessaloniki, Greece · Department of Computer Science, Kingston University London, UK