Multimodal large language models (MLLMs) can explain deepfake verdicts in natural language, but such explanations are not necessarily visually grounded in the visual evidence underlying the prediction. A model may describe plausible artifacts inferred from language priors rather than from image evidence. Existing grounding methods improve visual reliance through decoding or attention interventions, but they generally strengthen grounding over the entire image, making them ill-suited for forensic artifacts that are subtle, spatially localized, and image-dependent. We propose Look Before You Judge, a training-free framework that formulates explainable deepfake detection as a sequential evidence acquisition process. Instead of directly predicting image authenticity from holistic visual reasoning, our framework first identifies image-specific candidate evidence regions by contrasting the MLLM's decoder-to-visual attention between an original image and its Gaussian-blurred counterpart. The identified regions are then inspected individually, and the resulting local evidence is integrated with the global image context before reaching a final verdict. The framework operates without manipulation masks, external forensic models, or parameter updates, making it directly applicable to off-the-shelf MLLMs. Across five open-source MLLMs on TriDF and MMTD-Set, our framework improves detection accuracy by up to 12.8%, reduces CHAIR by up to 33.4% and hallucination rate by up to 21.3%, and outperforms representative training-free decoding and attention methods.
Figures & tables
Figure 1: Look Before You Judge. A general MLLM produces an ungrounded explanation and an incorrect authentic verdict. Our framework instead inspects localized candidate regions and aggregates the resulting evidence to reach the correct manipulated verdict.
Figure 2: Overview of Look Before You Judge . Stage 1 derives a blur-sensitive visual prior by contrasting output-to-visual attention between the input image and its blurred counterpart. Stage 2 constructs candidate forensic regions from high-response tokens. Stage 3 performs region-wise inspection with attention steering and aggregates the filtered evidence to produce the final prediction and explanation. The entire framework is training-free and requires only frozen MLLMs.
Method
TriDF
MMTD-Set
ACC ↑
Cover ↑
CHAIR ↓
Hal ↓
F0.5↑
DeepFake
AIGC-Editing
ACC ↑
F1 ↑
ACC ↑
F1 ↑
Open-source models
InternVL-3.5-8B
0.4176
0.0270
0.9745
1.0000
0.0296
0.5235
0.4936
0.5200
0.4419
InternVL-3.5-8B + Ours
0.5458
0.2239
0.6407
0.7875
0.2564
0.5591
0.6278
0.5260
0.5917
InternVL-3.5-14B
0.4213
0.0395
0.9763
0.9980
0.0217
0.5300
0.5488
0.5105
0.4872
Table 1: Overall Quantitative Comparison. Results across five open-source MLLMs on TriDF and MMTD-Set, with commercial MLLMs reported as performance references.
Method
TriDF
ACC ↑
Cover ↑
CHAIR ↓
Hal ↓
F0.5↑
Vanilla
0.4176
0.0270
0.9745
1.0000
0.0296
+ VCD
0.4241
0.0372
0.9777
1.0000
0.0212
+ AttnReal
0.4206
0.0719
0.9750
0.9993
0.0246
+ Ours
0.5458
0.2239
0.6407
0.7875
0.2564
Table 2: Comparison with Training-Free Grounding. Results on TriDF using InternVL-3.5-8B.
Method
TriDF
ACC ↑
Cover ↑
CHAIR ↓
Hal ↓
F0.5↑
Vanilla
0.4176
0.0270
0.9745
1.0000
0.0296
Prompt-only
0.5113
0.1369
0.8545
0.9927
0.1340
w/o blur contrast
0.4867
0.1920
0.8600
0.9568
0.1132
w/o component steering
0.4884
0.1890
0.8574
0.9548
0.1182
Ours
0.5458
0.2239
0.6407
0.7875
0.2564
Table 3: Ablation Study. Effects of evidence-before-judgment prompting, blur-contrastive region mining, and component-wise steering on TriDF.
Figure 3: Qualitative Spatial Grounding. Candidate regions discovered by Stages 1 and 2 align with annotated tampered areas without localization supervision.
Figure 4: From Global Error to Local Evidence. Top: Vanilla inference overlooks the annotated blending and texture artifacts and incorrectly predicts the image as authentic. Bottom: Our framework discovers three candidate regions, inspects them independently, and aggregates the retained local evidence under a global view to produce the correct manipulated verdict.
Method
TriDF
ACC ↑
Cover ↑
CHAIR ↓
Hal ↓
F0.5↑
Vanilla
0.4176
0.0270
0.9745
1.0000
0.0296
Prompt-only
0.5113
0.1369
0.8545
0.9927
0.1340
Self-Proposed Region Inspection
0.5076
0.0919
0.8870
0.9495
0.0836
Ours
0.5458
0.2239
0.6407
0.7875
0.2564
Table S1: Comparison with Self-Proposed Region Inspection. Results on TriDF using InternVL-3.5-8B.
Method
TriDF
ACC ↑
Cover ↑
CHAIR ↓
Hal ↓
F0.5↑
Vanilla
0.4176
0.0270
0.9745
1.0000
0.0296
+ Facial Regions
0.5176
0.2096
0.6727
0.8014
0.2343
+ Ours
0.5458
0.2239
0.6407
0.7875
0.2564
Table S2: Comparison with Alternative Region Selection Method. Results on TriDF using InternVL-3.5-8B.
Figure S1: Qualitative result on a manipulated image generation example. Vanilla inference overlooks the localized artifacts and predicts the image as authentic. Our framework identifies blur and texture/color inconsistencies in the selected regions and produces the correct manipulated verdict.
Figure S2: Qualitative result on an authentic face. Vanilla inference hallucinates facial artifacts and produces a false positive. Our framework assigns only weak, uncertain, or negative evidence to the inspected regions and recovers the correct authentic verdict.
Figure S3: Qualitative comparison when both methods correctly predict manipulation. Vanilla inference provides generic and weakly supported claims, whereas our framework grounds its explanation in blending and texture irregularities around the hairline and forehead.
Figure S4: Representative failure case on a manipulated outdoor scene. Our framework identifies a relevant background-blur cue, but the cue is also compatible with natural focus variation and lacks sufficient complementary evidence. Both methods therefore predict the image as authentic.
Method
Inference Efficiency
Avg. Invoc. ↓
Time (s/img) ↓
Rel. Latency ↓
Peak GPU Mem. (GB) ↓
Vanilla
1.000
13.480
1.000 ×
16.129
Ours
6.970
19.951
1.480 ×
16.624
Table S3: Inference efficiency comparison. inference efficiency on a balanced 100-image subset of TriDF. Both methods use the same InternVL-3.5-8B backbone under identical single-GPU settings. A model invocation denotes either an autoregressive generation or an explicit attention forward evaluation.
Department of Networks and Digital Media, Kingston University London, UK · Centre for Research and Technology Hellas (CERTH), Thessaloniki, Greece · Department of Computer Science, Kingston University London, UK