cs.CVSep 28, 2026

Look Before You Judge: Training-Free Region Mining for Grounded and Explainable Deepfake Detection

Authors: Chia-Ling Chen, Yu-Ting Ta, Jian-Yu Jiang-Lin, Tai-Ming Huang, Ling Lo, Po-Ching Chen, Yan-Tsung Wang, Pei-Heng Li, +3 more

Organizations: National Taiwan University · National Tsing Hua University · National Yang Ming Chiao Tung University

Abstract

Multimodal large language models (MLLMs) can explain deepfake verdicts in natural language, but such explanations are not necessarily visually grounded in the visual evidence underlying the prediction. A model may describe plausible artifacts inferred from language priors rather than from image evidence. Existing grounding methods improve visual reliance through decoding or attention interventions, but they generally strengthen grounding over the entire image, making them ill-suited for forensic artifacts that are subtle, spatially localized, and image-dependent. We propose Look Before You Judge, a training-free framework that formulates explainable deepfake detection as a sequential evidence acquisition process. Instead of directly predicting image authenticity from holistic visual reasoning, our framework first identifies image-specific candidate evidence regions by contrasting the MLLM's decoder-to-visual attention between an original image and its Gaussian-blurred counterpart. The identified regions are then inspected individually, and the resulting local evidence is integrated with the global image context before reaching a final verdict. The framework operates without manipulation masks, external forensic models, or parameter updates, making it directly applicable to off-the-shelf MLLMs. Across five open-source MLLMs on TriDF and MMTD-Set, our framework improves detection accuracy by up to 12.8%, reduces CHAIR by up to 33.4% and hallucination rate by up to 21.3%, and outperforms representative training-free decoding and attention methods.

Figures & tables

Explore similar work

CardsList
  1. FORGE: Forensic Reasoning with Grounded Evidence

    Mar 20, 2025Rohit Kundu, Shan Jia, Vishal Mohanty +2Deepfake DetectionDigital Forensics

  2. VIGIL: Part-Grounded Structured Reasoning for Generalizable Deepfake Detection

    Mar 23, 2026Xinghan Li, Junhao Xu, Jingjing ChenDeepfake DetectionDigital Forensics

  3. DF-CBM: Region-Aware Concept Bottleneck Models for Deepfake Detection

    Sep 28, 2026Georgios Tsoumplekas, Vazgken Vanian, Alexandros Doumanoglou +4Deepfake DetectionUnsupervised Detection