Organizations: CCDS Nanyang Technological University Singapore, Singapore · Ant Digital Technologies Ant Group Singapore, Singapore · Ant Digital Technologies Ant Group Hangzhou, China · Singapore University of Technology and Design Singapore, Singapore · IAIC Agency for Science, Technology and Research (A*STAR) Singapore, Singapore · Singapore Management University Singapore, Singapore
Conventional image forgery detection methods produce binary scores or pixel-level masks without interpretable evidence, while recent multimodal large language model (MLLM)-based approaches generate post-hoc explanations of predetermined classification results rather than reasoning from evidence. Inspired by the forensic workflow of human judicial experts, we propose Agentic Tool-Augmented Reasoning (ATAR), a framework integrating 22 specialized forensic tools across seven complementary domains to autonomously detect, localize, and explain image forgeries through multi-turn reasoning. A Dual-Stream Forensic Reasoning paradigm combines a high-level semantic anomaly path, which magnifies suspicious regions for fine-grained inspection, with a low-level forgery artifact path, which invokes forensic tools to extract objective evidence. We further introduce Forensics Curriculum Learning: during General Experience SFT, an automated teacher-student mentoring pipeline synthesizes multi-turn tool-usage reasoning trajectories; during Forensic Scene RL, a Tool Prior Curriculum guides early tool exploration and progressively transfers control to the agent, while a Structured Evidence Reward provides fine-grained process-level supervision. Experiments across IMDL, Deepfake detection, DMDL, and AIGC detection show that ATAR achieves 78.5% average image-level F1 on six zero-shot IMDL benchmarks, surpassing the strongest MLLM baseline by 11.8 percentage points, and remains competitive with specialized detectors on other tasks while producing substantially more faithful and grounded explanations.
Figures & tables
Figure 1. Three forgery detection paradigms. (a) Conventional IMDL methods produce pixel masks but lack logical reasoning and interpretable evidence. (b) MLLM-based methods generate post-hoc verbalizations of pre-determined detection results rather than genuine reasoning. (c) ATAR autonomously invokes forensic tools and reasons from their outputs to produce grounded conclusions. Comparison of three forgery detection paradigms.
Figure 2. Overview of ATAR. Top: multi-turn reasoning chain. Bottom: two-stage training pipeline (Stage 1: General Experience SFT; Stage 2: Forensic Scene RL). Overview of the ATAR method showing the reasoning chain and two-stage training pipeline.
Figure 3. General Experience trajectory synthesis pipeline. Left: instance-level annotation. Right: per-sample tool evaluation. Bottom: teacher-guided trajectory generation, where the student explores without ground truth and the teacher provides corrective guidance. General Experience SFT data synthesis pipeline with annotation, tool evaluation, and trajectory generation.
IMDL
Deepfake
DMDL
AutoSplice
CASIA2
Fant.Reality
IMD2020
NeXT-rpl.
OpenForen.
DocTamper
Compress. (4)
75.7
33.6
6.7
15.0
2.4
40.7
1.2
Noise (4)
7.7
18.6
29.7
22.5
30.6
7.7
39.8
Freq. (3)
0.8
3.9
16.9
7.4
7.9
2.9
3.9
Struct. (3)
1.1
6.8
9.4
7.3
13.9
6.0
4.2
Copy-Move (1)
0.9
8.8
5.2
5.7
7.6
9.2
16.8
Table 1. Domain-level optimal tool selection proportion (%) across training datasets. Per-tool breakdown is provided in the appendix due to space constraints.
CASIA v1+
CocoGlide
Coverage
Korus
Columbia
NIST16
Avg
Method
iF1
ACC
pF1
IoU
iF1
ACC
pF1
IoU
iF1
ACC
pF1
IoU
iF1
ACC
pF1
IoU
iF1
ACC
pF1
IoU
iF1
ACC
pF1
IoU
iF1
ACC
pF1
IoU
MVSS-Net ( Chen et al., 2021 ) ( ICCV’21 )
.734
.757
.471
.397
.665
.560
.492
.372
.669
.550
.467
.373
.632
.539
.181
.120
.749
.672
.670
.576
.682
.538
.299
.216
.689
.603
.430
.342
TruFor ( Guillaro et al., 2023 ) ( CVPR’23 )
.732
.652
.544
.464
.685
.576
.516
.405
.681
.570
.450
.346
.679
.559
.292
.217
.975
.975
.833
.769
.620
.466
.391
.297
.729
.633
.504
.416
ForMa ( Guo et al., 2025 ) ( SPL’25 )
.781
.769
.629
.574
.668
.502
.453
.362
.669
.525
.487
.409
.670
.509
.304
.235
.683
.650
.001
.000
.570
.487
.055
.030
.674
.574
.322
.268
SparseViT ( Su et al., 2025b ) ( AAAI’25 )
.697
.535
.153
.088
.667
.500
.355
.253
.667
.500
.197
.112
.667
.500
.103
.057
.607
.435
.000
.000
.563
.392
.552
.519
.645
.477
.227
.172
FakeShield ( Xu et al., 2025 ) ( ICLR’25 )
.891
.883
.594
.538
.669
.506
.521
.428
.619
.525
.242
.212
.467
.559
.125
.102
.920
.920
.744
.658
.438
.554
.260
.228
.667
.658
.414
.361
Table 2. IMDL benchmark results on six zero-shot datasets. iF1/ACC measure image-level classification; pF1/IoU measure pixel-level localization. Methods above the first rule are specialized detectors; below are MLLM-based and general-purpose MLLMs. Bold = best, underline = second.
T-SROIE
FSTS-1.5k
Avg
Method
pF1
IoU
pF1
IoU
pF1
IoU
MVSS-Net ( Chen et al., 2021 ) ( ICCV’21 )
.024
.012
.124
.077
.074
.045
TruFor ( Guillaro et al., 2023 ) ( CVPR’23 )
.061
.037
.406
.317
.234
.177
PSCC-Net ( Liu et al., 2022 ) ( TCSVT’22 )
.019
.009
.346
.264
.183
.137
APSC-Net ( Qu et al., 2024 ) ( CVPR’24 )
.186
.126
.231
.169
.209
.148
ForMa ( SPL’25 )
.037
.021
.189
.135
.113
.078
Table 3. Document manipulation detection and localization on T-SROIE and FSTS-1.5k (zero-shot, localization only). Bold = best, underline = second.
Advances in generative AI have made image falsification highly realistic, demanding trustworthy authentication systems. Existing forensic detectors can target certain forgery types but lack interpretability, while vision-language models (VLMs) provide explanations but cannot exploit forensic traces for reliable detection. We propose Forensic Knowledge Graphs (FKGs), a unified framework that integrates forensic evidence extraction, structured reasoning, and human-interpretable explanation. Our FKG structure encodes forensic traces along with their causal dependencies and links to scene content. To generate accurate FKGs, we introduce a novel forensic authentication network and an Iterative Context Refinement strategy that guides VLMs to produce faithful, grounded explanations. We also present FKG-50K, a dataset of 50,000 realistic forgeries with ground-truth FKGs. Experiments demonstrate that FKG outperforms both forensic detectors and VLMs in detection, forgery identification and localization, and forensic justification.
Forensic deepfake analysis demands more than binary classification: investigators need region-grounded natural language explanations they can verify against the image. Multimodal large language models (MLLMs) are a natural fit, but pretrained MLLMs fail systematically, producing globally coherent text that misses the small localized cues defining manipulations. We argue this is an inductive bias problem rather than a capacity issue: the image-text contrastive objective training MLLM visual encoders optimizes for whole-image semantic summaries, not patch-level forensic detail. The same mismatch explains why prior deepfake reasoning methods target either face manipulation or fully AI-generated content, never both. We propose FORGE, which addresses the mismatch by routing a second visual stream into the language model from a Vision-Only Model (VOM) trained on dense patch prediction rather than image-text alignment. The MLLM's native encoder and the VOM operate on a shared patch grid, which lets us interleave their tokens with preserved spatial correspondence; we show this beats naive concatenation. A two-stage adapter training protocol (generic image-caption alignment, then joint task-specific optimization) prevents the localized stream from overfitting to training-domain manipulations. Across face-manipulated and fully synthetic content, FORGE produces region-referential explanations answering fine-grained attribute queries ("Does the eyes/nose/mouth look real or fake?") and substantially outperforms in-domain baselines on cross-domain evaluations; region-specific evaluation and human studies confirm explanation faithfulness.
Existing vision-language forgery detection and grounding methods operate under a closed-world paradigm, assuming verification can be completed by the model alone. However, self-contained MLLMs are constrained by finite parametric knowledge, static training corpora, and limited perceptual resolution, creating a practical ceiling in dynamic open-world forensics -- particularly for real-time event verification requiring external clues and forgery segmentation demanding fine-grained scrutiny of local manipulations. To address these limitations, we shift from scaling up the self-contained model toward reaching beyond it. We propose \textbf{OmniVL-Guard Pro}, a tool-augmented agent that extends unified forensics from closed-world prediction to open-world clues-driven reasoning. OmniVL-Guard Pro integrates a tool environment spanning real-time event search, local cropping and zooming, edge-anomaly screening, face detection, video frame extraction, and SAM3-based segmentation. To generate high-quality tool-reasoning trajectories, we introduce \textbf{Tree-Structured Self-Evolving Tool Trajectory Generation}, which produces diverse trajectories through seed guidance, guider-free self-evolution, and weakly-hinted hard sample synthesis, yielding the Full-Spectrum Tool Reasoning (FSTR) dataset for training. We further propose \textbf{Checker-Guided Agentic Reinforcement Learning} (CGARL), which provides process-level supervision to penalize cases where the answer is correct but the reasoning is distorted. Extensive experiments demonstrate that OmniVL-Guard Pro achieves state-of-the-art performance across various tasks, and exhibits strong zero-shot generalization. The FSTR dataset and code for OmniVL-Guard Pro will be publicly released at https://github.com/shen8424/OmniVL-Guard-Pro.
Jinjie Shen, Zheng Huang, Yuchen Zhang +7
School of Computer Science and Information Engineering, Hefei University of Technology, Hefei, China · Wuhan University, Wuhan, China · Xi’an Jiaotong University +2