Organizations: CCDS Nanyang Technological University Singapore, Singapore · Ant Digital Technologies Ant Group Singapore, Singapore · Ant Digital Technologies Ant Group Hangzhou, China · Singapore University of Technology and Design Singapore, Singapore · IAIC Agency for Science, Technology and Research (A*STAR) Singapore, Singapore · Singapore Management University Singapore, Singapore
Conventional image forgery detection methods produce binary scores or pixel-level masks without interpretable evidence, while recent multimodal large language model (MLLM)-based approaches generate post-hoc explanations of predetermined classification results rather than reasoning from evidence. Inspired by the forensic workflow of human judicial experts, we propose Agentic Tool-Augmented Reasoning (ATAR), a framework integrating 22 specialized forensic tools across seven complementary domains to autonomously detect, localize, and explain image forgeries through multi-turn reasoning. A Dual-Stream Forensic Reasoning paradigm combines a high-level semantic anomaly path, which magnifies suspicious regions for fine-grained inspection, with a low-level forgery artifact path, which invokes forensic tools to extract objective evidence. We further introduce Forensics Curriculum Learning: during General Experience SFT, an automated teacher-student mentoring pipeline synthesizes multi-turn tool-usage reasoning trajectories; during Forensic Scene RL, a Tool Prior Curriculum guides early tool exploration and progressively transfers control to the agent, while a Structured Evidence Reward provides fine-grained process-level supervision. Experiments across IMDL, Deepfake detection, DMDL, and AIGC detection show that ATAR achieves 78.5% average image-level F1 on six zero-shot IMDL benchmarks, surpassing the strongest MLLM baseline by 11.8 percentage points, and remains competitive with specialized detectors on other tasks while producing substantially more faithful and grounded explanations.
Figures & tables
Figure 1. Three forgery detection paradigms. (a) Conventional IMDL methods produce pixel masks but lack logical reasoning and interpretable evidence. (b) MLLM-based methods generate post-hoc verbalizations of pre-determined detection results rather than genuine reasoning. (c) ATAR autonomously invokes forensic tools and reasons from their outputs to produce grounded conclusions. Comparison of three forgery detection paradigms.
Figure 2. Overview of ATAR. Top: multi-turn reasoning chain. Bottom: two-stage training pipeline (Stage 1: General Experience SFT; Stage 2: Forensic Scene RL). Overview of the ATAR method showing the reasoning chain and two-stage training pipeline.
Figure 3. General Experience trajectory synthesis pipeline. Left: instance-level annotation. Right: per-sample tool evaluation. Bottom: teacher-guided trajectory generation, where the student explores without ground truth and the teacher provides corrective guidance. General Experience SFT data synthesis pipeline with annotation, tool evaluation, and trajectory generation.
IMDL
Deepfake
DMDL
AutoSplice
CASIA2
Fant.Reality
IMD2020
NeXT-rpl.
OpenForen.
DocTamper
Compress. (4)
75.7
33.6
6.7
15.0
2.4
40.7
1.2
Noise (4)
7.7
18.6
29.7
22.5
30.6
7.7
39.8
Freq. (3)
0.8
3.9
16.9
7.4
7.9
2.9
3.9
Struct. (3)
1.1
6.8
9.4
7.3
13.9
6.0
4.2
Copy-Move (1)
0.9
8.8
5.2
5.7
7.6
9.2
16.8
Table 1. Domain-level optimal tool selection proportion (%) across training datasets. Per-tool breakdown is provided in the appendix due to space constraints.
CASIA v1+
CocoGlide
Coverage
Korus
Columbia
NIST16
Avg
Method
iF1
ACC
pF1
IoU
iF1
ACC
pF1
IoU
iF1
ACC
pF1
IoU
iF1
ACC
pF1
IoU
iF1
ACC
pF1
IoU
iF1
ACC
pF1
IoU
iF1
ACC
pF1
IoU
MVSS-Net ( Chen et al., 2021 ) ( ICCV’21 )
.734
.757
.471
.397
.665
.560
.492
.372
.669
.550
.467
.373
.632
.539
.181
.120
.749
.672
.670
.576
.682
.538
.299
.216
.689
.603
.430
.342
TruFor ( Guillaro et al., 2023 ) ( CVPR’23 )
.732
.652
.544
.464
.685
.576
.516
.405
.681
.570
.450
.346
.679
.559
.292
.217
.975
.975
.833
.769
.620
.466
.391
.297
.729
.633
.504
.416
ForMa ( Guo et al., 2025 ) ( SPL’25 )
.781
.769
.629
.574
.668
.502
.453
.362
.669
.525
.487
.409
.670
.509
.304
.235
.683
.650
.001
.000
.570
.487
.055
.030
.674
.574
.322
.268
SparseViT ( Su et al., 2025b ) ( AAAI’25 )
.697
.535
.153
.088
.667
.500
.355
.253
.667
.500
.197
.112
.667
.500
.103
.057
.607
.435
.000
.000
.563
.392
.552
.519
.645
.477
.227
.172
FakeShield ( Xu et al., 2025 ) ( ICLR’25 )
.891
.883
.594
.538
.669
.506
.521
.428
.619
.525
.242
.212
.467
.559
.125
.102
.920
.920
.744
.658
.438
.554
.260
.228
.667
.658
.414
.361
Table 2. IMDL benchmark results on six zero-shot datasets. iF1/ACC measure image-level classification; pF1/IoU measure pixel-level localization. Methods above the first rule are specialized detectors; below are MLLM-based and general-purpose MLLMs. Bold = best, underline = second.
T-SROIE
FSTS-1.5k
Avg
Method
pF1
IoU
pF1
IoU
pF1
IoU
MVSS-Net ( Chen et al., 2021 ) ( ICCV’21 )
.024
.012
.124
.077
.074
.045
TruFor ( Guillaro et al., 2023 ) ( CVPR’23 )
.061
.037
.406
.317
.234
.177
PSCC-Net ( Liu et al., 2022 ) ( TCSVT’22 )
.019
.009
.346
.264
.183
.137
APSC-Net ( Qu et al., 2024 ) ( CVPR’24 )
.186
.126
.231
.169
.209
.148
ForMa ( SPL’25 )
.037
.021
.189
.135
.113
.078
Table 3. Document manipulation detection and localization on T-SROIE and FSTS-1.5k (zero-shot, localization only). Bold = best, underline = second.
School of Computer Science and Information Engineering, Hefei University of Technology, Hefei, China · Wuhan University, Wuhan, China · Xi’an Jiaotong University +2