GIFTBench: Diagnosing Generalization in Image Forgery Localization and Informing Model Design
Abstract
Reliable evaluation of image forgery localization (IFL) requires assessing models under diverse distribution changes, yet existing benchmarks often cover limited manipulation conditions or entangle multiple factors in cross-dataset evaluation. Consequently, aggregate performance provides an incomplete view of localization generalization. We introduce GIFTBench, a multi-axis benchmark of 115,013 manipulated images with pixel-level annotations spanning manipulation source, semantic target, editing operation, and composition complexity. GIFTBench supports axis-specific transfer analysis and evaluation on twelve external datasets. Its diagnostic studies reveal asymmetric cross-source transfer, recall-dominated failures, and heterogeneous degradation across semantic, operational, and compositional changes. Beyond diagnosis, the scale and diversity of GIFTBench provide a substantially broader training distribution than conventional IFL datasets. Training representative localizers on GIFTBench consistently improves their aggregate transfer to external datasets, showing that the benchmark serves not only as an evaluation tool but also as an effective training resource for cross-domain localization. Guided by the diagnostic findings, we further develop ForenScope, a detection and localization framework combining classification-adapted representations with multi-depth, multi-scale spatial features, learned layer fusion, and selective coarse-scale conditioning. Experiments show improved cross-dataset localization while retaining image-level detection capability. The GIFTBench dataset showcase page is available at https://giftbench-preview.doudoudouya337.chatgpt.site.
Figures & tables
| Benchmark | Scale | Classical | AI-based | Semantic Target | Operation | Composition |
| CASIA v2 ( Dong and others, 2013 ) | 12.6K | ✓ | ||||
| DEFACTO ( Mahfoudi et al., 2019 ) | 159K | ✓ | ||||
| IMD2020 ( Novozamsky and others, 2020 ) | 40K | ✓ | ||||
| CocoGlide ( Guillaro et al., 2023 ) | 20K | ✓ | ||||
| AutoSplice ( Jia et al., 2023 ) | 110K | ✓ | ||||
| TGIF ( Mareen et al., 2024 ) | 60K | ✓ |
| Training Protocol | Model | IFL Dataset | Related Dataset | Average | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CASIA v1 | NIST16 | IMD2020 | Coverage | Columbia | MISD | AutoSplice | CocoGlide | Br-Gen | DDL | T-IC13 | OSTF | Avg. (All) | Avg. (Excl. IMD) | ||
| Protocol-MVSS | PSCC-Net | .375 | .184 | .259 | .239 | .605 | .690 | .659 | .285 | .267 | .358 | .254 | .190 | .3638 | – |
| TruFor | .721 | .321 | .323 | .424 | .865 | .756 | .320 | .204 | .085 | .486 | .364 | .230 | .4249 | – | |
| IML-ViT | .718 | .291 | .322 | .438 | .747 | .710 | .221 | .211 | .079 | .312 | .370 | .239 | .3882 | – | |
| Protocol-CAT | PSCC-Net | .569 | .344 | – | .386 | .856 | .763 | .564 | .509 | .133 | .348 | .173 | .240 | – | .4441 |
| TruFor | .821 | .301 | – | .483 | .876 | .772 | .327 | .283 | .078 | .596 | .259 | .218 | – | .4558 | |
| GIFTBench training | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| In-domain GIFTBench | Out-of-domain datasets | ||||||||||||
| Method | Single | Fusion | Multi | CASIAv1 | NIST16 | IMD2020 | COVERAGE | Columbia | MISD | AutoSplice | CocoGlide | Br-Gen | Avg |
| PSCC-Net | .610 | .538 | .658 | .320 | .189 | .271 | .339 | .649 | .641 | .856 | .567 | .563 | .4883 |
| TruFor | .837 | .816 | .819 | .642 | .164 | .490 | .344 | .801 | .763 | .917 | .603 | .574 | .5887 |
| IML-ViT | .730 | .656 | .793 | .628 | .251 | .552 | .241 | .742 | .749 | .868 | .306 | .568 | .5450 |
| SparseViT | .862 | .796 | .850 | .633 | .200 | .543 | .430 | .813 | .763 | .923 | .698 | .643 | .6273 |
| Variant | ADE Adapt. | Multi-depth | Multi-scale | CLS-FiLM | F1 | IoU |
|---|---|---|---|---|---|---|
| Retrained structural variants (direct output) | ||||||
| w/o ADE adaptation | ✓ | ✓ | ✓ | .632 | .535 | |
| F17 only | ✓ | ✓ | ✓ | .619 | .523 | |
| H37 only | ✓ | ✓ | ✓ | .641 | .544 | |
| H74 only | ✓ | ✓ | N/A | .644 | .540 | |
| w/o CLS-FiLM | ✓ | ✓ | ✓ | .659 | .567 | |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Axis | Labels and meaning |
|---|---|
| Manipulation source | Splicing, Copy–Move, Removal, and AI-Edit. This identifies the family of editing pipeline that produced the sample. |
| Semantic Target | Object, Component, and Background. This identifies the semantic target category or spatial scale of the manipulated content. |
| Editing operation | Add, Remove, and Replace. This identifies the intended editing objective independently of the source family. |
| Composition complexity | Single, Fusion, and Multi-Source. This identifies whether the edit is isolated, boundary-refined, or composed with other manipulation sources. |
| Release partition | Train | Test |
|---|---|---|
| GIFTBench-Core (single-operation) | ||
| Splicing | 20,673 | 2,327 |
| Copy–Move (including attribute) | 20,887 | 2,386 |
| Removal | 21,074 | 2,380 |
| AI-Edit | 25,758 | 2,851 |
| Core subtotal | 88,392 | 9,944 |
| Setting | Training source | Evaluation partition | Diagnostic question |
| Axis-specific GIFTBench protocols | |||
| Protocol I | One source group | All four source groups | Source transfer: does evidence transfer across manipulation families? |
| Protocol II | One semantic target group | All three semantic target groups | Semantic-target transfer: does behavior survive changes in target category? |
| Protocol III | One operation group | All three operation groups | Operation transfer: does evidence transfer across editing intent? |
| Protocol IV | Single-operation training | Single, Fusion, and Multi-Source groups | Composition: does isolated-edit evidence survive composition? |
| Cross-domain transfer settings | |||
| Benchmark observation | Evidence in benchmark analysis | Design requirement / hypothesis | Architectural realization |
|---|---|---|---|
| Source-transfer asymmetry | Protocol I matrices show strong dependence on the manipulation source | Avoid a single source-sensitive representation; test whether complementary evidence reduces this specialization | Multi-depth patch features |
| Transfer profiles vary across inputs and groups | Aggregate profiles are not uniform across examples or conditions | Combine complementary depth evidence with scale-specific learnable aggregation | Learnable scale-wise LayerMixer |
| Composition creates heterogeneous spatial evidence | Fusion/Multi-Source results and attention analysis show changes in extent, boundaries, and evidence distribution | Represent coarse extent and fine boundaries at multiple resolutions | H19/H37/H74 coarse-to-fine pathway |
| Recall-dominated errors under source shift | False-negative decomposition shows that unfamiliar local evidence is often missed | Test whether image-level manipulation context can complement ambiguous local evidence | Bounded CLS-conditioned FiLM on H19/H37 |
| Resolution-specific architectural choice | The benchmark does not identify a unique conditioning location | Preserve a high-resolution local pathway while conditioning broader-context features | No FiLM on H74 |
| Training Protocol | Model | Manipulation Source | Composition | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Splicing | Copy–Move | Removal | AI–Edit | Single | Fusion | Multi–Source | |||||||||
| F1 | IoU | F1 | IoU | F1 | IoU | F1 | IoU | F1 | IoU | F1 | IoU | F1 | IoU | ||
| Protocol-MVSS | CAT-Net | .376 | .293 | .239 | .171 | .128 | .084 | .151 | .106 | .128 | .084 | .151 | .106 | .219 | .160 |
| ObjectFormer | .170 | .106 | .077 | .043 | .070 | .041 | .168 | .105 | .070 | .041 | .168 | .105 | .174 | .103 | |
| PSCC-Net | .295 | .209 | .161 | .103 | .074 | .045 | .143 | .094 | .074 | .045 | .143 | .094 | .233 | .149 | |
| MVSS-Net | .307 | .224 | .219 | .152 | .105 | .067 | .142 | .095 | .105 | .067 | .142 | .095 | .286 | .194 | |
| Training Set | Model | Splicing | Copy-Move | Removal | AI-Edit | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| F1 | IoU | AUC | F1 | IoU | AUC | F1 | IoU | AUC | F1 | IoU | AUC | ||
| Splicing | PSCC-Net | 0.5143 | 0.4149 | 0.9022 | 0.2333 | 0.1590 | 0.8254 | 0.1723 | 0.1184 | 0.7189 | 0.2258 | 0.1686 | 0.6948 |
| TruFor | 0.9024 | 0.8524 | 0.9951 | 0.5732 | 0.4961 | 0.9440 | 0.4178 | 0.3565 | 0.8537 | 0.4870 | 0.4183 | 0.8925 | |
| IML-ViT | 0.8753 | 0.8443 | 0.9262 | 0.5931 | 0.5275 | 0.8454 | 0.3456 | 0.2959 | 0.7506 | 0.3500 | 0.2962 | 0.7655 | |
| Mesorch | 0.9155 | 0.8733 | 0.9935 | 0.5380 | 0.4795 | 0.9293 | 0.3130 | 0.2683 | 0.8054 | 0.2269 | 0.1896 | 0.7679 | |
| SparseViT | 0.9187 | 0.8703 | 0.9972 | 0.6730 | 0.5996 | 0.9717 | 0.4392 | 0.3791 | 0.8663 | 0.5964 | 0.5308 | 0.9247 | |
| Training Set | Model | Object | Component | Background | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| F1 | IoU | AUC | F1 | IoU | AUC | F1 | IoU | AUC | ||
| Object | PSCC-Net | 0.4452 | 0.3530 | 0.9078 | 0.3887 | 0.2895 | 0.8970 | 0.4244 | 0.3330 | 0.8357 |
| TruFor | 0.7161 | 0.6466 | 0.9723 | 0.6683 | 0.5794 | 0.9638 | 0.7125 | 0.6440 | 0.9484 | |
| IML-ViT | 0.6497 | 0.5948 | 0.8676 | 0.5595 | 0.4882 | 0.8452 | 0.6110 | 0.5481 | 0.8491 | |
| Mesorch | 0.6478 | 0.5807 | 0.9575 | 0.5266 | 0.4465 | 0.9280 | 0.5345 | 0.4677 | 0.9009 | |
| SparseViT | 0.7793 | 0.7077 | 0.9861 | 0.7019 | 0.6119 | 0.9733 | 0.7551 | 0.6885 | 0.9665 | |
| Training Set | Model | Add | Remove | Replace | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| F1 | IoU | AUC | F1 | IoU | AUC | F1 | IoU | AUC | ||
| Add | PSCC-Net | 0.5362 | 0.4318 | 0.9407 | 0.2908 | 0.2139 | 0.8105 | 0.3581 | 0.2709 | 0.8263 |
| TruFor | 0.8861 | 0.8316 | 0.9951 | 0.5083 | 0.4422 | 0.8842 | 0.7327 | 0.6630 | 0.9616 | |
| IML-ViT | 0.7332 | 0.6820 | 0.8741 | 0.3151 | 0.2651 | 0.7519 | 0.5335 | 0.4695 | 0.8559 | |
| Mesorch | 0.8551 | 0.7998 | 0.9883 | 0.3901 | 0.3309 | 0.8408 | 0.5927 | 0.5215 | 0.9282 | |
| SparseViT | 0.8878 | 0.8309 | 0.9956 | 0.4947 | 0.4301 | 0.8826 | 0.7550 | 0.6874 | 0.9683 | |
| Model | Single Fusion | Fusion | Single Multi | Multi-Source | ||
|---|---|---|---|---|---|---|
| PSCC-Net | .611/.511 | .538/.442 | .577/.480 | .486/.394 | -.073/-.069 | -.091/-.086 |
| TruFor | .884/.835 | .814/.729 | .817/.755 | .670/.609 | -.070/-.106 | -.147/-.146 |
| IML-ViT | .673/.649 | .657/.613 | .684/.634 | .606/.558 | -.016/-.036 | -.078/-.076 |
| Mesorch | .794/.719 | .733/.631 | .740/.657 | .707/.613 | -.061/-.088 | -.033/-.044 |
| SparseViT | .908/.862 | .793/.713 | .845/.781 | .722/.653 | -.115/-.149 | -.123/-.128 |
| Condition | Perturbation |
|---|---|
| JPEG | Re-encode the image at the evaluated JPEG quality levels |
| Gaussian blur | Apply the evaluated blur-strength levels to the image |
| Resize–JPEG | Resize before JPEG re-encoding |
| Double-JPEG | Apply two successive JPEG compression stages |
| Variant | Encoder / image-level branch | Trainable localization modules | FiLM | Prototype |
|---|---|---|---|---|
| ForenScope | ADE-adapted DINOv2; frozen image-level branch | Multi-depth projections, LayerMixers, decoder, auxiliary/output heads | On | Disabled in default direct output |
| ForenScope + Proto. | Same trained network; frozen image-level branch | Same trained localization branch | On | Enabled at inference only |
| Existing No-CLS | ADE-adapted encoder; frozen image-level branch | Retrained multi-depth branch without CLS-FiLM | Off | Disabled |
| H37-only | ADE-adapted encoder; frozen image-level branch | Retrained H37-only localization variant | Off in effective path | Disabled |
| w/o CLS-FiLM & ADE adaptation | Original DINOv2; no ADE adaptation | Retrained localization branch without CLS-FiLM | Off | Disabled |
| Primary-6 | Extra-6 | Summary | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | CASIA | CocoGlide | Columbia | IMD2020 | NIST16 | COVERAGE | MISD | AutoSplice | BRGen | SceneText | OSTF | DDL | Primary-6 | Extra-6 | All-12 |
| ForenScope | .679/.598 | .799/.697 | .546/.452 | .579/.488 | .548/.469 | .578/.472 | .696/.560 | .871/.789 | .746/.665 | .703/.596 | .595/.476 | .520/.412 | .621/.529 | .688/.583 | .655/.556 |
| Existing No-CLS | .660/.582 | .762/.657 | .534/.442 | .572/.482 | .516/.438 | .581/.467 | .694/.562 | .897/.827 | .759/.679 | .724/.617 | .615/.497 | .486/.386 | .604/.511 | .696/.595 | .650/.553 |
| H37-only | .658/.577 | .676/.564 | .495/.401 | .570/.477 | .547/.461 | .494/.384 | .688/.553 | .889/.816 | .748/.666 | .691/.592 | .578/.465 | .527/.416 | .573/.477 | .687/.584 | .630/.531 |
| Joint control | .567/.490 | .593/.481 | .601/.499 | .550/.456 | .500/.420 | .424/.313 | .686/.550 | .897/.831 | .748/.661 | .685/.580 | .574/.459 | .463/.363 | .539/.443 | .676/.574 | .607/.509 |
| ForenScope + Proto. | .695/.615 | .820/.724 | .564/.472 | .580/.487 | .566/.485 | .592/.487 | .726/.597 | .873/.793 | .753/.671 | .705/.598 | .600/.478 | .532/.425 | .636/.545 | .698/.594 | .667/.569 |
| Training Protocol | Model | CASIA v1 | CocoGlide | Columbia | IMD2020 | NIST16 | COVERAGE | MISD | AutoSplice | Avg. | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AUC | F1 | AUC | F1 | AUC | F1 | AUC | F1 | AUC | F1 | AUC | F1 | AUC | F1 | AUC | F1 | AUC | F1 | ||
| Protocol-MVSS | MVSS-Net | .639 | .697 | .507 | .667 | .468 | .663 | .564 | .907 | .500 | .943 | .519 | .667 | .704 | .489 | .543 | .761 | .556 | .724 |
| PSCC-Net | .743 | .521 | .440 | .639 | .806 | .732 | .629 | .898 | .547 | .887 | .495 | .662 | .994 | .869 | .827 | .822 | .685 | .754 | |
| TruFor | .148 | .392 | .446 | .625 | .049 | .100 | .392 | .796 | .425 | .819 | .352 | .592 | .034 | .075 | .459 | .646 | .288 | .506 | |
| ForenScope | .875 | .732 | .702 | .675 | .819 | .731 | .727 | .815 | .563 | .688 | .551 | .418 | .892 | .779 | .691 | .400 | .727 | .655 | |
| GIFTBench | ForenScope | .902 | .844 | .989 | .963 | .994 | .967 | .899 | .834 | .837 | .705 | .891 | .809 | .993 | .911 | .977 | .934 | .935 | .871 |
| Variant | Definition |
|---|---|
| ForenScope | Full multi-depth network with learned LayerMixer and CLS-FiLM; direct output only. |
| Existing No-CLS | ADE-adapted encoder retained, but the localization branch is retrained without CLS-FiLM. |
| H37-only | H37 is the only effective spatial scale; the remaining fusion topology is retained. |
| w/o ADE adaptation | Original DINOv2 encoder with the full multi-scale, CLS-conditioned localization architecture. |
| Joint control | Original DINOv2 initialization, no ADE adaptation, and no CLS-FiLM; retrained jointly. |
| Equal fixed mixer | Fixed-checkpoint intervention replacing learned four-depth weights by equal weights. |
| Variant | In-domain | OOD |
|---|---|---|
| ForenScope | .886/.823 | .655/.556 |
| Existing No-CLS | .881/.816 | .650/.553 |
| H37-only | .871/.801 | .630/.531 |
| w/o ADE adaptation | – | .618/.519 |
| Joint control | .873/.805 | .607/.509 |
| Reference checkpoint (EMA) | Inference condition | F1/IoU |
|---|---|---|
| ForenScope (Static) | Equal layer weights | .653/.557 |
| Existing No-CLS | Learned layer fusion | .650/.553 |
| Existing No-CLS | Equal layer weights | .553/.450 |
| Original Full-Late | Normal CLS-FiLM | .643/.548 |
| Original Full-Late | FiLM off | .622/.518 |
| Original Full-Late | CLS shuffled | .633/.536 |