Less Supervision, Better Generalization: Weakly Supervised Fake Region Localization in Diffusion-Edited Images
Organizations: Kyung Hee University · NAVER Cloud
Abstract
Localizing AI-edited regions is essential for interpretable forensic analysis, but remains challenging due to subtle and spatially distributed artifacts that are misaligned with semantic or object boundaries. Existing approaches rely on pixel-level supervision from controlled editing pipelines, which is difficult to scale and can introduce misleading signals: artifacts frequently extend beyond annotated regions, while out-of-mask pixels are treated as authentic. This limits models' ability to capture transferable evidence and generalize across generators and datasets. To address these issues, we propose ReGFLoW, a Reconstruction-Guided Fake Localization framework under Weak supervision, which is the first weakly supervised approach for diffusion-edited fake region localization. ReGFLoW requires only real/fake labels at the image level and uses diffusion reconstruction errors as dense spatial guidance to inject them into both feature and score spaces. Furthermore, by artifact-centric multiple instance learning, ReGFLoW utilizes localized diffusion evidence without relying on semantic-affinity or boundary-based pseudo-mask priors. Extensive experiments show competitive cross-generator localization, while ReGFLoW outperforms all evaluated fully supervised baselines when evaluation includes both partially edited and fully synthetic images and in cross-dataset tests, without target-domain adaptation.
Figures & tables
| Supervision | Method | Metric | In-domain | Cross-domain | Average | ||||
|---|---|---|---|---|---|---|---|---|---|
| SD1.5 | SD2.1 | SDXL | SD3 | Flux.1 | Avg. | OOD Avg. | |||
| Full | MaskCLIP [ 9 ] | F1 | 75.8 | 63.2 | 35.2 | 49.8 | 18.4 | 48.5 | 41.7 |
| IoU | 68.6 | 56.1 | 29.5 | 42.7 | 14.8 | 42.4 | 35.8 | ||
| TruFor [ 19 ] | F1 | 71.0 | 61.9 | 31.9 | 38.5 | 9.7 | 42.6 | 35.5 | |
| IoU | 63.4 | 54.7 | 26.5 | 32.2 | 7.6 | 36.9 | 30.3 | ||
| IML-ViT [ 18 ] | F1 | 73.6 | 50.6 | 26.0 | 28.3 | 7.9 | 37.3 | 28.2 | |
| Supervision | Method | In-domain | Cross-domain | Average | ||||
|---|---|---|---|---|---|---|---|---|
| SD1.5 | SD2.1 | SDXL | SD3 | Flux.1 | Avg. | OOD Avg. | ||
| Full | TruFor [ 19 ] | 28.8 | 20.1 | 15.9 | 16.4 | 8.4 | 17.9 | 15.2 |
| IML-ViT [ 18 ] | 30.2 | 24.7 | 17.6 | 13.5 | 7.7 | 18.7 | 15.9 | |
| MaskCLIP [ 9 ] | 95.1 | 73.2 | 20.8 | 17.7 | 4.8 | 42.3 | 29.1 | |
| Weak | Ours | 35.6 | 33.4 | 30.1 | 32.2 | 29.4 | 32.2 | 31.3 |
| Supervision | Method | COCO-GLIDE | DOLOS | |
|---|---|---|---|---|
| P | P | P+F | ||
| Full | TruFor [ 19 ] | 25.17 | 8.21 | 12.70 |
| IML-ViT [ 18 ] | 15.23 | 1.16 | 1.74 | |
| MaskCLIP [ 9 ] | 8.34 | 13.35 | 14.65 | |
| Weak | Ours | 42.01 | 21.09 | 21.19 |
| Approach | Method | SD1.5 | SD2.1 | SDXL | SD3 | Flux.1 | Avg | OOD Avg | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| F1 | ACC | F1 | ACC | F1 | ACC | F1 | ACC | F1 | ACC | F1 | ACC | F1 | ACC | ||
| Image only | CNNDet [ 54 ] | 84.60 | 85.04 | 71.56 | 75.94 | 59.70 | 68.72 | 56.27 | 67.08 | 35.72 | 57.57 | 61.57 | 70.87 | 55.81 | 67.33 |
| GramNet [ 55 ] | 80.51 | 80.35 | 74.01 | 76.66 | 65.28 | 70.76 | 64.35 | 70.29 | 52.00 | 63.37 | 67.23 | 72.29 | 63.91 | 70.27 | |
| FreqNet [ 56 ] | 75.88 | 77.70 | 60.97 | 68.37 | 53.15 | 64.02 | 53.50 | 64.37 | 38.47 | 57.08 | 56.39 | 66.31 | 51.52 | 63.46 | |
| NPR [ 57 ] | 79.41 | 79.28 | 81.67 | 81.84 | 72.12 | 74.28 | 73.43 | 75.47 | 67.62 | 71.36 | 74.85 | 76.45 | 73.71 | 75.74 | |
| Image &Pixel | TruFor [ 19 ] | 90.12 | 97.73 | 35.93 | 55.62 | 58.04 | 66.41 | 59.73 | 67.51 | 49.12 | 61.62 | 58.59 | 69.78 | 50.70 | 62.79 |
| Method | Bias Predictor | Feature-aligned Fusion | Artifact-Centric MIL | Pixel F1 |
|---|---|---|---|---|
| ReGFLoW | ✓ | ✓ | ✓ | 49.38 |
| w/o Bias Predictor | ✓ | ✓ | 46.06 | |
| w/o Reconstruction Prior | ✓ | 38.08 | ||
| Score-Pooling BCE | 32.34 |
| Train | SD2 | SD3 | SDXL | FLUX | Avg. |
|---|---|---|---|---|---|
| SD1.5 | 40.96 | 38.09 | 29.36 | 24.12 | – |
| + SD2 | – | 42.34 (+4.25) | 33.73 (+4.37) | 28.71 (+4.60) | +4.40 |
| + SD3 | 42.75 (+1.79) | – | 34.48 (+5.12) | 30.96 (+6.84) | +4.58 |
| + SDXL | 40.88 (-0.08) | 41.84 (+3.75) | – | 31.00 (+6.88) | +3.52 |
| + FLUX | 41.57 (+0.61) | 43.82 (+5.73) | 34.47 (+5.11) | – | +3.82 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Workflow category | Examples | Implication for localization supervision |
|---|---|---|
| Mask-conditioned inpainting pipelines | Stable Diffusion/SDXL inpainting, Kandinsky inpainting, FLUX.1 Fill [ 58 , 59 ] | The intended edit mask is available and useful for controlled benchmark construction, but it does not necessarily cover the full artifact support. |
| Consumer-facing tools with UI-level selection | ChatGPT Images, Adobe Firefly, Midjourney Editor, Canva Magic Edit, CapCut [ 32 , 33 , 34 , 35 , 36 ] | The selection or brush region is a user-facing edit control. It is not usually available for images collected after editing or sharing, and it is not a model-internal artifact mask. |
| Instruction- or reference-based editing models | Gemini/Nano Banana, FLUX.2, Seedream, Qwen-Image-Edit, Hunyuan Image, OmniGen [ 38 , 39 , 40 , 41 , 42 , 43 ] | The workflow may expose only the input image, instruction, reference images, and final output. A dense pixel-level edit mask is not naturally produced as a supervision signal. |
| Asset | Usage in this work | License / terms |
|---|---|---|
| OpenSDID / OpenSDI [ 9 ] | Training and evaluation dataset for diffusion-generated and diffusion-edited image detection/localization. | CC BY-SA 4.0; academic use. |
| IMDLBenCo [ 61 ] | Training, evaluation, logging, and metric computation framework. | CC BY 4.0. |
| CLIP [ 47 ] | Frozen global visual encoder. | MIT License. |
| MAE [ 48 ] | Initialization for the local artifact encoder. | CC BY-NC 4.0. |
| Stable Diffusion VAE / latent diffusion model [ 2 ] | Frozen reconstruction model used to compute the reconstruction residual prior. | CreativeML OpenRAIL-M. |
| Hugging Face Diffusers | Software library used for loading or running diffusion/VAE components. | Apache License 2.0. |