Localizing AI-edited regions is essential for interpretable forensic analysis, but remains challenging due to subtle and spatially distributed artifacts that are misaligned with semantic or object boundaries. Existing approaches rely on pixel-level supervision from controlled editing pipelines, which is difficult to scale and can introduce misleading signals: artifacts frequently extend beyond annotated regions, while out-of-mask pixels are treated as authentic. This limits models' ability to capture transferable evidence and generalize across generators and datasets. To address these issues, we propose ReGFLoW, a Reconstruction-Guided Fake Localization framework under Weak supervision, which is the first weakly supervised approach for diffusion-edited fake region localization. ReGFLoW requires only real/fake labels at the image level and uses diffusion reconstruction errors as dense spatial guidance to inject them into both feature and score spaces. Furthermore, by artifact-centric multiple instance learning, ReGFLoW utilizes localized diffusion evidence without relying on semantic-affinity or boundary-based pseudo-mask priors. Extensive experiments show competitive cross-generator localization, while ReGFLoW outperforms all evaluated fully supervised baselines when evaluation includes both partially edited and fully synthetic images and in cross-dataset tests, without target-domain adaptation.
Figures & tables
Figure 1: Motivation of ReGFLoW. (A) Conventional localization tasks, such as semantic segmentation and image manipulation localization, often rely on object- or boundary-aligned targets. (B) Diffusion-edited fake regions can be non-semantic and weakly bounded, making similarity-based expansion unreliable. (C) ReGFLoW introduces reconstruction-guided MIL for weakly supervised fake region localization.
Figure 2: Conceptual illustration of mask incompleteness in diffusion-based editing. The guidance mask indicates the intended edit region, but latent processing and blending can introduce weak global or boundary-localized traces beyond the mask. Thus, the manipulation mask should not be treated as a complete label for all diffusion artifacts.
Figure 3: Failure case on an out-of-domain fully synthetic image. Although the entire image is fake, the mask-supervised baseline predicts only an object-shaped region and misses large fake areas.
Figure 4: Overview of the proposed ReGFLoW framework. ReGFLoW integrates dual-encoder features and a diffusion reconstruction prior for patch-level fake evidence estimation under weak supervision. The resulting score map is optimized with artifact-centric MIL and used for fake-region localization and image-level classification.
Supervision
Method
Metric
In-domain
Cross-domain
Average
SD1.5
SD2.1
SDXL
SD3
Flux.1
Avg.
OOD Avg.
Full
MaskCLIP [ 9 ]
F1
75.8
63.2
35.2
49.8
18.4
48.5
41.7
IoU
68.6
56.1
29.5
42.7
14.8
42.4
35.8
TruFor [ 19 ]
F1
71.0
61.9
31.9
38.5
9.7
42.6
35.5
IoU
63.4
54.7
26.5
32.2
7.6
36.9
30.3
IML-ViT [ 18 ]
F1
73.6
50.6
26.0
28.3
7.9
37.3
28.2
Table 1: P setting (partial edited images only). Pixel-level performance on OpenSDID benchmark. OOD Avg. denotes the average over four cross-domain generators: SD2.1, SDXL, SD3, and Flux.1.
Supervision
Method
In-domain
Cross-domain
Average
SD1.5
SD2.1
SDXL
SD3
Flux.1
Avg.
OOD Avg.
Full
TruFor [ 19 ]
28.8
20.1
15.9
16.4
8.4
17.9
15.2
IML-ViT [ 18 ]
30.2
24.7
17.6
13.5
7.7
18.7
15.9
MaskCLIP [ 9 ]
95.1
73.2
20.8
17.7
4.8
42.3
29.1
Weak
Ours
35.6
33.4
30.1
32.2
29.4
32.2
31.3
Table 2: P+F setting (partial edited and fully synthetic images). Pixel-level F1 evaluated on OpenSDID. OOD Avg. denotes the average over four cross-domain generators.
Supervision
Method
COCO-GLIDE
DOLOS
P
P
P+F
Full
TruFor [ 19 ]
25.17
8.21
12.70
IML-ViT [ 18 ]
15.23
1.16
1.74
MaskCLIP [ 9 ]
8.34
13.35
14.65
Weak
Ours
42.01
21.09
21.19
Table 3: Cross-dataset pixel-level F1 on COCO-GLIDE [ 53 ] and DOLOS [ 22 ] . P denotes partially edited images, while P+F includes both partially edited and fully synthetic images. All methods are evaluated without target-domain adaptation.
Approach
Method
SD1.5
SD2.1
SDXL
SD3
Flux.1
Avg
OOD Avg
F1
ACC
F1
ACC
F1
ACC
F1
ACC
F1
ACC
F1
ACC
F1
ACC
Image only
CNNDet [ 54 ]
84.60
85.04
71.56
75.94
59.70
68.72
56.27
67.08
35.72
57.57
61.57
70.87
55.81
67.33
GramNet [ 55 ]
80.51
80.35
74.01
76.66
65.28
70.76
64.35
70.29
52.00
63.37
67.23
72.29
63.91
70.27
FreqNet [ 56 ]
75.88
77.70
60.97
68.37
53.15
64.02
53.50
64.37
38.47
57.08
56.39
66.31
51.52
63.46
NPR [ 57 ]
79.41
79.28
81.67
81.84
72.12
74.28
73.43
75.47
67.62
71.36
74.85
76.45
73.71
75.74
Image &Pixel
TruFor [ 19 ]
90.12
97.73
35.93
55.62
58.04
66.41
59.73
67.51
49.12
61.62
58.59
69.78
50.70
62.79
Table 4: Image-level detection performance on OpenSDID. OOD Avg denotes the average over four cross-domain generators: SD2.1, SDXL, SD3, and Flux.1.
Method
Bias Predictor
Feature-aligned Fusion
Artifact-Centric MIL
Pixel F1
ReGFLoW
✓
✓
✓
49.38
w/o Bias Predictor ×
✓
✓
46.06
w/o Reconstruction Prior
×
×
✓
38.08
Score-Pooling BCE
×
×
×
32.34
Table 5: Ablation studies of ReGFLoW modules. The module names follow Figure 4 .
Train
SD2
SD3
SDXL
FLUX
Avg. Δ
SD1.5
40.96
38.09
29.36
24.12
–
+ SD2
–
42.34 (+4.25)
33.73 (+4.37)
28.71 (+4.60)
+4.40
+ SD3
42.75 (+1.79)
–
34.48 (+5.12)
30.96 (+6.84)
+4.58
+ SDXL
40.88 (-0.08)
41.84 (+3.75)
–
31.00 (+6.88)
+3.52
+ FLUX
41.57 (+0.61)
43.82 (+5.73)
34.47 (+5.11)
–
+3.82
Table 6: Effect of source-domain diversity on ReGFLoW’s cross-domain localization. Each row adds one source domain to SD1.5 while keeping the total training size fixed.
Figure 5: Robustness to Gaussian blur and JPEG compression on SD3 and SDXL. Baseline curves are digitized from the robustness results reported in OpenSDI [ 9 ] , and ReGFLoW is evaluated under the same degradation settings.
Figure 6: Qualitative comparison on out-of-domain fake-region localization. MaskCLIP [ 9 ] and TruFor [ 19 ] are fully supervised, whereas Ours is weakly supervised. The examples include (a) boundary-aligned edits and (b) boundary-unaligned edits, where boundaries provide limited guidance.
The selection or brush region is a user-facing edit control. It is not usually available for images collected after editing or sharing, and it is not a model-internal artifact mask.
The workflow may expose only the input image, instruction, reference images, and final output. A dense pixel-level edit mask is not naturally produced as a supervision signal.
Appendix
Table 7: Mask availability across representative image editing workflows. The key distinction is between an edit control and a reliable pixel-level supervision mask.
Figure 7: Before–after differencing does not provide reliable pixel-level supervision. For each editing system, we show the edited output and the grayscale difference map computed against the original image. The difference map captures all rendering discrepancies, including geometric misalignment, illumination shifts, color changes, texture harmonization, and global scene adjustments. It is therefore a noisy discrepancy map rather than a reliable fake-region mask.
Asset
Usage in this work
License / terms
OpenSDID / OpenSDI [ 9 ]
Training and evaluation dataset for diffusion-generated and diffusion-edited image detection/localization.
CC BY-SA 4.0; academic use.
IMDLBenCo [ 61 ]
Training, evaluation, logging, and metric computation framework.
CC BY 4.0.
CLIP [ 47 ]
Frozen global visual encoder.
MIT License.
MAE [ 48 ]
Initialization for the local artifact encoder.
CC BY-NC 4.0.
Stable Diffusion VAE / latent diffusion model [ 2 ]
Frozen reconstruction model used to compute the reconstruction residual prior.
CreativeML OpenRAIL-M.
Hugging Face Diffusers
Software library used for loading or running diffusion/VAE components.
Apache License 2.0.
Appendix
Table 8: Existing assets used in this work.
Figure 8: Additional qualitative comparisons across diverse out-of-domain cases. Each triplet shows the input image, ground-truth mask, and ReGFLoW’s prediction (left to right). We include examples with mask-confined edits, boundary or out-of-mask traces, non-semantic background changes, real-image false-positive controls, and challenging failure cases such as tiny edits, compression artifacts, or high-frequency real textures.