Generative models can now synthesize highly realistic images, simultaneously increasing the risks of misinformation and visual forgery. Therefore, detecting AI-generated images becomes more essential, and a reliable detector must generalize to unseen generators and stay robust to unseen perturbations in the wild. Existing detectors are typically trained on either independently collected real and generated images or aligned real-generated pairs designed to mitigate content bias. Building on aligned pairs, recent methods form a mixed view by replacing some patches of the real image with their generated counterparts. However, we find that self-attention lets real and generated patches interact, so the feature of each patch no longer reflects its own source alone. This contextual shift makes a per-patch source label an imprecise target. To this end, we propose Relative Patch Response Learning (PRL). Instead of labeling each patch, PRL compares the same patch across two mixed views of an aligned pair and learns from its patch response, the change of its score between the views. (i) To give precise supervision under the contextual shift, a relative response objective measures the responses of source-changed patches against those of source-unchanged patches, which respond to the shift alone. (ii) To provide a reliable reference for the shift, a reference coherence objective keeps each group of source-unchanged patches moving as a whole. (iii) Since the two views contain different amounts of generated content, an area ranking objective asks the view with the larger generated area to have a higher mean patch score. Extensive experiments demonstrate the superior performance of PRL, which surpasses the best prior methods by 4.3% and 5.9% in average balanced accuracy across eight standard and three in-the-wild benchmarks, respectively.
Figures & tables
Figure 1: Three training paradigms and the contextual shift. (a) Detectors built on hand-crafted features or on vision foundation models (VFMs) learn from real and generated images collected separately. (b) Alignment-based methods pair each real image with a reconstruction of matched content, but treat the two as independent training samples with their own image-level labels; the pair itself is not used by the supervision. (c) Patch-level methods replace some patches of a real image by their reconstruction and label each patch by its own source, in addition to the image-level label. (d) Contextual shift: a real patch q has the same input in the real image and in a mixed view, yet its score rises once other patches are replaced, while the per-patch label of (c) still marks q as real.
Figure 2: Contextual shift in trained detectors. Each mixed input takes the outlined square from the aligned reconstruction and the rest from the real image. The maps show how much each patch score changes from the real image to the mixed input, darker for a larger change. For DINOv3 and PPL, we train a patch head on them. Although only the square changed, the DINOv3 detector and PPL respond outside it almost as much as inside it. PRL keeps a shift outside the square and responds inside it.
Figure 3: Overview of PRL. ❶ Aligned region mixing : two masks mix the real and the synthetic image into a small and a large view. ❷ Patch responses : the patch scoring head scores every position in both views, and the patch response Δp (Eq. ( 3 )) is read on the changed sets Sr,s , Ss,r and on their references Sr,r , Ss,s ; the gating MLP gives the image-level decision. ❸ Training objectives : relative response, reference coherence and area ranking.
Method
Standard Benchmarks
In-the-Wild Benchmarks
GenImage
AIGCDetect
DRCT-2M
DDA-COCO
EvalGEN
Synthbuster
ForenSynths
UnivFD
Chameleon
SynthWildX
WildRF
Avg.
UnivFD CVPR’23
64.05
72.50
59.64
52.32
15.39
55.31
77.69
78.95
50.71
52.27
55.08
57.63
FatFormer CVPR’24
61.03
74.78
50.82
51.63
27.17
53.62
86.13
82.07
51.11
53.31
59.62
59.21
NPR CVPR’24
46.74
47.78
49.62
49.07
3.43
41.83
37.88
41.92
52.26
49.80
59.34
43.61
DRCT ICML’24
80.42
67.67
90.41
66.89
50.93
73.04
54.86
64.27
68.93
74.47
74.22
69.65
C2P-CLIP AAAI’25
62.09
70.16
54.29
50.90
13.82
68.08
87.00
84.88
50.60
53.07
59.14
59.46
Table 1: Cross-benchmark comparison. Balanced accuracy (%) on eight standard and three in-the-wild benchmarks; EvalGEN has no real counterpart and is reported as accuracy on its synthetic images. Every compared method uses its official checkpoint and inference code. Bold: best per column; underlined: second best.
Method
BigGAN
CRN
CycleGAN
DeepFake
GauGAN
IMLE
ProGAN
SAN
SITD
StarGAN
StyleGAN
DALL-E
Glide-50-27
Glide-100-10
Glide-100-27
Guided
LDM-100
LDM-200
LDM-200-cfg
Avg.
UnivFD CVPR’23
87.50
55.67
96.92
69.39
98.79
68.11
99.38
58.22
62.22
95.12
79.99
79.55
77.80
77.40
77.80
67.15
91.20
90.75
67.05
78.95
FatFormer CVPR’24
97.90
69.32
99.32
86.03
99.17
69.45
99.21
62.10
80.83
99.75
87.51
88.40
70.10
69.60
70.60
61.45
88.40
89.45
70.75
82.07
NPR CVPR’24
43.25
0.00
59.83
33.30
35.62
0.00
50.45
48.40
22.22
49.47
50.18
52.30
50.50
50.85
51.00
45.75
51.15
51.10
51.05
41.92
DRCT ICML’24
52.45
49.80
48.83
56.38
49.45
49.26
53.02
86.76
54.44
54.38
54.45
79.05
60.05
66.45
63.15
59.00
94.80
94.80
94.65
64.27
C2P-CLIP AAAI’25
92.07
92.75
99.39
91.85
98.20
97.68
97.10
74.20
87.22
99.70
76.97
84.90
78.90
78.60
76.35
57.05
82.70
83.65
63.40
84.88
AIDE ICLR’25
43.68
19.96
51.73
45.54
39.73
19.38
49.31
51.37
30.56
51.43
49.57
51.80
51.25
51.55
51.50
50.25
60.80
62.25
62.35
47.05
Table 2: Comparison of balanced accuracy (%) on the UnivFD benchmark.
Table 6
Figure 4: Robustness to post-processing. Balanced accuracy on debiased GenImage-JPEG96 dataset, when every test image is JPEG-compressed at decreasing quality (a), resized by a scale factor (b) or blurred with a Gaussian of increasing σ (c).
Figure 5: Ablation. Balanced accuracy (%) on GenImage, AIGCDetect, Chameleon and WildRF. (a) Three ways of using the aligned pairs: independent image-level label without and with the mixed views, the per-patch label on the mixed views, and the patch response. (b) PRL with one objective removed.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Method
ADM
BigGAN
Glide
Midjourney
SDv1.4
SDv1.5
VQDM
Wukong
Avg.
UnivFD CVPR’23
62.51
84.37
61.28
55.06
55.57
55.67
76.88
61.10
64.05
FatFormer CVPR’24
59.82
81.77
60.28
51.99
52.81
52.76
69.97
58.86
61.03
NPR CVPR’24
46.73
44.25
46.06
43.86
46.73
47.06
50.00
49.22
46.74
DRCT ICML’24
60.74
58.38
67.86
88.68
99.37
99.21
70.03
99.10
80.42
C2P-CLIP AAAI’25
53.62
92.67
65.91
50.60
55.02
55.31
63.58
60.00
62.09
AIDE ICLR’25
50.22
50.42
52.26
57.46
75.96
76.19
50.78
69.47
60.35
Appendix
Table 5: Comparison of balanced accuracy (%) on GenImage.
Method
ADM
BigGAN
CycleGAN
DALLE2
GauGAN
Glide
Midjourney
ProGAN
SDXL
SDv1.4
SDv1.5
StarGAN
StyleGAN
StyleGAN2
VQDM
WFIR
Wukong
Avg.
UnivFD CVPR’23
62.51
87.50
96.92
49.95
98.79
61.28
55.06
99.38
58.15
55.57
55.67
95.12
79.99
69.41
76.88
69.20
61.10
72.50
FatFormer CVPR’24
59.82
97.90
99.32
49.00
99.17
60.28
51.99
99.21
63.75
52.81
52.76
99.75
87.51
81.63
69.97
87.50
58.86
74.78
NPR CVPR’24
46.73
43.25
59.83
46.65
35.62
46.06
43.86
50.45
47.48
46.73
47.06
49.47
50.18
50.02
50.00
49.65
49.22
47.78
DRCT ICML’24
60.74
52.45
48.83
68.25
49.45
67.86
88.68
53.02
80.65
99.37
99.21
54.38
54.45
53.74
70.03
50.15
99.10
67.67
C2P-CLIP AAAI’25
53.62
92.07
99.39
50.20
98.20
65.91
50.60
97.10
51.20
55.02
55.31
99.70
76.97
58.51
63.58
65.35
60.00
70.16
AIDE ICLR’25
50.22
43.68
51.73
50.95
39.73
52.26
57.46
49.31
50.55
75.96
76.19
51.43
49.57
49.20
50.78
49.20
69.47
53.98
Appendix
Table 6: Comparison of balanced accuracy (%) on AIGCDetectBenchmark.
Method
DALL-E 2
DALL-E 3
Firefly
GLIDE
Midjourney
SD1.3
SD1.4
SD2
SDXL
Avg.
UnivFD CVPR’23
70.98
34.83
77.38
40.73
39.93
57.88
57.38
63.13
55.48
55.31
FatFormer CVPR’24
53.53
35.33
59.83
67.28
43.88
52.43
53.18
51.03
66.08
53.62
NPR CVPR’24
55.76
12.36
9.46
53.01
40.81
55.76
56.41
35.86
57.11
41.83
DRCT ICML’24
47.54
53.39
48.59
60.19
89.29
94.09
94.09
90.79
79.39
73.04
C2P-CLIP AAAI’25
54.89
43.64
81.54
81.59
52.74
79.69
79.94
65.64
72.99
68.08
AIDE ICLR’25
37.83
36.63
27.73
67.93
60.48
77.03
76.68
56.18
71.38
56.87
Appendix
Table 7: Comparison of balanced accuracy (%) on Synthbuster.
Method
SDXL-Ref
SDXL-Ctrl
LCM-SDv1.5
LCM-SDXL
LDM
SD-Ctrl
SD-Turbo
SD2.1-Ctrl
SDXL-Turbo
SD2.1
SDv2-DR
SDv1-DR
SDv1.4
SDv1.5
SDXL-DR
SDXL
Avg.
UnivFD CVPR’23
54.30
72.80
53.48
65.20
81.27
66.26
55.25
63.97
52.42
56.08
52.56
56.47
55.16
54.78
52.18
62.04
59.64
FatFormer CVPR’24
48.60
61.06
48.61
49.97
52.93
49.97
48.59
50.20
48.54
48.56
53.71
54.98
48.55
48.55
51.78
48.55
50.82
NPR CVPR’24
47.72
57.20
49.16
47.72
48.85
48.11
48.28
48.57
48.07
48.50
50.78
52.60
48.84
48.77
53.03
47.76
49.62
DRCT ICML’24
85.86
80.19
99.65
78.73
99.90
99.89
91.94
95.04
70.54
96.42
92.63
99.88
99.90
99.88
72.27
83.90
90.41
C2P-CLIP AAAI’25
50.92
72.09
50.67
53.56
63.61
57.02
51.07
54.28
50.44
50.95
50.91
55.89
51.88
51.64
50.68
52.98
54.29
AIDE ICLR’25
66.05
53.42
69.95
54.13
60.74
64.80
53.03
52.95
53.07
58.09
50.02
54.52
74.82
74.28
50.01
53.11
58.94
Appendix
Table 8: Comparison of balanced accuracy (%) on DRCT-2M.
Method
DDA-COCO
EvalGEN
VAE-ema
VAE-mse
SDXL
SD2.1
SD3.5
Flux
Avg.
Flux
GoT
Infinity
NOVA
OmniGen
Avg.
UnivFD CVPR’23
54.40
53.24
51.52
53.29
51.28
50.21
52.32
3.94
10.27
15.30
39.72
7.71
15.39
FatFormer CVPR’24
51.49
53.17
50.77
53.31
51.63
49.42
51.63
1.59
16.14
19.03
91.85
7.22
27.17
NPR CVPR’24
49.39
48.88
48.30
48.85
48.52
50.46
49.07
2.04
1.10
6.01
2.42
5.57
3.43
DRCT ICML’24
83.31
76.99
62.49
76.82
51.47
50.24
66.89
36.26
49.59
63.40
49.68
55.70
50.93
C2P-CLIP AAAI’25
51.42
51.27
50.22
51.23
51.28
49.98
50.90
1.78
6.71
11.39
46.57
2.66
13.82
Appendix
Table 9: Comparison of balanced accuracy (%) on DDA-COCO and of accuracy (%) on the synthetic images of EvalGEN.
Method
BigGAN
CRN
CycleGAN
DeepFake
GauGAN
IMLE
ProGAN
SAN
SITD
StarGAN
StyleGAN
StyleGAN2
WFIR
Avg.
UnivFD CVPR’23
87.50
55.67
96.92
69.39
98.79
68.11
99.38
58.22
62.22
95.12
79.99
69.41
69.20
77.69
FatFormer CVPR’24
97.90
69.32
99.32
86.03
99.17
69.45
99.21
62.10
80.83
99.75
87.51
81.63
87.50
86.13
NPR CVPR’24
43.25
0.00
59.83
33.30
35.62
0.00
50.45
48.40
22.22
49.47
50.18
50.02
49.65
37.88
DRCT ICML’24
52.45
49.80
48.83
56.38
49.45
49.26
53.02
86.76
54.44
54.38
54.45
53.74
50.15
54.86
C2P-CLIP AAAI’25
92.07
92.75
99.39
91.85
98.20
97.68
97.10
74.20
87.22
99.70
76.97
58.51
65.35
87.00
AIDE ICLR’25
43.68
19.96
51.73
45.54
39.73
19.38
49.31
51.37
30.56
51.43
49.57
49.20
49.20
42.36
Appendix
Table 10: Comparison of balanced accuracy (%) on ForenSynths.
Method
FLUX-Fill (rand.)
FLUX-Fill (obj.)
FLUX.2 -klein
Qwen- Edit
Step1X- Edit
Nano Banana 2
GPT- Image 2
Avg.
UnivFD CVPR’23
50.16
50.83
50.72
50.39
50.72
50.26
50.23
50.47
FatFormer CVPR’24
49.86
49.95
50.11
50.05
50.26
50.00
56.96
51.02
NPR CVPR’24
45.09
46.02
49.81
55.17
48.77
55.53
80.33
54.38
DRCT ICML’24
51.00
50.86
50.56
52.17
50.13
57.21
63.99
53.70
AIDE ICLR’25
51.13
50.65
49.63
58.63
57.54
54.43
60.43
54.63
AlignedForensics ICLR’25
50.51
50.44
50.42
73.18
50.43
64.14
56.96
56.58
Appendix
Table 11: Comparison of balanced accuracy (%) on the ManipulationBench subset of DailyBench; rows other than ours are as reported by DailyBench, including its own detector FPD.
Method
GenImage
AIGCDetect
Chameleon
WildRF
Avg.
Image-level BCE on the aligned pairs
94.58
96.06
85.89
94.06
92.65
+ mixed views, image-level BCE
97.33
97.38
82.84
93.17
92.68
+ mixed views, per-patch label
97.45
96.92
79.29
93.62
91.82
PRL (Ours)
96.83
97.75
90.01
97.59
95.55
w/o Lrel
96.32
96.77
83.20
95.82
93.03
w/o Lcoh
96.05
96.43
86.83
96.89
94.05
Appendix
Table 12: Ablation. Balanced accuracy (%) of the variants of Figure 5 on GenImage, AIGCDetect, Chameleon and WildRF. All rows use the same DINOv3 backbone with LoRA and the same aligned training pairs. Top: three ways of using the aligned pairs; bottom: PRL with one objective removed.
Method
Standard Benchmarks
In-the-Wild Benchmarks
GenImage
AIGCDetect
DRCT-2M
DDA-COCO
EvalGEN
Synthbuster
ForenSynths
UnivFD
Chameleon
SynthWildX
WildRF
Avg.
DDA, released model (DINOv2)
91.74
87.83
98.47
92.89
96.86
85.52
81.41
82.58
82.41
90.87
90.20
89.16
DDA on DINOv3
97.93
96.09
99.04
95.98
99.40
88.77
87.50
85.52
91.82
89.01
94.17
93.20
PRL (Ours)
96.83
97.75
99.27
94.39
99.14
97.00
92.57
94.22
90.01
93.60
97.59
95.67
Appendix
Table 13: DDA and PRL on the same backbone. Balanced accuracy (%) on the eleven benchmarks of Table 1 for the released DDA model (DINOv2), for DDA adapted from its official code to our DINOv3 backbone with LoRA, and for PRL, which uses the same backbone, adaptation and aligned pairs.
Figure 6: Sensitivity to the loss weights. Mean balanced accuracy (%) over GenImage, AIGCDetect, Chameleon and WildRF when one weight varies and the other two keep their default values; the red bar is the default setting.
MAIS & NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA), Beijing, China · ShanghaiTech University, Shanghai, China · Anhui University, Hefei, China