Look Closer: Patch-wise Supervision for AI-Generated Image Detection
Authors: Zhida Zhang, Tao Wu, Siyu Liu, Jie Cao
Organizations: MAIS & NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA), Beijing, China · ShanghaiTech University, Shanghai, China · Anhui University, Hefei, China
How much of an image does a detector need to see? Small RGB regions can retain useful evidence of image synthesis even when they reveal little of the full scene. Motivated by single-patch detection, we study patch-wise supervision: a shared backbone classifies explicit crops, each crop receives its own loss, and patch probabilities are averaged only at inference. The procedure requires neither handcrafted residual filtering nor a learned image-level fusion module. Experiments span single-patch selection, multiple generator collections, and four CNN and Transformer backbones. On GenImage, the reported patch-wise variants improve average accuracy over their whole-image counterparts across all four backbones. Comparisons of supervision granularity, source resolution, crop size, and inference coverage further characterize the approach, while post-processing tests and difficult-image evaluation reveal its limitations. The historical experiments include evaluation-based model selection, so their scores are not presented as a uniformly selected leaderboard comparison. Overall, the study identifies explicit local input and patch-level supervision as a simple, useful combination for investigating generalizable AI-generated image detection.
Figures & tables
Figure 1: GenImage results across four backbones. Blue dashed curves denote whole-image baselines; red curves denote PWS. The reported PWS mean is higher for each backbone, although this ordering does not hold for every generator subset. Exact values and the model-selection qualifications are given in Table 2 and Section 5.1 .
Figure 2: Motivation for examining local pixel relationships. Upsampling can map a low-resolution representation into structured neighborhoods. The illustrated mapping and pixel values are schematic, not measurements establishing a universal artifact or a model of all generators.
Figure 3: The single-patch study, illustrated with minimum-complexity selection. We also test middle, maximum, and random selection. The original label Image Entropy denotes a filter-based complexity score, not Shannon entropy. Each instance uses one RGB crop without aggregation.
Figure 4: Patch-wise supervision (PWS). All patches use the same backbone and head. The single head symbol denotes parameter sharing, not feature fusion: each patch produces its own logits and loss. Only probabilities are averaged during inference. The spatial score display visualizes predictions; it is not an input to the classifier.
Input / selection
Accuracy (%)
Whole image
99.01
Single patch, minimum complexity
98.66
Single patch, middle complexity
97.93
Single patch, maximum complexity
94.91
Single patch, random
98.78
Table 1: Single-patch selection on the full-data DIFF setting. One crop occupies 6.25% of the resized image canvas. This is an input-area fraction, not a measured compute ratio.
Figure 5: Recorded training and validation accuracy over 200 epochs, shown with the original 10-epoch moving-average visualization. Ori denotes whole-image learning, SPD the single-patch study, and PWS patch-wise supervision. These trajectories illustrate observed training behavior, not a causal test of the learned features.
Backbone / detector
Input / method
Midj.
SD1.4
SD1.5
ADM
GLIDE
Wukong
VQDM
BigGAN
Mean
DIRE [ 14 ]
60.20
99.90
99.80
50.90
55.00
99.20
50.10
50.20
70.66
GenDet [ 19 ]
89.60
96.10
96.10
58.00
78.40
92.80
66.50
75.00
81.56
PatchCraft [ 18 ]
79.00
89.50
89.30
77.30
78.40
89.30
83.70
72.40
82.30
AIDE [ 15 ]
79.38
99.74
99.76
78.54
97.82
98.65
80.26
66.89
86.88
ResNet-18
Whole image
62.33
99.75
99.31
56.00
60.75
96.17
55.67
51.08
72.63
PWS
89.06
99.89
99.89
89.58
99.62
99.88
85.00
88.85
93.97
Table 2: GenImage accuracy (%) across eight generator subsets, with SD v1.4 as the training source. Whole-image and PWS rows use the indicated backbone; external-method values are retained from the original comparison. Means include SD v1.4. T marks documented target-based selection; other whole-image/PWS rows have incompletely recovered selection histories. Selection-policy differences are described in Section 5.1 . † marks an inconsistency between the reported mean and the displayed subset scores (Appendix A.5 ).
Detector
Accuracy
AP
CNNSpot [ 13 ]
70.78
80.41
LNP [ 8 ]
83.84
91.58
LGrad [ 12 ]
75.34
81.91
DIRE-G [ 14 ]
68.68
78.78
DIRE-D [ 14 ]
71.53
–
UnivFD [ 10 ]
–
91.74 †
Table 3: AIGCD mean accuracy and AP (%) across sixteen subsets. PWS models are trained on ProGAN; their complete selection histories have not been recovered (Section 5.1 ). External rows preserve the original comparison, including DIRE’s separate G/D variants. A dash denotes an unavailable metric, not zero. Per-generator values are in Appendix B ; † is explained in Appendix A.5 .
Backbone
Image-level loss
Patch-level loss
ResNet-50
72.40
95.40 T
Xception
80.19
95.21
Swin-T
77.70
93.58
Table 4: Supervision granularity on GenImage, mean accuracy (%). Image-level supervision averages patch logits before the loss; PWS applies a loss to each patch. Per-generator values appear in Appendix C . T marks the documented target-selected result; other entries have incompletely recovered selection histories.
Backbone
Whole image
Aligned PWS
Default PWS
ResNet-50
75.06
85.21
95.40 T
Xception
79.42
85.43
95.21
Table 5: Resolution-aligned GenImage comparison, mean accuracy (%). Aligned patches are extracted from a 224-pixel canvas for ResNet-50 and a 299-pixel canvas for Xception. T marks the documented target-selected result; other entries have incompletely recovered selection histories.
Patch size
GenImage
AIGCD
32×32
94.30
91.77
64×64
95.21
93.02
128×128
91.45
91.54
Table 6: Xception patch-size comparison, accuracy (%). The 128-pixel entries follow the detailed per-generator tables; the earlier main-text table interchanged their dataset columns.
Test patches
WFR
Mean
Recorded time (s)
1
92.95
87.33
6.04
4
92.95
91.75
6.19
16
96.05
92.45
6.69
64
96.45
93.02
9.06
128
96.40
92.67
16.73
256
96.65
92.70
30.08
Table 7: Inference patch-count study on AIGCD. WFR and mean accuracies are percentages. The original elapsed-time observations are retained, but their workload and timing boundary are unspecified; they should not be converted to per-image latency or compared with other implementations.
Detector
JPEG
Downsample
Blur
CNNSpot [ 13 ]
64.03
58.85
68.39
LNP [ 8 ]
53.74
63.55
67.20
LGrad [ 12 ]
51.54
60.86
71.29
DIRE-G [ 14 ]
66.57
56.09
68.92
DIRE-D [ 14 ]
70.27
62.26
70.69
UnivFD [ 10 ]
74.25
70.87
72.98
Table 8: AIGCD mean accuracy (%) after image processing. Downsampling uses ratio 0.5 and blur uses σ=1 in the experiment description. JPEG results are retained from the original report. The corresponding perturbation-evaluation script has not been located in the retained code, so the quality factor cannot be verified and this column cannot currently be reproduced from the available materials. AIDE’s downsampling value was not reported.
Training source
Detector
Backbone
Syn.
Real
Overall
BAcc
ProGAN
AIDE [ 15 ]
ResNet-50
0.63
98.46
56.45
49.55
ProGAN
PWS
ResNet-50
28.79
98.69
68.70
63.74
ProGAN
PWS
Xception
28.15
97.92
67.98
63.04
SD v1.4
AIDE [ 15 ]
ResNet-50
16.82
94.38
61.10
55.60
SD v1.4
PWS
ResNet-50
27.00
93.59
65.02
60.30
SD v1.4
PWS T
Xception
37.58
96.63
71.29
67.11
Table 9: Chameleon accuracy (%). Syn. and Real report class-wise accuracy; BAcc is their arithmetic mean, calculated here from the displayed values. AIDE rows are retained from the original comparison. T marks documented target-based selection; other PWS rows have incompletely recovered selection histories. The SD v1.4/Xception PWS row uses a checkpoint selected on Chameleon (Section 5.1 ).
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Single-patch transfer dynamics with ResNet-50 on DiffusionForensics. Solid lines denote SPD and dashed lines the whole-image baseline (ORI). Blue is ADM validation in the LSUN-Bedroom domain; orange and green are ADM and SD v1 test results in the ImageNet domain. The historical plot uses a 20-epoch moving average. It illustrates a gap between source validation and target-domain performance, not a confidence interval or a multi-seed estimate.
Figure 7: Exploratory GCM entropy ratios across 12 synthetic-image subsets. The dashed reference at one means equality of the two sample means, not equality of their full distributions. The original plot’s Natural Baseline label denotes this reference. Ratios both above and below one argue against a universal low-entropy detection rule.
Figure 8: GCM visualizations under linear-range quantization (left) and dynamic stratification (right). Each group shows the real-image mean matrix, synthetic-image mean matrix, and absolute difference. Color ranges vary between panels. The structures reflect both image statistics and quantization; their appearance is not evidence that the neural detector uses the same features.
Figure 9: Exploratory image-level GCM entropy distributions across the examined subsets: box plots in the upper two rows and histograms in the lower two rows. Real and synthetic distributions overlap, and their relative positions vary by subset. These plots complement the ratios in Figure 7 ; they are not a detection benchmark or a statistical test of the learned representation.
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China · Australian National University, Canberra, Australia · Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China +1