Look Closer: Patch-wise Supervision for AI-Generated Image Detection
Authors: Zhida Zhang, Tao Wu, Siyu Liu, Jie Cao
Organizations: MAIS & NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA), Beijing, China · ShanghaiTech University, Shanghai, China · Anhui University, Hefei, China
How much of an image does a detector need to see? Small RGB regions can retain useful evidence of image synthesis even when they reveal little of the full scene. Motivated by single-patch detection, we study patch-wise supervision: a shared backbone classifies explicit crops, each crop receives its own loss, and patch probabilities are averaged only at inference. The procedure requires neither handcrafted residual filtering nor a learned image-level fusion module. Experiments span single-patch selection, multiple generator collections, and four CNN and Transformer backbones. On GenImage, the reported patch-wise variants improve average accuracy over their whole-image counterparts across all four backbones. Comparisons of supervision granularity, source resolution, crop size, and inference coverage further characterize the approach, while post-processing tests and difficult-image evaluation reveal its limitations. The historical experiments include evaluation-based model selection, so their scores are not presented as a uniformly selected leaderboard comparison. Overall, the study identifies explicit local input and patch-level supervision as a simple, useful combination for investigating generalizable AI-generated image detection.
Figures & tables
Figure 1: GenImage results across four backbones. Blue dashed curves denote whole-image baselines; red curves denote PWS. The reported PWS mean is higher for each backbone, although this ordering does not hold for every generator subset. Exact values and the model-selection qualifications are given in Table 2 and Section 5.1 .
Figure 2: Motivation for examining local pixel relationships. Upsampling can map a low-resolution representation into structured neighborhoods. The illustrated mapping and pixel values are schematic, not measurements establishing a universal artifact or a model of all generators.
Figure 3: The single-patch study, illustrated with minimum-complexity selection. We also test middle, maximum, and random selection. The original label Image Entropy denotes a filter-based complexity score, not Shannon entropy. Each instance uses one RGB crop without aggregation.
Figure 4: Patch-wise supervision (PWS). All patches use the same backbone and head. The single head symbol denotes parameter sharing, not feature fusion: each patch produces its own logits and loss. Only probabilities are averaged during inference. The spatial score display visualizes predictions; it is not an input to the classifier.
Input / selection
Accuracy (%)
Whole image
99.01
Single patch, minimum complexity
98.66
Single patch, middle complexity
97.93
Single patch, maximum complexity
94.91
Single patch, random
98.78
Table 1: Single-patch selection on the full-data DIFF setting. One crop occupies 6.25% of the resized image canvas. This is an input-area fraction, not a measured compute ratio.
Figure 5: Recorded training and validation accuracy over 200 epochs, shown with the original 10-epoch moving-average visualization. Ori denotes whole-image learning, SPD the single-patch study, and PWS patch-wise supervision. These trajectories illustrate observed training behavior, not a causal test of the learned features.
Backbone / detector
Input / method
Midj.
SD1.4
SD1.5
ADM
GLIDE
Wukong
VQDM
BigGAN
Mean
DIRE [ 14 ]
60.20
99.90
99.80
50.90
55.00
99.20
50.10
50.20
70.66
GenDet [ 19 ]
89.60
96.10
96.10
58.00
78.40
92.80
66.50
75.00
81.56
PatchCraft [ 18 ]
79.00
89.50
89.30
77.30
78.40
89.30
83.70
72.40
82.30
AIDE [ 15 ]
79.38
99.74
99.76
78.54
97.82
98.65
80.26
66.89
86.88
ResNet-18
Whole image
62.33
99.75
99.31
56.00
60.75
96.17
55.67
51.08
72.63
PWS
89.06
99.89
99.89
89.58
99.62
99.88
85.00
88.85
93.97
Table 2: GenImage accuracy (%) across eight generator subsets, with SD v1.4 as the training source. Whole-image and PWS rows use the indicated backbone; external-method values are retained from the original comparison. Means include SD v1.4. T marks documented target-based selection; other whole-image/PWS rows have incompletely recovered selection histories. Selection-policy differences are described in Section 5.1 . † marks an inconsistency between the reported mean and the displayed subset scores (Appendix A.5 ).
Detector
Accuracy
AP
CNNSpot [ 13 ]
70.78
80.41
LNP [ 8 ]
83.84
91.58
LGrad [ 12 ]
75.34
81.91
DIRE-G [ 14 ]
68.68
78.78
DIRE-D [ 14 ]
71.53
–
UnivFD [ 10 ]
–
91.74 †
Table 3: AIGCD mean accuracy and AP (%) across sixteen subsets. PWS models are trained on ProGAN; their complete selection histories have not been recovered (Section 5.1 ). External rows preserve the original comparison, including DIRE’s separate G/D variants. A dash denotes an unavailable metric, not zero. Per-generator values are in Appendix B ; † is explained in Appendix A.5 .
Backbone
Image-level loss
Patch-level loss
ResNet-50
72.40
95.40 T
Xception
80.19
95.21
Swin-T
77.70
93.58
Table 4: Supervision granularity on GenImage, mean accuracy (%). Image-level supervision averages patch logits before the loss; PWS applies a loss to each patch. Per-generator values appear in Appendix C . T marks the documented target-selected result; other entries have incompletely recovered selection histories.
Backbone
Whole image
Aligned PWS
Default PWS
ResNet-50
75.06
85.21
95.40 T
Xception
79.42
85.43
95.21
Table 5: Resolution-aligned GenImage comparison, mean accuracy (%). Aligned patches are extracted from a 224-pixel canvas for ResNet-50 and a 299-pixel canvas for Xception. T marks the documented target-selected result; other entries have incompletely recovered selection histories.
Patch size
GenImage
AIGCD
32×32
94.30
91.77
64×64
95.21
93.02
128×128
91.45
91.54
Table 6: Xception patch-size comparison, accuracy (%). The 128-pixel entries follow the detailed per-generator tables; the earlier main-text table interchanged their dataset columns.
Test patches
WFR
Mean
Recorded time (s)
1
92.95
87.33
6.04
4
92.95
91.75
6.19
16
96.05
92.45
6.69
64
96.45
93.02
9.06
128
96.40
92.67
16.73
256
96.65
92.70
30.08
Table 7: Inference patch-count study on AIGCD. WFR and mean accuracies are percentages. The original elapsed-time observations are retained, but their workload and timing boundary are unspecified; they should not be converted to per-image latency or compared with other implementations.
Detector
JPEG
Downsample
Blur
CNNSpot [ 13 ]
64.03
58.85
68.39
LNP [ 8 ]
53.74
63.55
67.20
LGrad [ 12 ]
51.54
60.86
71.29
DIRE-G [ 14 ]
66.57
56.09
68.92
DIRE-D [ 14 ]
70.27
62.26
70.69
UnivFD [ 10 ]
74.25
70.87
72.98
Table 8: AIGCD mean accuracy (%) after image processing. Downsampling uses ratio 0.5 and blur uses σ=1 in the experiment description. JPEG results are retained from the original report. The corresponding perturbation-evaluation script has not been located in the retained code, so the quality factor cannot be verified and this column cannot currently be reproduced from the available materials. AIDE’s downsampling value was not reported.
Training source
Detector
Backbone
Syn.
Real
Overall
BAcc
ProGAN
AIDE [ 15 ]
ResNet-50
0.63
98.46
56.45
49.55
ProGAN
PWS
ResNet-50
28.79
98.69
68.70
63.74
ProGAN
PWS
Xception
28.15
97.92
67.98
63.04
SD v1.4
AIDE [ 15 ]
ResNet-50
16.82
94.38
61.10
55.60
SD v1.4
PWS
ResNet-50
27.00
93.59
65.02
60.30
SD v1.4
PWS T
Xception
37.58
96.63
71.29
67.11
Table 9: Chameleon accuracy (%). Syn. and Real report class-wise accuracy; BAcc is their arithmetic mean, calculated here from the displayed values. AIDE rows are retained from the original comparison. T marks documented target-based selection; other PWS rows have incompletely recovered selection histories. The SD v1.4/Xception PWS row uses a checkpoint selected on Chameleon (Section 5.1 ).
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Single-patch transfer dynamics with ResNet-50 on DiffusionForensics. Solid lines denote SPD and dashed lines the whole-image baseline (ORI). Blue is ADM validation in the LSUN-Bedroom domain; orange and green are ADM and SD v1 test results in the ImageNet domain. The historical plot uses a 20-epoch moving average. It illustrates a gap between source validation and target-domain performance, not a confidence interval or a multi-seed estimate.
Figure 7: Exploratory GCM entropy ratios across 12 synthetic-image subsets. The dashed reference at one means equality of the two sample means, not equality of their full distributions. The original plot’s Natural Baseline label denotes this reference. Ratios both above and below one argue against a universal low-entropy detection rule.
Figure 8: GCM visualizations under linear-range quantization (left) and dynamic stratification (right). Each group shows the real-image mean matrix, synthetic-image mean matrix, and absolute difference. Color ranges vary between panels. The structures reflect both image statistics and quantization; their appearance is not evidence that the neural detector uses the same features.
Figure 9: Exploratory image-level GCM entropy distributions across the examined subsets: box plots in the upper two rows and histograms in the lower two rows. Real and synthetic distributions overlap, and their relative positions vary by subset. These plots complement the ratios in Figure 7 ; they are not a detection benchmark or a statistical test of the learned representation.
AI-generated image detectors generalize poorly when their training and test images originate from different generators or datasets. Despite the rich spatial representations produced by vision foundation models like DINO, existing detectors typically classify images using only the globally aggregated CLS token. We hypothesize that globally aggregating DINO features into a single CLS token obscures spatially distributed generation traces. To test this hypothesis, we introduce PatchHead, a lightweight spatial aggregation head that preserves the two-dimensional organization of DINO patch tokens and integrates evidence across neighboring regions. During training, we freeze the pretrained DINO backbone and optimize only the inserted LoRA adapters, PatchHead, and auxiliary projection head. Across nine cross-dataset benchmarks spanning manually curated and in-the-wild settings, PatchHead ranks first on seven datasets and second on the remaining two. It improves the strongest prior method from 91.6% to 94.6% in average balanced accuracy (+3.0 points) and raises the worst-case accuracy from 82.4% to 89.4% (+6.9 points), while introducing only 8.6% more trainable parameters and 0.08% additional FLOPs. Further qualitative analysis suggests that PatchHead (i) reduces class-conditional domain discrepancy, and (ii) redirects the representation from content-dominated saliency toward spatially distributed authenticity evidence. Together, these observations provide a representation-level account of why spatial patch aggregation transfers more reliably across generators and datasets than a single CLS-based global representation. Our code and models will be made available upon acceptance.
Shengbo Qi, Hongyi Fang, Benjia Zhou +1
Beijing Institute of Technology, Zhuhai · Shenzhen University
Detecting AI-generated images (AIGI) remains challenging because detectors often fail to generalize to unseen generators. Although existing methods are trained on large datasets, their performance still degrades when generation settings change, indicating that data scale alone is insufficient and that limited coverage of generative variations during training is a key factor. Studies on generative model editing show that small changes in internal representations can produce diverse and meaningful image variations, many of which are not explored under standard sampling. Leveraging this insight, we propose PROBE (Probing Robustness via Boundary Exploration), a framework that improves detector generalization by actively exploring challenging regions of the generative process. Instead of treating the generator as a fixed data source, PROBE uses the detector as a critic to steer the generator through manifold-level modifications, producing realistic samples that are difficult to classify. These samples expose failure cases that are uncommon under standard data sampling strategies and are used to refine the detector. Experimental results across multiple benchmarks indicate that PROBE enhances generalization to unseen generators, resulting in more generalizable AIGI detection performance. Code and models are available at https://github.com/Amamiya-C/PROBE-AIGI-Detection
Zijie Cao, Weijie Tu, Yao Xiao +3
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China · Australian National University, Canberra, Australia · Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China +1
Generalization remains a critical bottleneck in AI-generated image detection. Because many modern generators are proprietary or adversarially modified, existing detectors overfit to the low-level textural patterns of accessible training data, resulting in severe failures on unseen domains. Conventional regularization techniques (e.g., L1/L2 norms, Dropout) apply indiscriminate parametric constraints and fail to provide the domain-invariant structure necessary for cross-generator robustness. To address this, we propose Feature-Augmented Implicit Regularization (FAIR). FAIR introduces an orthogonal, macro-structural prior, specifically, Scene Composition Structure (SCS), during training to geometrically constrain the model's optimization trajectory. By augmenting the primary feature space with domain-invariant SCS features, FAIR explicitly penalizes texture-biased shortcut learning. Crucially, this structural prior is entirely discarded at inference, yielding a smoothed, generalized decision boundary with zero architectural or computational overhead. Extensive evaluations across five massive benchmarks demonstrate that integrating FAIR into state-of-the-art detectors significantly improves cross-generator generalization, boosting accuracy by up to 8.04% and establishing new state-of-the-art robustness in zero-shot transfer scenarios.
Md Redwanul Haque, Manzur Murshed, Manoranjan Paul +1
Deakin University, Burwood, VIC, Australia · Charles Sturt University, Bathurst, NSW, Australia