The rapid evolution of image generators calls for forensic cues that generalize beyond known generation mechanisms. Existing detectors often rely on static image representations or endpoint reconstruction discrepancies, leaving the evolution of intermediate reconstruction stages underexplored. We observe that the relative token predictability of real and generated images can reverse across reconstruction scales, suggesting that intermediate stages may expose forensic evidence overlooked by endpoint comparisons. Motivated by this observation, we propose RED (Reconstruction Evolution Dynamics), a framework that captures transferable forensic cues from coarse-to-fine reconstruction evolution. To our knowledge, RED is the first framework to use scale-wise token predictability to guide forensic evidence aggregation across intermediate reconstruction states. It represents the reconstruction trajectory produced by a frozen multiscale VQ-VAE in the shared feature space of a frozen CLIP encoder. To connect the observed predictability variations with visual evidence, RED learns image-adaptive stage weights from scale-wise token negative log-likelihoods provided by a frozen VAR model. A cross-stage evidence aggregation module then jointly models the original-image representation and the weighted reconstruction features, capturing complementary forensic cues through interactions along the reconstruction trajectory. Experiments on six diverse benchmarks demonstrate that RED achieves the highest average accuracy of 92.5% and average precision of 97.5% among the evaluated methods. Further evaluations show strong robustness to common image degradations, supporting the value of reconstruction evolution for generalizable AI-generated image detection. The code will be made publicly available upon acceptance of this paper.
Figures & tables
Figure 1
Figure 3: Overview of RED for generalizable AI-generated image detection. (a) Coarse-to-fine reconstruction states form a trajectory of visual evidence. (b) Scale-wise token NLLs guide image-specific weighting of reconstruction features. (c) Cross-stage evidence aggregation jointly models original-image and reconstruction representations, then combines a dedicated original-image readout with a trajectory summary pooled using the same stage weights.
Method
Venue
In Domain
Out of Domain
Avg.
GenImage
ARForensics
EvalGEN
QIB
Chameleon
G2-T
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
CNNSpot Wang et al. (2020)
CVPR
61.2
66.9
61.9
66.3
75.8
84.8
54.8
58.1
56.8
61.3
55.4
57.5
61.0
65.8
UnivFD Ojha et al. (2023)
CVPR
87.9
94.9
82.7
90.3
93.8
98.9
89.1
95.4
92.7
98.2
85.2
93.1
88.6
95.1
DIRE Wang et al. (2023)
ICCV
99.9
99.9
74.7
82.3
50.7
77.9
78.1
86.5
50.0
42.7
50.4
57.6
67.3
74.5
NPR Tan et al. (2024)
CVPR
90.9
97.6
81.1
91.8
50.5
42.5
62.2
70.6
51.6
43.3
50.9
47.9
64.5
65.6
Table 1: Comparison with state-of-the-art detectors on in-domain and out-of-domain benchmarks (%). The best, second-best, and third-best results in each column are indicated by bold , solid underline , and dashed underline , respectively. † denotes MLLM-based detectors.
Method
ARForensics (ACC.)
EvalGEN (ACC.)
Inf
JP
LG
OMV2
RAR
Switti
VAR
Avg.
Flux
GoT
Inf
NOVA
OG
Avg.
CNNSpot
75.9
74.3
59.0
49.6
54.0
72.2
48.4
61.9
72.4
79.9
67.4
79.1
80.2
75.8
UnivFD
94.7
94.1
73.5
67.1
91.9
92.0
65.7
82.7
92.6
94.0
94.1
94.5
93.7
93.8
DIRE
73.5
74.5
76.0
73.3
76.7
75.0
74.3
74.7
50.6
51.0
50.5
50.7
50.6
50.7
NPR
89.6
96.6
82.8
62.3
57.1
92.8
86.3
81.1
48.5
49.6
51.7
50.5
52.3
50.5
DRCT
98.0
97.9
88.9
74.4
66.8
97.6
73.6
85.3
87.1
97.3
86.1
87.7
96.1
90.9
Table 2: Comparison of detection accuracy (ACC) on ARForensics and EvalGEN (%).
Method
Qwen-Image-Bench (A.P.)
F2-P
F2-M
GLM
G1
G1.5
HY3
I4
I4-U
QI
Q2-P
Q2512
S4
S4.5
S5
G2
K2.1
NB2
NB-P
Avg.
CNNSpot
57.7
51.7
67.0
62.1
58.4
55.4
63.0
61.6
68.9
53.1
60.6
56.4
56.6
52.9
55.9
62.4
51.4
51.5
58.1
UnivFD
96.0
94.4
97.0
95.8
93.9
97.8
96.1
96.1
96.5
96.4
96.0
95.4
92.6
95.6
95.2
96.1
94.2
92.0
95.4
DIRE
90.7
64.6
81.0
99.9
99.9
92.6
95.3
95.0
82.1
99.9
66.0
75.9
77.0
85.7
99.9
79.2
72.2
99.9
86.5
NPR
53.8
59.3
45.2
95.1
83.3
80.8
82.3
83.0
69.8
93.7
43.8
59.2
70.4
42.5
86.1
59.6
74.1
89.3
70.6
DRCT
77.6
70.0
98.5
95.8
93.1
98.5
97.4
98.8
95.4
94.4
94.2
87.6
80.3
80.0
86.9
94.5
83.3
92.3
89.9
Table 3: Comparison of average precision (AP) on Qwen-Image-Bench (%).
Figure 4: Robustness to common image degradations on Qwen-Image-Bench. Dashed lines mark the 90% and 50% mAP reference levels.
Variant
Modification / Setting
In Domain
Out of Domain
Avg.
GenImage
ARForensics
EvalGEN
QIB
Chameleon
G2-T
Reconstruction Trajectory Representation
Original image only
Remove all reconstruction inputs
87.9
82.7
93.8
89.1
92.7
85.2
88.6
Final reconstruction only
Use the final reconstruction alone
88.0
79.1
93.2
79.5
85.8
72.7
83.0
Original + final reconstruction
Exclude intermediate reconstruction states
95.5
91.8
96.4
90.3
84.1
81.3
89.9
All stages, uniform weights
Use all reconstruction stages with equal weights
95.4
92.0
96.3
92.0
87.8
82.2
91.0
Table 4: Component-wise ablation of RED across six evaluation sets (ACC, %).
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Evaluation collection
Groups
Real/group
Generated/group
GenImage
8
1,000
1,000
ARForensics
7
1,000
1,000
EvalGEN
5
500
500
Qwen-Image-Bench
18
1,000
1,000
Chameleon
1
1,000
1,000
GPT-Image-2-Twitter
1
1,000
1,000
Appendix
Table 5: Per-group evaluation sample counts. Real images are drawn from a fixed pool of 1,000 ImageNet validation images; EvalGEN uses a fixed 500-image subset.
Method
GenImage
ADM
BigGAN
Glide
Midjourney
SD1.4
SD1.5
VQDM
Wukong
Avg.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
CNNSpot
53.6
56.1
62.4
63.8
63.6
69.0
59.0
64.7
63.1
71.9
65.1
73.8
58.7
62.8
63.9
73.1
61.2
66.9
UnivFD
75.6
86.3
90.9
97.3
91.2
97.2
83.4
92.4
91.1
96.9
92.1
97.2
89.9
96.3
89.2
95.7
87.9
94.9
DIRE
99.9
99.9
99.9
99.9
99.9
99.9
99.9
99.9
99.9
99.9
99.9
99.9
99.8
99.9
99.9
99.9
99.9
99.9
NPR
93.2
98.6
91.4
97.7
96.9
99.5
80.2
92.9
92.0
98.1
92.0
98.1
92.1
98.3
89.4
97.8
90.9
97.6
Appendix
Table 6: Per-generator accuracy (ACC) and average precision (AP) on GenImage. All results are reported in percentages, and Avg. denotes the mean across eight generators.
Method
ARForensics (A.P.)
EvalGEN (A.P.)
Inf.
JP
LG
OMV2
RAR
Switti
VAR
Avg.
Flux
GoT
Inf.
NOVA
OG
Avg.
CNNSpot
84.9
82.6
62.5
50.7
55.4
80.2
47.8
66.3
80.9
90.3
74.8
87.2
90.5
84.8
UnivFD
98.9
99.0
84.5
79.1
97.4
97.6
75.8
90.3
98.2
99.6
99.1
98.8
98.8
98.9
DIRE
81.2
81.7
83.1
81.3
84.0
82.9
81.8
82.3
77.9
80.4
71.1
80.7
79.5
77.9
NPR
97.7
99.5
95.1
83.1
72.0
98.5
96.4
91.8
35.5
35.2
49.0
44.7
48.2
42.5
DRCT
99.5
98.5
98.5
97.8
98.0
99.1
97.0
98.3
97.9
98.4
98.4
99.7
99.1
98.7
Appendix
Table 7: Per-generator average precision (AP, %) on ARForensics and EvalGEN. Avg. denotes the mean across generators within each benchmark.
Method
Qwen-Image-Bench (ACC.)
F2-P
F2-M
GLM
G1
G1.5
HY3
I4
I4-U
QI
Q2-P
Q2512
S4
S4.5
S5
G2
K2.1
NB2
NB-P
Avg.
CNNSpot
54.9
50.5
61.9
58.4
54.9
53.2
58.6
57.1
62.1
51.3
56.5
53.6
53.5
51.1
53.5
57.3
48.9
49.6
54.8
UnivFD
89.4
87.6
91.8
89.2
86.4
92.9
90.7
90.4
91.2
90.9
90.2
89.4
84.5
90.1
89.0
90.0
87.3
83.9
89.1
DIRE
88.1
61.9
62.4
99.9
99.8
88.2
88.6
88.2
62.5
99.9
50.8
62.2
62.6
63.2
99.9
62.3
65.2
99.9
78.1
NPR
52.5
54.4
51.1
84.9
65.1
65.0
70.3
71.1
58.5
81.7
50.1
52.4
55.2
49.3
71.2
55.5
58.5
73.9
62.2
DRCT
51.6
51.0
84.7
68.5
56.2
90.9
81.3
83.1
62.0
60.4
57.6
53.4
51.9
52.4
55.1
62.3
52.6
56.1
62.8
Appendix
Table 8: Per-generator accuracy (ACC, %) on Qwen-Image-Bench. Avg. denotes the mean across 18 generators.
Method
ImageNet-val
Chameleon
LSUN
MS COCO
Avg.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
ACC.
A.P.
UnivFD
92.7
98.2
60.2
75.3
88.2
92.2
82.7
88.6
80.95
88.58
D 3 QE
86.9
92.2
62.9
67.6
88.3
94.1
87.3
94.2
81.35
87.03
MF 2 DA
93.3
99.2
69.0
85.4
90.0
93.7
89.9
93.9
85.55
93.05
Ours
90.7
97.1
88.5
95.5
89.6
96.3
90.1
96.9
89.73
96.45
Appendix
Table 9: Comparison across real-image sources with a fixed generated-image source. ACC and AP are reported in percentages.
Figure 5: Scale-wise NLL difference profiles relative to ImageNet validation.
Figure 6: Scale-wise token-NLL profiles under the frozen VAR model.
Detecting AI-generated images is only half the task: a deployed detector must also justify its verdict, yet existing detectors inherit three failure modes from their training data: real and fake images collected from different sources invite provenance shortcuts, supervised explanation corpora teach templated rationales, and a static forgery corpus leaves the decision boundary standing still while generators keep moving. We introduce \methodname{}, an adversarial reinforcement learning framework that pits two heterogeneous models against each other. A diffusion image editor learns to edit real photographs into fake counterparts of those same photographs that fool the current detector, while a reasoning MLLM learns to expose them with a verdict grounded in free-form reasoning. Both rewards are shortcut-proof by design: the attacker is credited only when its edit is faithfully executed, and the defender only when its verdict is correct. As the two models alternate, each round's attacker regenerates a harder training pool aimed at the current detector's blind spots, so the detector must generalize rather than memorize any fixed artifact distribution. Although the explanation is never rewarded, its quality rises round over round as a side effect of accuracy-only training. A detector trained within this loop improves monotonically across rounds on each of three external benchmarks.
Yicheng Bao, Xiahui Guo, Xuhong Wang +1
1East China Normal University · 2Shanghai Artificial Intelligence Laboratory
Detecting AI-generated images (AIGI) remains challenging because detectors often fail to generalize to unseen generators. Although existing methods are trained on large datasets, their performance still degrades when generation settings change, indicating that data scale alone is insufficient and that limited coverage of generative variations during training is a key factor. Studies on generative model editing show that small changes in internal representations can produce diverse and meaningful image variations, many of which are not explored under standard sampling. Leveraging this insight, we propose PROBE (Probing Robustness via Boundary Exploration), a framework that improves detector generalization by actively exploring challenging regions of the generative process. Instead of treating the generator as a fixed data source, PROBE uses the detector as a critic to steer the generator through manifold-level modifications, producing realistic samples that are difficult to classify. These samples expose failure cases that are uncommon under standard data sampling strategies and are used to refine the detector. Experimental results across multiple benchmarks indicate that PROBE enhances generalization to unseen generators, resulting in more generalizable AIGI detection performance. Code and models are available at https://github.com/Amamiya-C/PROBE-AIGI-Detection
Zijie Cao, Weijie Tu, Yao Xiao +3
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China · Australian National University, Canberra, Australia · Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China +1
AI-generated image detectors achieve high accuracy on in-distribution data but often fail on unseen generators. A key obstacle to understanding this failure is the black-box nature of current detectors: they do not reveal which evidence drives their decisions. We propose ForensicConcept, a framework that extracts explicit forensic concepts from detectors and enables their transfer across backbones. Our method localizes decision-critical patches via Transformer attribution, clusters them into a compact concept codebook, and uses a concept-aligned projection to produce auditable evidence readouts. Motivated by prior studies showing that DINO representations can guide diffusion generation and exhibit concept-level correspondence with diffusion features, we introduce a generation-trace reference based on CleanDIFT diffusion features and quantify backbone-trace alignment via neighborhood-structure consistency (CKNNA). We further propose concept codebook injection to transfer diffusion-derived concepts into target backbones. Experiments on GenImage, GAN-family, and Chameleon benchmarks show consistent improvements over prior methods. We also find that CKNNA alignment predicts transfer effectiveness, providing a principled explanation for why some backbones yield more transferable forensic evidence than others.
Menyanshu Zhou, Ziyin Zhou, Ke Sun +4
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, School of Informatics, Xiamen University, Xiamen 361005, P.R.China · Sino-Russian Research Center for Digital Economy