AI-generated image detectors are commonly trained on fixed generator domains and become difficult to maintain as new generative models emerge. Continual adaptation is challenging because replaying historical generated images is costly, whereas updating shared parameters with limited current-domain data can overwrite prior forensic knowledge. We propose EvoKnow, a replay-free framework that formulates continual AI-generated image detection as forensic knowledge evolution. EvoKnow preserves a shared forensic basis learned from base domains, incrementally adds isolated residual experts for complementary generator-relevant evidence, and retrieves expertise through an Analytical Incremental Router (AIR) updated in closed form from current-stage generated images and accumulated sufficient statistics. Experiments demonstrate effective cross-generator generalization, few-shot expansion, and long-horizon continual adaptation. With ten generated images per arriving generator, EvoKnow achieves 96.70% average accuracy on non-base GenImage generators and 94.48% accuracy on Chameleon without target-benchmark adaptation. Under a strict replay-free continual learning protocol, EvoKnow achieves state-of-the-art continual learning performance, attaining 96.32% mean stage-wise accuracy and 4.32% average forgetting.
Figures & tables
Figure 1: Continual AIGC detection with EvoKnow. Instead of repeatedly reshaping a unified forensic space, EvoKnow freezes transferable shared evidence and incrementally adds isolated residual expertise for newly arriving generators. Historical synthetic images are not replayed.
Figure 2: Overview of EvoKnow. A global view and K local views first pass through a frozen shared forensic learner. AIR performs expert retrieval over the base expert and isolated incremental residual experts. Within each expert, LoRA adapts semantic evidence shared by both branches, GPB emphasizes structural evidence in the global branch, and LFB emphasizes frequency evidence in the local branch. Cross-attention aggregates local evidence, which is fused with the global score for real-versus-synthetic prediction. When a new generator arrives, only its new incremental residual expert is trained and AIR is updated in closed form; the shared forensic learner and previously learned experts remain frozen, and historical synthetic images are not replayed.
Dataset
Generator / Metric
Existing Detectors Base: ProGAN
EvoKnow Base: SD v1.4
CNNSpot
UnivFD
NPR
FatFormer
C2P-CLIP
AIDE ∗
SAFE ∗
MiraGe ∗
OmniDFA
Base
Exp.
Transfer
GenImage
SD v1.4 (base)
96.3
55.6
55.1
53.6
77.5
77.2
99.9
98.8
97.47 †
99.70
98.18
–
BigGAN
46.8
84.4
57.7
82.2
85.9
50.6
77.4
96.5
97.33 †
93.80
93.77
–
Midjourney
52.8
55.1
53.4
52.1
56.6
58.2
95.7
83.2
97.58 †
93.10
96.06
–
SD v1.5
95.9
55.7
55.0
53.8
76.9
77.4
99.8
98.5
97.75 †
99.70
98.53
–
ADM
50.1
62.5
43.8
61.4
71.6
50.4
59.5
82.7
85.50 †
78.10
96.21
–
Table 1: Cross-generator and cross-benchmark detection accuracy (%). Gray rows denote base domains and orange rows denote summary metrics; base domains are excluded from averages. Ave. averages non-base generators, while GAN Ave. and Diffusion Ave. summarize the corresponding UniversalFakeDetect groups. Existing-detector results follow their reported ProGAN-based protocols, whereas EvoKnow is initialized from GenImage SD v1.4; this is therefore a reported-result comparison rather than a controlled same-source evaluation. Base denotes the zero-shot shared forensic learner and base expert, Exp. denotes generator-based GenImage expansion with ten images per arriving generator, and Transfer evaluates Exp. directly on external benchmarks without target adaptation. EvoKnow columns are shaded in blue. Best and second-best summary values are shown in bold and underlined ; † denotes OmniDFA models trained with UniversalFakeDetect data.
Method
Observed-domain average accuracy ( AAt , %)
Continual summary
ProGAN
Deepfake
BigGAN
StyleGAN2
DDPM
ADM
DALL-E
GLIDE
SD v1.4
Midjourney
VQDM
SD v2.1
SDXL v1.0
SD v3.0
AA↑
AF↓
Seq
99.99
94.46
81.01
94.05
90.35
80.14
79.66
76.68
84.50
73.82
74.31
51.90
57.29
73.85
79.43
23.04
Joint
99.99
98.77
99.34
99.57
98.15
98.97
98.81
98.54
98.80
98.30
98.38
95.69
95.17
96.01
98.18
1.77
ER Chaudhry et al. (2019)
99.99
98.26
96.77
96.96
90.92
85.97
86.18
84.65
84.11
83.94
83.73
84.51
86.02
90.68
89.48
12.10
EWC Kirkpatrick et al. (2017)
99.99
92.57
91.63
93.73
89.44
82.87
79.85
70.47
82.35
80.11
77.10
67.97
69.39
78.11
82.54
21.52
OSLA Ritter et al. (2018)
99.99
98.37
90.25
94.98
87.69
76.65
85.05
75.53
83.32
77.86
85.92
69.61
73.33
79.08
84.12
23.77
Table 2: Continual detection on the 14-stage multi-source benchmark of Wang et al. Wang et al. (2026a) , covering GANs, deepfakes, and diffusion models. ProGAN is the base domain, and each subsequent generator provides ten generated images for adaptation. Baselines follow their reported continual-learning protocols; Joint explicitly uses accumulated historical data. In contrast, EvoKnow neither retains nor replays historical generated images or their instance-level features, and uses task-agnostic real-image supervision for binary detection. The first 14 columns report stage-wise observed-domain accuracy AAt using the complete AIR pipeline; AA averages AAt over the trajectory, and AF is in percentage points. Baseline AF signs are converted for consistent comparison. Best non-Joint results are bolded, with ties highlighted once.
Figure 3: Expert retrieval and adaptation-sample analysis. (a) Generator-to-expert routing assignments. (b) Effect of the GenImage adaptation budget on cross-generator evaluation over UniversalFakeDetect, reported by average accuracy and across-generator standard deviation. UniversalFakeDetect is used only for evaluation.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Phase
Hyperparameter
Value
Shared knowledge learning
Backbone
DINOv3-L/16
Input resolution
224×224
Global / local views
1 / 4
Optimizer
AdamW
Learning rate
5×10−6
Weight decay
10−2
Appendix
Table 4: Implementation details and hyperparameter settings.
Figure 4: Analytical illustration of the structural comparison in GPB. (a) A hard neighbor–center comparison and its sigmoid relaxations. (b) Derivatives with respect to the feature difference. The initial sharpness is k=10 ; the other values illustrate how sharpness changes the transition and gradient concentration. The hard comparison has zero derivative for d=0 and no ordinary derivative at d=0 , indicated by an open marker.
Figure 5: Analytical illustration of spectral selection in LFB. (a) A hard high-pass mask at c=0.25 and smooth masks with α=30 at three cutoff values. (b) Their derivatives with respect to the cutoff. The initial cutoff is c=0.25 . The smooth mask provides cutoff gradients near the transition; the hard mask has zero cutoff derivative away from ρ=c . Frequency is measured on the patch-feature grid.
Stage
SD v1.4
BigGAN
VQDM
SD v1.5
Wukong
ADM
GLIDE
Midjourney
Seen
All
T1
99.69
93.78
95.67
99.69
99.74
78.05
95.52
93.10
99.69
94.41
T2
100.00
97.54
97.78
99.93
99.96
83.19
97.59
94.43
98.77
96.30
T3
100.00
97.54
97.77
99.93
99.96
83.19
97.59
94.42
98.44
96.30
T4
100.00
97.54
97.52
99.93
99.96
82.87
97.59
94.31
98.75
96.22
T5
100.00
97.32
97.42
99.93
99.96
81.54
97.03
94.08
98.93
95.91
T6
100.00
97.41
97.54
99.93
99.96
83.96
97.47
94.72
96.47
96.37
Appendix
Table 5: Complete generator-wise detection trajectory on GenImage. Generators arrive in the column order. Seen averages accuracy over the generators observed by the current stage; All averages accuracy over all eight generators, including those not yet observed. Bold marks the highest values in the Seen and All columns. All entries are percentages; averages are computed from the displayed generator-wise results.
Method
SD v1.4
BigGAN
VQDM
SD v1.5
Wukong
ADM
GLIDE
Midjourney
AA ↑
AF ↓
AA ↑
AF ↓
AA ↑
AF ↓
AA ↑
AF ↓
AA ↑
AF ↓
AA ↑
AF ↓
AA ↑
AF ↓
AA ↑
AF ↓
EvoKnow
99.69
–
98.77
−0.31
98.44
0.00
98.75
0.08
98.93
0.14
96.47
0.07
96.53
0.13
96.31
0.11
Appendix
Table 6: Stage-wise continual detection performance on GenImage. Each generator denotes its arrival stage. AA averages accuracy over the generators observed so far; AF measures the average signed performance drop relative to the accuracy obtained immediately after each generator is learned. AA is reported in percentages and AF in percentage points. AF is undefined at the first stage. Final-stage results are bolded.
Table 11
Method
Adaptation budget
GAN
Guided
LDM
GLIDE
DALL-E
Ave. ± Std.
ProG.
Cycle
BigG.
Style
GauG.
StarG.
200
200-cfg
100
100/27
50/27
100/10
CNNSpot
CVPR’20
99.99
85.20
70.20
85.70
78.95
91.70
65.66
60.07
54.03
54.96
54.14
60.78
63.80
55.58
70.05 ± 14.90
PatchFor.
ECCV’20
75.03
68.97
68.47
79.16
64.23
63.94
68.52
67.41
76.50
76.10
75.77
74.81
73.28
67.91
71.44 ± 4.73
Co-occur.
EI’20
97.70
63.15
53.75
92.50
51.10
54.70
69.90
60.50
70.70
70.55
71.00
70.25
69.60
67.55
68.78 ± 12.69
Freq-spec
WIFS’19
49.90
99.90
50.50
49.90
50.30
99.70
50.40
50.90
50.40
50.40
50.30
51.70
51.40
50.00
57.55 ± 17.26
F3Net
ECCV’20
99.38
76.38
65.33
92.56
58.10
100.00
83.05
69.20
68.15
75.35
68.80
81.65
83.25
66.30
77.68 ± 12.48
Appendix
Table 9: UniversalFakeDetect detection accuracy (%) under zero-shot and few-shot knowledge expansion. Existing-detector results follow their reported ProGAN-based training protocols. For EvoKnow, the adaptation budget denotes the number of generated images per arriving generator in the GenImage stream; UniversalFakeDetect is used only for evaluation. The knowledge-unit organization and arrival sequence are fixed across adaptation budgets. The final column reports average accuracy and population standard deviation across the 14 displayed generators; the standard deviation measures cross-generator variation rather than variation across repeated runs.
Organization
MidJ.
SD1.4
SD1.5
ADM
GLIDE
Wuk.
VQDM
BigGAN
Ave.
Generator-based
96.06
98.18
98.53
96.21
96.27
98.64
97.41
93.77
96.88
Category-based
97.39
99.32
99.62
97.30
98.32
98.61
99.21
99.64
98.68
Δ
+1.33
+1.14
+1.09
+1.09
+2.05
-0.03
+1.80
+5.87
+1.79
Appendix
Table 10: Effect of incremental residual knowledge organization on GenImage (%). Generator-based organization assigns one incremental residual expert per generator, whereas category-based organization groups generators into GAN, diffusion, and unknown-source categories. Both settings use ten generated images per arriving generator. Unlike Table 1 , Ave. here includes all eight GenImage domains, including the SD v1.4 base domain.
Method
Real
Big
Mid
Wuk
SD14
SD15
ADM
GLI
VQ
Avg.
ResNet
16.1
10.6
56.3
27.6
12.1
22.6
17.1
20.1
10.1
21.4
DIRE
22.6
11.6
0.0
26.6
22.1
20.6
18.1
22.6
10.6
17.2
ESSP
17.6
13.6
49.7
27.1
19.6
17.6
16.6
27.1
13.1
22.4
LIDA
83.4
98.5
69.3
13.1
23.6
50.3
47.2
55.3
45.7
54.0
EvoKnow (router)
98.5
93.1
29.5
35.7
2.1
65.0
63.5
68.0
38.1
54.8
Appendix
Table 11: Auxiliary source-attribution diagnostic (rank-1 accuracy, %). The evaluation includes one real-image class and eight generator classes. The real label is determined by the binary detection head, whereas generator labels are supplied by the analytical router. Avg. is the unweighted mean over all nine classes. Big, Mid, Wuk, GLI, and VQ abbreviate BigGAN, Midjourney, Wukong, GLIDE, and VQDM, respectively.
As generative models continue to evolve, AI-generated image detectors must incrementally adapt to emerging generative domains while preserving knowledge acquired from previous ones. This continual learning setting is particularly challenging because forensic traces are often subtle and generator-specific, making detectors highly vulnerable to catastrophic forgetting. Existing methods primarily address this problem by stabilizing feature representations, implicitly treating forgetting as a representation-level issue. In this paper, we show that this perspective is incomplete. We demonstrate that even when feature representations remain discriminative, the decision boundary can progressively drift as the classification head is continually optimized on new domains. These two effects jointly give rise to a compound failure mode, termed Dual Degradation. To overcome this challenge, we propose DECODE, a decoupled continual detection framework that jointly mitigates representation- and decision-level forgetting. Specifically, we introduce Subspace Diversity Regularization (SDR) to preserve diverse forensic representations and Closed-Form Decision Alignment (CDA) to recalibrate the shared classification head after each adapter merge without manual hyperparameter tuning. Extensive experiments on 19 generative domains show that DECODE achieves an average accuracy of 99.36% with only 0.39% forgetting, while further generalizing to 11 unseen generators with 95.36% accuracy.
Zihao Cai, Xinghan Li, Ruiyan Yang +3
Visual Laboratory of Fudan University, College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China · CEC GienTech Technology Co., Ltd., Shanghai, China
The rapid advancement of generative Artificial Intelligence (AI) has introduced significant challenges for reliable AI-generated image detection. Existing detectors often suffer from performance degradation under distribution shifts and when encountering newly emerging generative models. In this work, we propose a data-centric continual adaptation framework for updating detectors in evolving environments. We show that both in-the-wild data and generator-driven data are essential for adapting detectors. We introduce an automated, weakly supervised pipeline for constructing in-the-wild datasets through fact-check article retrieval. Additionally, we demonstrate that incorporating even a small amount of generator-driven data during training enables effective adaptation to newly emerging models, while combining it with in-the-wild data within a continual learning framework enables robust adaptation and mitigates catastrophic forgetting. Extensive experiments on two state-of-the-art detectors show significant improvements of +9.14% and +8% in average accuracy, respectively.
The rapid evolution of image generators calls for forensic cues that generalize beyond known generation mechanisms. Existing detectors often rely on static image representations or endpoint reconstruction discrepancies, leaving the evolution of intermediate reconstruction stages underexplored. We observe that the relative token predictability of real and generated images can reverse across reconstruction scales, suggesting that intermediate stages may expose forensic evidence overlooked by endpoint comparisons. Motivated by this observation, we propose RED (Reconstruction Evolution Dynamics), a framework that captures transferable forensic cues from coarse-to-fine reconstruction evolution. To our knowledge, RED is the first framework to use scale-wise token predictability to guide forensic evidence aggregation across intermediate reconstruction states. It represents the reconstruction trajectory produced by a frozen multiscale VQ-VAE in the shared feature space of a frozen CLIP encoder. To connect the observed predictability variations with visual evidence, RED learns image-adaptive stage weights from scale-wise token negative log-likelihoods provided by a frozen VAR model. A cross-stage evidence aggregation module then jointly models the original-image representation and the weighted reconstruction features, capturing complementary forensic cues through interactions along the reconstruction trajectory. Experiments on six diverse benchmarks demonstrate that RED achieves the highest average accuracy of 92.5% and average precision of 97.5% among the evaluated methods. Further evaluations show strong robustness to common image degradations, supporting the value of reconstruction evolution for generalizable AI-generated image detection. The code will be made publicly available upon acceptance of this paper.
Wenpeng Mu, Junshan Jin, Tanfeng Sun +2
School of Computer Science, Shanghai Jiao Tong University, Shanghai, China