Is Real-World Training Data Necessary for Generalist Graph Anomaly Detection?
Authors: Yujing Liu, Yixin Liu, Yue Tan, Xiaofeng Cao, Alan Wee-Chung Liew, Heng Tao Shen, Shirui Pan
Organizations: School of Information and Communication Technology, Griffith University, Australia · School of Computer Science and Technology, Tongji University, China
Generalist graph anomaly detection (GAD) aims to build a foundation model that detects anomalies on arbitrary unseen graphs without retraining or fine-tuning. Sufficient data are essential for foundation model training, yet generalist GAD still faces a data shortage, as real-world anomalous graphs are scarce and costly to collect and annotate. To fill this gap, we propose AG-FORGE, an Anomalous Graph generation Forge for automatic synthesis of anomalous graphs, exploring the feasibility of synthetic data-driven training for generalist GAD. Empirically, we find that synthetic data can achieve performance comparable to real-world training, but fail to push the performance boundary further due to the limited capacity of existing methods. To further unlock model capacity as training data scale up, we develop TS-GGAD, a Topology-Semantic coordinated Generalist GAD that captures complementary topological and semantic anomaly evidence, together with a curriculum learning strategy tailored to large-scale synthetic training. Extensive experiments on 14 real-world datasets demonstrate that TS-GGAD, trained on data generated by AG-FORGE, significantly outperforms state-of-the-art methods.
Figures & tables
Figure 1: Real-world vs. synthetic data for model training.
Figure 2: The framework of the proposed AG-Forge , TS-GGAD , and curriculum learning strategy.
Method
Cite
CS
ACM
Blog
Amz
Photo
Weibo
Cora
Pubmed
Flickr
FB
Yelp
Quest
Reddit
Rank
GAD Methods
GCN
48.32
56.19
50.43
43.06
58.69
47.70
46.40
32.41
33.68
38.07
76.35
51.21
41.37
46.80
16.07
GAT
63.11
59.25
61.48
60.50
54.84
45.63
73.51
56.07
67.21
56.05
55.33
50.43
57.37
43.39
14.64
CoLA
73.81
65.99
54.94
60.58
65.02
65.18
41.08
66.02
70.09
60.64
70.88
52.41
53.23
50.60
12.43
SmoothGNN
90.42
78.65
78.00
73.06
50.75
44.80
85.96
88.39
78.22
76.95
53.62
61.30
59.01
54.09
8.64
ANEMONE
41.36
42.18
45.53
35.18
43.51
49.22
35.57
47.08
37.65
37.07
45.55
50.88
51.55
52.41
17.29
Table 1: Detection performance in terms of AUROC (%). Highlighted are the results ranked first , second , and third . Purple-shaded rows denote models trained with AG-Forge -generated data.
Figure 3: Ablation for AG-Forge .
Figure 4: Ablation for TS-GGAD .
Figure 5: Scaling analysis.
Figure 6: Sensitivity to K .
Figure 7: Visualization.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Method
ARC
UNPrompt
IA-GGAD
ReFi-GAD
Ours
Avg. Time (s)
1.4
5.62
3.34
1.88
1.96
Appendix
Table 2: Average inference time across 14 real-world datasets.
Dataset
#Nodes
#Edges
#Features
Avg. Degree
#Anomaly
%Anomaly
Citation Networks
Cora
2,708
11,604
1,433
4.29
150
5.54
Citeseer
3,327
10,154
3,703
3.05
150
4.51
ACM
16,484
147,866
8,337
8.97
597
3.62
Pubmed
19,717
92,846
500
4.71
600
3.04
CS
18,333
167,986
6,805
9.16
600
3.27
Appendix
Table 3: The statistics of datasets organized by application domains.
Method
Citeseer
CS
ACM
BlogCatalog
Amazon
Photo
Weibo
Rank
GAD Methods
GCN
48.32 ± 1.03
56.19 ± 1.23
50.43 ± 1.04
43.06 ± 0.92
58.69 ± 0.17
47.70 ± 2.10
46.40 ± 1.82
16.29
GAT
63.11 ± 1.92
59.25 ± 0.94
61.48 ± 0.94
60.50 ± 0.73
54.84 ± 2.68
45.63 ± 1.39
73.51 ± 2.66
14.00
CoLA
73.81 ± 2.87
65.99 ± 2.30
54.94 ± 4.14
60.58 ± 5.07
65.02 ± 10.86
65.18 ± 1.52
41.08 ± 7.06
12.57
SmoothGNN
90.42 ± 8.81
78.65 ± 5.25
78.00 ± 4.92
73.06 ± 1.91
50.75 ± 3.01
44.80 ± 5.02
85.96 ± 4.50
10.00
ANEMONE
41.36 ± 2.40
42.18 ± 1.19
45.53 ± 3.24
35.18 ± 3.43
43.51 ± 4.29
49.22 ± 4.87
35.57 ± 4.75
18.29
Appendix
Table 4: Anomaly detection performance in terms of AUROC (%) ± std. Highlighted are the results ranked first , second , and third . Purple-shaded rows denote models trained with AG-Forge -generated data.
Method
Citeseer
CS
ACM
BlogCatalog
Amazon
Photo
Weibo
Rank
GAD Methods
GCN
7.63 ± 2.98
8.35 ± 3.53
7.10 ± 3.33
4.95 ± 0.16
8.52 ± 0.16
5.48 ± 0.39
9.95 ± 1.01
15.71
GAT
7.37 ± 0.61
6.85 ± 0.53
8.06 ± 0.55
13.29 ± 0.55
7.95 ± 0.45
4.92 ± 0.14
53.35 ± 3.83
14.29
CoLA
15.34 ± 2.30
21.09 ± 3.70
18.16 ± 2.39
25.23 ± 2.36
12.99 ± 4.17
10.98 ± 1.56
24.57 ± 6.31
9.71
SmoothGNN
41.39 ± 13.91
17.06 ± 3.50
16.81 ± 4.82
20.33 ± 2.75
8.70 ± 2.24
5.16 ± 0.73
43.60 ± 7.87
12.14
ANEMONE
4.43 ± 0.77
3.18 ± 0.25
3.58 ± 0.17
4.54 ± 0.28
6.15 ± 0.52
6.76 ± 0.53
9.25 ± 0.84
18.29
Appendix
Table 5: Anomaly detection performance in terms of AUPRC (%) ± std. Highlighted are the results ranked first , second , and third . Purple-shaded rows denote models trained with AG-Forge -generated data.
Variant
Cite
CS
ACM
Blog
Amz
Photo
Weibo
Cora
Pubmed
Flickr
FB
Yelp
Quest
Reddit
Average
Structure Generation
w/o CS
96.17
99.11
95.03
73.64
71.36
74.47
90.79
95.77
97.19
80.04
96.64
74.60
60.47
57.75
83.07
w/o SF
95.27
99.20
94.69
71.71
83.78
76.77
90.16
95.88
97.21
76.81
96.90
74.44
64.78
62.24
84.27
w/o PB
96.37
98.71
95.08
72.72
81.40
76.81
92.43
96.14
97.54
79.43
96.13
74.71
64.06
62.52
84.58
w/o SW
95.23
99.17
95.63
74.53
84.28
77.91
92.32
95.91
97.34
81.83
97.15
74.75
63.86
63.13
85.22
Feature Generation
Appendix
Table 6: Ablation study of different components in AG-Forge . Performance is evaluated in terms of AUROC (%).
Variant
Cite
CS
ACM
Blog
Amz
Photo
Weibo
Cora
Pubmed
Flickr
FB
Yelp
Quest
Reddit
Average
- Topo
62.34
63.25
57.01
64.37
91.04
69.88
83.50
64.16
57.59
58.80
69.12
57.91
62.19
64.63
66.13
- Attr
95.43
98.95
95.24
71.90
68.06
70.75
89.25
95.64
97.06
79.09
95.51
74.61
59.79
57.18
82.03
- Curr
95.60
98.69
93.94
69.96
65.98
73.78
88.52
95.06
97.28
81.55
92.79
69.20
56.91
57.21
81.18
- Gat
95.84
98.92
94.58
75.47
90.25
79.25
92.43
95.50
96.96
79.14
96.20
73.90
65.55
64.90
85.64
Ours
96.52
99.43
96.24
76.19
87.10
82.15
93.53
96.41
97.77
86.46
97.41
74.46
64.62
64.31
86.61
Appendix
Table 7: Ablation study of different components in TS-GGAD . Performance is evaluated in terms of AUROC (%).
Figure 8: Sensitivity to γ of TS-GGAD .
Figure 9: Visualization of node embeddings.
Figure 10: Visualization of anomaly evidence coordination.
ID
Structure
Feature
Anomaly
Difficulty
N
D
Na
Ratio (%)
0
CS
Sparse
LC
0.3
13,058
1,292
956
7.32
1
SF
Sparse
LC
0.7
8,151
1,048
619
7.59
2
CS
Dense
Mix
0.9
12,888
560
813
6.31
3
PB
Binary
Mix
0.9
12,549
217
891
7.10
4
SW
Binary
AS
0.1
11,443
1,888
649
5.67
5
SW
Dense
Mix
0.5
10,720
1,619
755
7.04
Appendix
Table 8: Configurations and statistics of example synthetic graphs generated by AG-Forge .
Driven by the pressing demand for graph anomaly detection (GAD) in high-stakes domains, the generalist GAD paradigm, which trains a single detector transferable across new graphs, has recently gained growing attention. However, existing methods often rely on scarce and costly annotations for training and sometimes even require few-shot support at inference, which limits their robustness to diverse and unseen anomaly patterns. To address this limitation, we introduce ProMoS, the first unsupervised generalist GAD framework, which detects anomalies by modeling the abundant normality in unlabeled data. ProMoS adopts a knowledge-distillation paradigm to distill normality priors from a frozen self-supervised graph neural network (GNN) teacher to a mixture-of-students model with shared global and lightweight personalized branches, enabling efficient and expressive normality modeling without learning from scratch. We further propose prototype-guided soft-label distillation to align teacher and student in a shared prototype space, enhancing cross-graph generalizability. During inference, ProMoS performs zero-shot anomaly detection on unseen graphs via distillation bias and prototype geometric deviation. Extensive experiments show the effectiveness and efficiency of ProMoS, charting a practical path toward label-free, zero-shot generalist GAD.
Yiming Xu, Zihan Chen, Zhen Peng +4
School of Computer Science and Technology, Xi’an Jiao-tong University, Xi’an, China · National Engineering Research Center for Visual Information and Applications, Xi’an, China · University of Virginia, Charlottesville, USA +3
Cross-domain graph anomaly detection (GAD) aims to identify abnormal nodes in unseen target graphs, showing strong potential in real-world applications with heterogeneous graph data. However, existing methods often depend on dataset-specific feature semantics and structural patterns, which limits their ability to generalize across different domains. To address this challenge, we propose AlignGAD, a zero-shot generalized graph anomaly detection framework. Our framework is built upon three key components: a Global Unification Module that aligns heterogeneous node features and normalizes graph signals in the spectral domain; a Clustering Module that constructs cluster-aware graph views to capture group-level abnormal patterns; and a Node Discrepancy Scoring Module that measures reconstruction discrepancy and aggregates anomaly evidence from different graph views. Experiments on multiple real-world datasets demonstrate the effectiveness of AlignGAD under the zero-shot GAD setting.
Graph Anomaly Detection (GAD) is increasingly shifting to Generalist GAD (GGAD) for cross-domain "one-for-all" detection, but existing GGAD methods predominantly rely on the neighbor consistency principle, falling into the \textbf{Node-to-Neighbor Consistency Paradigm} for anomaly quantification. These methods suffer from complex training pipelines, heavy training data dependency, high computational costs, and unstable cross-domain generalization. To address these limitations, we propose NeighborDiv, a training-free generalist graph anomaly detection framework based on neighbor diversity. Departing from the dominant Node-to-Neighbor Consistency Paradigm, we shift the focus to the \textbf{Neighbor-to-Neighbor Diversity Paradigm}, and uncover that the internal structural dispersion of a node's neighbor set is a powerful, independently discriminative anomaly signal. We quantify neighbor diversity via the variance of inter-neighbor feature similarities, which captures how a node organizes its local graph environment, and operates independently of conventional node-to-neighbor consistency frameworks. Extensive experiments under two standard GGAD evaluation paradigms show NeighborDiv achieves state-of-the-art performance, with relative gains of 10.25% in average AUC and 17.78% in average AP over the second-best baseline under Single-Domain Independent Training (SDIT), and 6.89%/9.58% in AUC/AP under Unified Multi-Domain Training (UMDT), respectively. Notably, NeighborDiv yields zero performance volatility across all datasets, eliminating training-set dependency and establishing a lightweight and highly practical GGAD framework.
Kaifeng Wei, Teng Liu, Liang Dong +2
School of Software Technology Zhejiang University Ningbo, China · School of Advanced Technology Xi’an Jiaotong-Liverpool University Suzhou, China