Graph Anomaly Detection as Finite-Horizon Control: Training-Free Scoring via Empirical Bayes
Organizations: University of California, Los Angeles · Block · Mila – Quebec AI Institute
Abstract
Node-level graph anomaly detection (GAD) identifies nodes whose attributes and interactions deviate from dominant graph regularities. Existing GAD models encode normality and anomaly scoring indirectly through architectures, message passing, reconstruction or contrastive objectives, and tuned score families. This entangles graph trust (how strongly graph structure should define normality), graph-spectral weighting, and anomaly-score choice, yielding scores that are costly, opaque, and unstable across graph regimes. We propose EB-GAD (Empirical-Bayes GAD), a training-free framework that models normality as graph-aware generalized Ornstein-Uhlenbeck (GOU) relaxation toward a graph-filtered template. Empirical Bayes fits the graph precision from the residual-field likelihood; the GOU then turns scoring into a closed-form finite-horizon control energy, the minimum effort to steer a feature-neutral node to its observed endpoint along graph-spectral relaxation. Sweeping relaxation horizon and endpoint tolerance yields a bank of scores that share one fitted prior: equilibrium Mahalanobis scoring is one limit, while finite-horizon control-energy and scale-normalized ratio scores reveal anomalies that static equilibrium scoring can mask. A label-free selector chooses the score family from feature homophily, edge density, and feature dimension, then ranks candidates by fitted-null deviation and rank stability. On 11 benchmarks and without labels at any step, EB-GAD has the best or tied-best AUROC on 9: the four financial fraud networks (up to 3.7M nodes), the YelpChi and Amazon review graphs, Weibo, Reddit and Facebook, with margins of up to 21.7 points. It is second on BlogCatalog and ACM.
Figures & tables
| Dataset | LOF | DIF | ANOM. | DOMINANT | AnomDAE | CONAD | CoLA | DiffGAD | TAM | EB-GAD |
|---|---|---|---|---|---|---|---|---|---|---|
| ∗ | ∗ | |||||||||
| Amazon | ∗ | ∗ | ∗ | ∗ | ||||||
| YelpChi | ∗ | ∗ | ∗ | ∗ | ||||||
| BlogCatalog | ∗ | ∗ | ∗ | ∗ | ||||||
| ∗ | ∗ | ∗ | ∗ |
| Amazon | YelpChi | BlogCat. | ACM | Elliptic | Ell.++ | DGraph | T-Fin. | ||||
| Family | anchor | anchor | energy | anchor | energy | profile | ratio | anchor | anchor | anchor | profile |
| Equilibrium | 95.1 | 60.6 | 77.9 | 71.7 | 78.4 | 87.8 | 81.3 | 74.5 | 72.5 | 66.8 | 80.1 |
| Selected | 95.1 | 60.6 | 78.0 | 71.7 | 78.6 | 91.4 | 86.0 | 74.5 | 72.5 | 66.8 | 84.8 |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | #Nodes | #Edges | #Feat | Anom.% |
| Standard benchmarks (Table 1 ) | ||||
| 8,405 | 407,963 | 400 | 4.1 | |
| 10,984 | 168,016 | 64 | 3.3 | |
| Amazon | 11,944 | 4,398,392 | 25 | 6.9 |
| YelpChi | 45,954 | 3,846,979 | 32 | 4.9 |
| BlogCatalog | 5,196 | 171,743 | 8,189 | 5.8 |
| Dataset | Baseline | Native AUROC | Oracle-flipped AUROC | Mean score normal/anomaly |
|---|---|---|---|---|
| Elliptic | DIF | |||
| Elliptic++ | DIF | |||
| Elliptic | DiffGAD | |||
| Elliptic++ | DiffGAD |
| Dataset | Template | PCA | Rule | Selected family (score) | ||||
|---|---|---|---|---|---|---|---|---|
| low-pass | all | 1.000 | 1 | 0.5 | 64 | 7 | equilibrium anchor ( ) | |
| low-pass | all | 1.000 | 1 | 2 | no | 5 | equilibrium anchor ( ) | |
| Amazon | low-pass | all | 0.458 | 0.5 | 5 | no | 5 | control energy (top-2 fusion) |
| YelpChi | low-pass | 500 | 1.000 | 20 | 1 | no | 5 | equilibrium anchor ( ) |
| BlogCatalog | affinity | all | 1.000 | 1 | 0.18 | 48 | 2 | control energy (stability-ranked fusion) |
| low-pass | all | 0.727 | 3 | 0.7 | 64 | 6 | profile aggregation (scale-entropy) |
| Dataset | NullKS( ) | NullKS( ) | BJ | NullKS(profile) | |||||
|---|---|---|---|---|---|---|---|---|---|
| 0.068 | 24.27 | 400 | all | 0.663 | 0.156 | 78.5 | 0.258 | 0.786 | |
| 0.993 | 7.65 | 64 | all | 0.945 | 0.138 | 57.6 | 0.206 | 0.878 | |
| Amazon | 0.644 | 17.18 | 25 | all | 0.277 | 0.071 | 236.2 | 0.208 | 0.740 |
| YelpChi | 0.876 | 2.07 | 32 | 0.000 | 0.000 | 1225.2 | 0.334 | 0.024 | |
| BlogCatalog | 0.009 | 33.25 | 8189 | all | 0.673 | 0.089 | 169.9 | 0.246 | 0.077 |
| 0.375 | 25.49 | 576 | all | 0.036 | 0.044 | 8.3 | 0.227 | 0.811 |
| Quantity | Class | Source or tested range |
|---|---|---|
| ; within its grid | model-derived | maximum of the residual-space likelihood (Alg. 1 ) |
| , , , | model-derived | closed forms of § 3.3 |
| plug-in nulls of the scores | model-derived | chi-squared laws under the fitted prior |
| Kolmogorov–Smirnov, Berk–Jones | canonical | [ 2 ] |
| ACAT, minimum- , mean combiners | canonical | [ 20 ] ; Fisher’s method |
| two-groups mixture of -values | canonical | [ 8 ] |
| Amazon | YelpChi | BlogCat. | ACM | Elliptic | Ell.++ | DGraph | T-Fin. | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0.5 | 2 | 1 | 1 | 0.2 | 0.5 | 0.7 | 1 | 2 | 1 | 0.2 | |
| 95.1 | 62.0 | 75.8 | 71.7 | 54.0 | 47.4 | 75.6 | 58.3 | 30.2 | 60.5 | 81.6 | |
| 47.3 | 60.6 | 48.2 | 64.6 | 61.0 | 89.7 | 72.1 | 74.5 | 72.5 | 66.8 | 80.1 | |
| Table 1 | 95.1 | 60.6 | 78.0 | 71.7 | 78.6 | 91.4 | 86.0 | 74.5 | 72.5 | 66.8 | 84.8 |
| Dataset | ANOM. | DOMINANT | AnomDAE | CONAD | CoLA | DiffGAD | TAM | EB-GAD (CPU) |
|---|---|---|---|---|---|---|---|---|
| 78 | 63 | 1 | 95 | 130 | 134 | 858 | 84 | |
| 38 | 35 | 2 | 57 | 8 | 110 | 5,467 | 136 | |
| Amazon | 25 | 37 | 153 | 308 | 14 | 106 | 10,739 | 708 |
| YelpChi | 170 | 185 | 1,152 | 1,359 | 5 | 118 | 5,106 | 27 |
| BlogCatalog | 447 | 430 | 92 | 3,047 | 18 | 3,272 | 8,764 | 675 |
| 3 | 20 | 9 | 56 | 3 | 118 | 479 | 8 |
| Dataset | Training-free | Best trained encoder | Gap |
|---|---|---|---|
| 95.1 | 95.7 (MLP) | ||
| 60.6 | 58.9 (MLP) | ||
| Amazon | 78.0 | 60.3 (MLP) | |
| YelpChi | 71.7 | 68.0 (SAGE) | |
| BlogCatalog | 78.6 | 77.6 (MLP) | |
| 91.4 | 95.4 (MLP) |
| Dataset | Spectrum | Orders | min | max | std | Spearman |
|---|---|---|---|---|---|---|
| full | 2 | 95.12 | 95.12 | 0.00 | 1.0000 | |
| full | 2 | 60.56 | 60.56 | 0.00 | 0.9995 | |
| Amazon | full | 2 | 78.05 | 78.05 | 0.00 | 1.0000 |
| YelpChi | truncated | 5 | 70.72 | 72.65 | 0.62 | 0.8368 |
| BlogCatalog | full | 2 | 78.62 | 78.62 | 0.00 | 1.0000 |
| full | 2 | 91.40 | 91.40 | 0.00 | 1.0000 |
| ties by node index | averaged ties | |||||
| Dataset | Score | released | relabeled | released | relabeled | node index |
| YelpChi | profile, mixture posterior | 93.7 | 60.3 | 67.3 | 66.8 | 94.4 |
| profile, mean | 93.5 | 60.6 | 68.7 | 68.3 | ||
| profile, ACAT | 92.8 | 61.9 | 69.3 | 69.8 | ||
| 70.9 | 70.7 | 70.9 | 70.7 | |||
| 64.2 | 64.0 | 64.2 | 64.0 | |||
| Dataset | Protocol | Anchor, likelihood | Anchor, averaged | Nullspace dropped | Jacobian |
|---|---|---|---|---|---|
| 95.1 | 92.5 | 95.3 | 95.2 | 92.5 | |
| 60.6 | 48.5 | 60.7 | 60.4 | 61.9 | |
| Amazon | 78.0 | = | = | 76.5 | = |
| YelpChi | 70.7 | = | 64.0 | – | 69.1 |
| BlogCatalog | 78.6 | = | = | 72.0 | = |
| 91.4 | = | = | 90.8 | 85.8 |
| Dataset | : AUROC (removed share), in increasing | |||
|---|---|---|---|---|
| 0.36 | 0.2: 95.6 (0) | 0.5: 95.1 (59) | 0.7: 92.5 ℓ (86) | |
| 1.49 | 0.7: 48.5 ℓ (89) | 1: 51.8 (87) | 2: 60.6 (39) | |
| Amazon | 0.87 | 0.5: 53.5 (12) | 0.7: 63.8 ℓ (79) | 1: 75.8 (77) |
| YelpChi | 1.38 | 0.7: 63.3 (0) | 1: 70.9 ℓ (3) | 2: 69.1 (1) |
| BlogCatalog | 0.28 | 0.1: 44.9 (22) | 0.2: 61.0 (92) | 0.5: 74.8 ℓ (95) |
| 0.55 | 0.5: 47.4 (67) | 0.7: 87.8 ℓ (86) | 1: 85.9 (77) | |
| Method | Amazon | YelpChi | BlogCat. | ACM | Elliptic | Ell.++ | DGraph | T-Fin. | |||
|---|---|---|---|---|---|---|---|---|---|---|---|
| LOF | |||||||||||
| DIF | |||||||||||
| ANOMALOUS | ∗ | ∗ | ∗ | ∗ | ∗ | OOM | OOM | OOM | |||
| DOMINANT | ∗ | ∗ | ∗ | ∗ | ∗ | OOM | |||||
| AnomalyDAE | OOM | ||||||||||
| CONAD | n/a |
| Dataset | GOU-EB | GOU fixed | heat | polynomial | identity |
|---|---|---|---|---|---|
| 92.5 (92.7) | 92.0 (92.7) | 85.9 (92.5) | 80.1 (91.4) | 92.7 (92.7) | |
| 48.2 (55.8) | 48.3 (55.9) | 48.6 (57.4) | 51.5 (56.3) | 48.2 (48.2) | |
| Amazon | 62.9 (67.7) | 62.9 (67.8) | 50.8 (67.4) | 69.7 (70.1) | 62.8 (62.8) |
| YelpChi | 69.1 (71.3) | 69.1 (71.3) | 66.1 (70.5) | 68.7 (70.5) | 68.7 (70.5) |
| BlogCatalog | 78.2 (78.5) | 78.2 (78.5) | 78.3 (78.3) | 77.6 (77.6) | 78.5 (78.5) |
| 80.0 (88.0) | 79.8 (88.0) | 91.6 (92.8) | 80.9 (87.7) | 48.1 (62.0) |
| Dataset (type) | Equilibrium | Finite, selected | Finite, best | selected | best |
|---|---|---|---|---|---|
| Weibo ( ) | 95.1 | 95.1 | 95.1 | ||
| Reddit ( ) | 60.6 | 60.6 | 60.6 | ||
| Amazon ( ) | 77.9 | 78.0 | 78.1 | ||
| YelpChi ( ) | 70.7 | 70.7 | 70.7 | ||
| BlogCatalog ( ) | 78.4 | 78.4 | 78.6 | ||
| Facebook ( ) | 87.8 | 87.8 | 88.0 |
| Cutoff | Factor | Dataset | Selection | before | after | change |
|---|---|---|---|---|---|---|
| (rule 1) | Elliptic++ | anchor ( ) anchor ( ) | 72.5 | 30.3 | ||
| (rule 1) | DGraph | anchor ( ) anchor ( ) | 66.8 | 60.5 | ||
| (rule 1) | Elliptic++ | anchor ( ) anchor ( ) | 72.5 | 30.3 | ||
| profile ( entropy ) anchor ( ) | 91.4 | 47.4 | ||||
| profile ( entropy ) anchor ( ) | 91.4 | 47.4 | ||||
| ACM | ratio ( aff. ) anchor ( ) | 86.0 | 75.6 |
| Dataset | Bank | NullKS | stability | tail-stab. | corr-stab. | rule | best |
|---|---|---|---|---|---|---|---|
| equilibrium | 95.1 | 74.7 | 90.8 | 89.6 | 95.1 | 95.1 | |
| equilibrium | 60.6 | 62.9 | 62.8 | 62.9 | 60.6 | 62.9 | |
| YelpChi | equilibrium | 68.7 | 67.9 | 70.7 | 70.7 | 70.7 | 70.7 |
| two-groups | 17.6 | 25.5 | 91.4 | 69.4 | 91.4 | 91.4 | |
| Elliptic | equilibrium | 74.4 | 67.2 | 58.4 | 67.4 | 74.4 | 74.4 |
| Elliptic++ | equilibrium | 30.3 | 52.3 | 30.3 | 54.5 | 72.5 | 72.5 |
| Dataset | Scores | Selector | Best single | Stacked (CV) | Gap |
|---|---|---|---|---|---|
| 224 | 95.1 | 95.1 | 95.2 | ||
| 224 | 60.6 | 62.0 | 64.9 | ||
| Amazon | 887 | 78.0 | 78.1 | 78.3 | |
| YelpChi | 158 | 70.7 | 73.0 | 74.9 | |
| BlogCatalog | 4391 | 78.6 | 78.7 | 79.5 | |
| 224 | 91.4 | 91.4 | 92.3 |
| Dataset | Injected | Total contamination | Fitted | Original anomalies | Injected anomalies |
|---|---|---|---|---|---|
| 4.1% | (0.697, 3) | 92.5 | – | ||
| 9.1% | (0.710, 3) | 92.0 | 57.1 | ||
| 14.1% | (0.720, 3) | 92.1 | 57.2 | ||
| 24.1% | (0.737, 3) | 91.4 | 56.1 | ||
| 3.3% | (0.910, 3) | 48.3 | – | ||
| 8.3% | (0.997, 3) | 48.3 | 79.9 |
| Residual distribution | Selected | Best member | Fitted |
|---|---|---|---|
| Gaussian (control) | 96.3 / 95.7 / 95.5 | 98.8 / 98.1 / 97.6 | 0.05 / 0.05 / 0.05 |
| Student- , 3 d.o.f. | 95.6 / 95.7 / 95.3 | 97.4 / 97.5 / 96.8 | 0.05 / 0.05 / 0.05 |
| two-component mixture | 80.2 / 86.7 / 76.2 | 87.0 / 91.6 / 82.7 | 0.05 / 0.05 / 0.05 |
| binarized | 59.3 / 56.1 / 56.4 | 60.8 / 58.1 / 61.1 | 0.05 / 0.05 / 0.05 |
| binarized, bit-flip anomalies | 59.1 / 60.5 / 59.4 | 63.4 / 60.9 / 63.7 | 0.05 / 0.05 / 0.05 |
| Placement | Selected, seeds 0–4 | Range | Best member, range |
|---|---|---|---|
| scattered | 79.9 / 81.3 / 82.7 / 80.5 / 83.0 | 79.9–83.0 | 79.9–83.0 |
| connected groups of at most 10 | 92.8 / 92.2 / 93.6 / 92.9 / 92.0 | 92.0–93.6 | 92.0–93.6 |
| Edge score | Seeds 0–4 | Range |
|---|---|---|
| smaller endpoint energy | 95.9 / 96.0 / 94.4 / 96.7 / 94.4 | 94.4–96.7 |
| standardized cross-term | 83.0 / 84.0 / 77.5 / 83.4 / 80.2 | 77.5–84.0 |