Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
Organizations: NSF AI Institute for Artificial Intelligence and Fundamental Interactions, Cambridge, 02139, MA · Laboratory for Nuclear Science, Massachusetts Institute of Technology, Cambridge, 02139, MA · School of Engineering and Applied Sciences, Harvard University, Allston, 02134, MA · Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University, Allston, 02134, MA
Abstract
Modern artificial intelligence has revolutionized how we extract representations from scientific data, yet the statistical properties of these representations remain poorly controlled, causing misspecified anomaly detection methods to falter. The hardest anomalies to detect are the rare, weakly separable ones hiding within the nominal distribution-a regime that grows in importance as models mature and easily separable signals are exhausted. We identify structural desiderata for detection in this regime under minimal prior information: sparsity, to enforce parsimony; locality, to preserve geometric sensitivity; and competition, to promote efficient allocation of model capacity. These principles define a class of self-organizing local kernels that adaptively partition the representation space around regions of statistical imbalance. As an instantiation, we introduce SparKer, a sparse ensemble of Gaussian kernels trained in a semi-supervised Neyman-Pearson framework to locally model the likelihood ratio between a sample that may contain anomalies and an anomaly-free reference. We provide theoretical insights into the mechanisms driving detection and self-organization, and demonstrate the approach on realistic high-dimensional problems in scientific discovery, open-world novelty detection, intrusion detection, and generative-model validation. Ensembles of only a handful of kernels identify statistically significant anomalies in representation spaces of thousands of dimensions while remaining sensitive across regimes, underscoring the interpretability, efficiency, and scalability of the approach.
Figures & tables
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| label | N(S)/N(R) | |||
|---|---|---|---|---|
| Gaussian bumps | bulk | 1.6 | 0.16 | |
| tail | 6.4 | 0.16 | ||
| extreme tail | 9 | 0.16 | ||
| Tail excess | excess |
| label | N(S)/N(R) | |||
|---|---|---|---|---|
| Gaussian bumps | bulk | (-6, -6) | 1 | |
| ood | (0, 0) | 1 | ||
| Scale distortion | smearing | 2.2 | ||
| squeezing | 1.8 |
| mass | N(S)/N(R)* | |
|---|---|---|
| resonant signal | Z’ ( GeV) | |
| Z’ ( GeV) | ||
| Z’ ( GeV) | ||
| non-resonant signal | ||
| *after physics cuts applied |
| Method | bulk | tail | extreme tail | tail excess |
|---|---|---|---|---|
| Distance-based two-sample tests | ||||
| Energy distance | ||||
| Nyström-MMD ( ) | ||||
| Nyström-MMD ( ) | ||||
| -NN two-sample balanced (aggregated) | ||||
| -NN distance (aggregated) | ||||
| Method | bulk | ood | smearing | squeezing |
|---|---|---|---|---|
| Distance-based two-sample tests | ||||
| Energy distance | ||||
| Nyström-MMD ( ) | ||||
| Nyström-MMD ( ) | ||||
| -NN two-sample balanced (aggregated) | ||||
| -NN distance (aggregated) | ||||
| Method | Z’200 (bulk) | Z’300 (tail) | Z’600 (extreme tail) | EFT (non-resonant) |
|---|---|---|---|---|
| Distance-based two-sample tests | ||||
| Nyström-MMD ( ) | ||||
| Nyström-MMD ( ) | ||||
| -NN two-sample balanced (aggregated) | ||||
| -NN distance (aggregated) | ||||
| Density-based methods (unsup.) | ||||
| Method | bulk | tail | extreme_tail | tail_excess |
|---|---|---|---|---|
| Unsupervised point-wise detectors | ||||
| ECOD | ||||
| Isolation Forest | ||||
| -NN distance | ||||
| SOM quant. error ( ) | ||||
| Semi-supervised | ||||
| Method | bulk | ood |
|---|---|---|
| Unsupervised point-wise detectors | ||
| ECOD | ||
| Isolation Forest | ||
| -NN distance | ||
| SOM quant. error ( ) | ||
| Semi-supervised | ||
| 1D benchmarks | ||||
|---|---|---|---|---|
| Method | bulk | tail | extreme tail | tail excess |
| SOM quant. error ( ) | ||||
| SOM quant. error ( ) | ||||
| SOM quant. error ( ) | ||||
| SOM quant. error ( ) | ||||
| SOM quant. error ( -aggregated) | ||||
| NP | |
| BCE | |
| MSE | |
| Loss function | MSE | BCE | NP |
|---|---|---|---|
| Power@5% | 0.61 | 0.68 | 0.86 |
| Model | w/o | w/ |
|---|---|---|
| Power@5% | 0.72 | 0.86 |
| Benchmark | Signal | ||||
|---|---|---|---|---|---|
| 1D exponential | bulk | 0.94 [0.91, 0.96] | 0.92 [0.88, 0.94] | 0.83 [0.79,0.86] | 0.85 [0.81, 0.88] |
| tail | 0.67 [0.62, 0.71] | 0.80 [0.75, 0.83] | 0.74 [0.69,0.78] | 0.76 [0.71, 0.80] | |
| extreme tail | 0.62 [0.57, 0.66] | 0.86 [0.82, 0.89] | 0.82 [0.78,0.85] | 0.86 [0.82, 0.89] | |
| tail excess | 0.98 [0.95, 0.99] | 0.94 [0.91, 0.96] | 0.92 [0.88,0.94] | 0.95 [0.92, 0.96] | |
| 2D mixture | bulk | 0.25 [0.20, 0.29] | 0.27 [0.22, 0.31] | 0.31 [0.18, 0.40] | – |
| ood | 0.58 [0.53, 0.63] | 0.50 [0.44, 0.54] | 0.22 [0.18, 0.26] | – |
| Application | L2 coeff. | Amplitude clip | |
|---|---|---|---|
| Gravitational waves detection | 100k | 1 | 100 |
| New particle jets discovery | 100k | 100 | |
| Hybrid butterfly discovery | 100k | 200 | |
| intrusion detection in Dante’s text | 100k | 500 | |
| OOD in ImageNet | 40k | 0 | 10k |
| CIFAR-10 vs. CIFAR-5m | 40k | 0 | 10k |