cs.LGNov 5, 2025

Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies

Authors: Gaia Grosso, Sai Sumedh R. Hindupur, Thomas Fel, Samuel Bright-Thonney, Philip Harris, Demba Ba

Organizations: NSF AI Institute for Artificial Intelligence and Fundamental Interactions, Cambridge, 02139, MA · Laboratory for Nuclear Science, Massachusetts Institute of Technology, Cambridge, 02139, MA · School of Engineering and Applied Sciences, Harvard University, Allston, 02134, MA · Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University, Allston, 02134, MA

Abstract

Modern artificial intelligence has revolutionized how we extract representations from scientific data, yet the statistical properties of these representations remain poorly controlled, causing misspecified anomaly detection methods to falter. The hardest anomalies to detect are the rare, weakly separable ones hiding within the nominal distribution-a regime that grows in importance as models mature and easily separable signals are exhausted. We identify structural desiderata for detection in this regime under minimal prior information: sparsity, to enforce parsimony; locality, to preserve geometric sensitivity; and competition, to promote efficient allocation of model capacity. These principles define a class of self-organizing local kernels that adaptively partition the representation space around regions of statistical imbalance. As an instantiation, we introduce SparKer, a sparse ensemble of Gaussian kernels trained in a semi-supervised Neyman-Pearson framework to locally model the likelihood ratio between a sample that may contain anomalies and an anomaly-free reference. We provide theoretical insights into the mechanisms driving detection and self-organization, and demonstrate the approach on realistic high-dimensional problems in scientific discovery, open-world novelty detection, intrusion detection, and generative-model validation. Ensembles of only a handful of kernels identify statistically significant anomalies in representation spaces of thousands of dimensions while remaining sensitive across regimes, underscoring the interpretability, efficiency, and scalability of the approach.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Isolation-based Spherical Ensemble Representations for Tabular Anomaly Detection

    Oct 15, 2025Yang Cao, Sikun Yang, Hao Tian +5Graph Anomaly DetectionInterpretable Anomaly Detection

  2. Kurtosis-Guided Denoising Score Matching for Tabular Anomaly Detection

    May 7, 2026Victor Livernoche, Jie Zan, Reihaneh RabbanyGraph Anomaly DetectionInterpretable Anomaly Detection

  3. Can We Predict Anomaly Detection Performance from Embedding-Space Geometry?

    Sep 22, 2026Kevin Wilkinghoff, Zheng-Hua TanArea Under The Receiver Operating Characteristic CurveDetection Performance