cs.LGNov 5, 2025

Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies

Authors: Gaia Grosso, Sai Sumedh R. Hindupur, Thomas Fel, Samuel Bright-Thonney, Philip Harris, Demba Ba

Organizations: NSF AI Institute for Artificial Intelligence and Fundamental Interactions, Cambridge, 02139, MA · Laboratory for Nuclear Science, Massachusetts Institute of Technology, Cambridge, 02139, MA · School of Engineering and Applied Sciences, Harvard University, Allston, 02134, MA · Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University, Allston, 02134, MA

Abstract

Modern artificial intelligence has revolutionized how we extract representations from scientific data, yet the statistical properties of these representations remain poorly controlled, causing misspecified anomaly detection methods to falter. The hardest anomalies to detect are the rare, weakly separable ones hiding within the nominal distribution-a regime that grows in importance as models mature and easily separable signals are exhausted. We identify structural desiderata for detection in this regime under minimal prior information: sparsity, to enforce parsimony; locality, to preserve geometric sensitivity; and competition, to promote efficient allocation of model capacity. These principles define a class of self-organizing local kernels that adaptively partition the representation space around regions of statistical imbalance. As an instantiation, we introduce SparKer, a sparse ensemble of Gaussian kernels trained in a semi-supervised Neyman-Pearson framework to locally model the likelihood ratio between a sample that may contain anomalies and an anomaly-free reference. We provide theoretical insights into the mechanisms driving detection and self-organization, and demonstrate the approach on realistic high-dimensional problems in scientific discovery, open-world novelty detection, intrusion detection, and generative-model validation. Ensembles of only a handful of kernels identify statistically significant anomalies in representation spaces of thousands of dimensions while remaining sensitive across regimes, underscoring the interpretability, efficiency, and scalability of the approach.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Oct 15, 2025cs.LG

Isolation-based Spherical Ensemble Representations for Tabular Anomaly Detection

Unsupervised tabular anomaly detection is a critical task with applications spanning offensive language detection, network security, and quality control. Despite extensive research, existing unsupervised anomaly detection methods still face fundamental challenges including conflicting distributional assumptions, computational inefficiency, and difficulty handling different anomaly types. To address these problems, we propose ISER (Isolation-based Spherical Ensemble Representations) that extends existing isolation-based methods by using hypersphere radii as a monotonic transformation of local density characteristics while maintaining linear time and constant space complexity w.r.t. the dataset size. ISER constructs ensemble representations where hypersphere radii encode local sparsity through a monotonic transformation of density: smaller radii correspond to dense regions while larger radii correspond to sparse regions. We introduce a novel similarity-based scoring method that measures pattern consistency by comparing ensemble representations against a theoretical anomaly reference pattern. Additionally, we enhance the performance of Isolation Forest by using ISER and adapting the scoring function to address axis-parallel bias and local anomaly detection limitations. Comprehensive experiments on 20 real-world datasets demonstrate ISER's competitive performance over 12 SOTA methods.
May 7, 2026cs.LG

Kurtosis-Guided Denoising Score Matching for Tabular Anomaly Detection

Denoising score matching (DSM) provides a way to learn data distributions by training a neural network to recover the score function, defined as the gradient of the log density, from noise-corrupted samples. Once trained, the score magnitude at a test point reflects how consistent that point is with the learned distribution, making it a natural anomaly signal. The key practical challenge is selecting the perturbation scale: too little noise yields unstable score estimates in sparse regions, while too much erases local structure and weakens anomaly sensitivity. This is compounded by the difficulty of hyperparameter tuning when anomalies are unknown and no validation set is available. We introduce kurtosis-based noise scaling (K-DSM), a per-feature scheme that sets noise levels from the shape of each marginal distribution, improving coverage of low-density regions and precision in high-density regions without extra model complexity. Contrary to prior claims that multi-scale or noise-conditioned training is necessary, we find that a carefully trained single-scale model is already a strong anomaly detector. On standard tabular anomaly detection benchmarks, K-DSM achieves state-of-the-art performance in the semi-supervised setting. When combined with a lightweight EMA-teacher filtering rule that removes low-density training points before each gradient step, it also achieves strong performance in the fully unsupervised (contaminated) setting, suggesting that simple, data-adaptive noise scaling enables robust anomaly detection while reducing reliance on hyperparameter tuning.
Sep 22, 2026cs.LG

Can We Predict Anomaly Detection Performance from Embedding-Space Geometry?

Anomaly detection systems are often trained using normal data alone, while model selection and evaluation typically require labeled anomalies. We study whether anomaly detection performance can be predicted without access to anomalous data. For kNN-based detectors, we derive a lower bound on the area under the ROC curve (AUC) that relates detection performance to the separation between inlier and outlier scores and to their respective variances. Under a local scaling model, we use this bound to characterize how density variation, intrinsic-dimensional heterogeneity, and cross-domain mismatch contribute to score variability. We then investigate anomaly-free model selection and show that inlier score variance alone does not reliably predict performance across different representations. To address this limitation, we introduce simple pseudo-anomaly probes that provide a reference for estimating relative score separation. Experiments on the DCASE 2022-2025 benchmarks, spanning four embedding models and 208 candidate systems, show that pseudo-anomaly-based estimators substantially improve anomaly-free model selection. In particular, diverse pseudo-anomalies enable anomaly-free model selection to outperform conventional development-set selection under domain shift. These results show that embedding-space geometry contains predictive information about anomaly detection performance while also highlighting the representation-dependent nature of inlier-only performance estimates.