stat.MLApr 16, 2026

PRIM-cipal components analysis

Authors: Tianhao LiuDaniel Andrés Díaz-PachónJ. Sunil Rao

Organizations: Division of Biostatistics, University of Miami, Miami, FL, 33136 USA · University of Minnesota, Minneapolis, MN, 55414 USA

Abstract

Supervised No Free Lunch Theorems (NFLTs) are well studied, yet unsupervised NFLTs remain underexplored. For elliptical distributions, we prove that there exist two equally optimal, scientifically meaningful bump-hunting strategies that are exact opposites, with no universal winner. Specifically, peeling kk orthogonal dimensions from Rd\mathbb{R}^d (dkd \ge k), retaining an inter-quantile region of probability 1α1-α per peeled dimension, maximizes total variance and Frobenius norm when the kk smallest principal components (called pettiest components) are selected, and minimizes them when the selected dimensions are the kk leading principal components. These optima inspire PRIM-based bump-hunting algorithms either by minimizing variance or by minimizing volume, thereby motivating an NFLT. We test our results on the Fashion-MNIST database, showing that peeling the largest principal components captures multiplicity, while peeling the smallest principal components isolates popular styles.

Explore similar work

CardsList
  1. Bandit PCA with Minimax Optimal Regret

    Jul 12, 2026Moïse Blanchard, Dmitrii Ostrovskii, Aadirupa SahaPrincipal Component AnalysisRegret