stat.MLJul 10, 2026

Influence Diagnostics in High-dimensional M-estimation: Precise Asymptotics

Authors: Hugo Cui

Organizations: Université Paris-Saclay, CNRS, Laboratoire de mathématiques d’Orsay, 91405, Orsay, France

Abstract

The impact of a given training point on a statistical model is classically measured through its leave-one-out influence, which quantifies the effect of its removal from the training set on the model accuracy. While the statistics of leave-one-out influences are well understood in the low-dimensional, large sample limit n,d=O(1)n\to \infty, d=O(1), they become more intricate in high dimensions, as the influence of a given sample develops non-trivial dependencies on all other training samples. For convex M-estimation under Gaussian design, in the high-dimensional limit ndn\asymp d, we show that the distribution of the influences across the training set converges to a limiting measure which we sharply characterize. Building on these results, we provide evidence that influential samples tend to lie close to the decision boundary, thereby making contact with a standard data selection heuristic in active learning.

Explore similar work

CardsList
  1. Extending Kernel Trick to Influence Functions

    May 11, 2026Zhenhuan Sun, Shahrokh ValaeeInfluence

  2. Finding Most Influential Sets

    Jun 4, 2026Lucas D. Konrad, Nikolas KuschnigEstimatorsCross Validation