cs.AIMay 4, 2026

First-Order Efficiency for Probabilistic Value Estimation via A Statistical Viewpoint

Authors: Ziqi LiuKiljae LeeYuan ZhangWeijing Tang

Organizations: Department of Statistics & Data Science, Carnegie Mellon University. · Department of Statistics, The Ohio State University.

Abstract

Probabilistic values, including Shapley values and semivalues, provide a model-agnostic framework to attribute the behavior of a black-box model to data points or features, with a wide range of applications including explainable artificial intelligence and data valuation. However, their exact computation requires utility evaluations over exponentially many coalitions, making Monte Carlo approximation essential in modern machine learning applications. Existing estimators are often developed through different representation strategies, including weighted averages, self-normalized weighting, regression adjustment, and weighted least squares. Our key observation is that these seemingly distinct constructions share a common first-order expansion, in which the leading term is determined by the sampling law and a working surrogate function. This first-order representation yields an explicit expression for the leading mean squared error (MSE), which characterizes how the sampling law and the surrogate jointly determine statistical efficiency. Guided by this criterion, we propose an Efficiency-Aware Surrogate-adjusted Estimator (EASE) that directly chooses the sampling law and surrogate to minimize the first-order MSE. We demonstrate that EASE consistently outperforms existing estimators for various probabilistic values.

Explore similar work

CardsList
  1. Beyond Shapley: Efficient Computation of Asymmetric Shapley Values

    Jun 23, 2026Ezequiel Companeetz, Santiago Cifuentes, Sergio AbriolaShapley ValueInterpretable Models

  2. Proxy-Based Approximation of Shapley and Banzhaf Interactions

    May 21, 2026Santo M. A. R. Thies, Hubert Baniecki, R. Teal Witter +3Shapley ValueApproximation