cs.LGJan 19, 2026

Explanation Multiplicity in SHAP: Characterization and Assessment

Authors: Hyunseung Hwang, Seungeun Lee, Lucas Rosenblatt, Steven Euijong Whang, Julia Stoyanovich

Organizations: KAIST, Daejeon, Republic of Korea · New York University, New York, NY, USA

Abstract

SHAP explanations are widely used in high-stakes settings to justify decisions, yet they can differ substantially across repeated runs, even when the model, the input instance, and the prediction are held fixed. Prior work has documented disagreement between explanation methods; we show that substantial disagreement arises even within SHAP across reruns of the same estimator on the same trained model and instance. We call this phenomenon explanation multiplicity and develop an evaluation methodology for characterizing it under deployment-realistic computational budgets, combining a dual-seed protocol that compares model-induced and explainer-induced variability, a hierarchy of magnitude-based, rank-based, and set-based metrics, and randomized Dirichlet and Mallows null models that provide reference scales for observed disagreement. Across multiple datasets, models, and sampling strategies, we find that explanation multiplicity is pervasive and persists even for high-confidence predictions. The relative contribution of each source depends on the data regime and model: model-induced disagreement is generally greater on smaller datasets, while explainer-induced disagreement is greater on larger datasets. Commonly used L2 distance can understate this instability, while rank-based metrics reveal substantial changes in top-ranked features, including the leading feature. Improved sampling methods such as CTE do not eliminate rank-level multiplicity, and K-Means reduces run-to-run variation while its compressed-background explanations can diverge from the empirical-distribution reference. Practitioners should treat single-run SHAP outputs as realizations of a distribution rather than as authoritative artifacts.

Figures & tables

Appendix figures & tables25 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Evaluating Explanation Methods by the Predictors They Induce

    Sep 17, 2026Jacob Selbæk, Hugo L. HammerFeature AttributionSHAP Feature Attribution

  2. MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees

    Oct 1, 2026Poushali Sengupta, Sabita Maharjan, Frank Eliassen +2Feature AttributionExplainability Evaluation

  3. RelShap: Relationally Consistent Shapley Explanations

    Aug 11, 2026Seungeun Lee, Joao Fonseca, Julia StoyanovichFeature AttributionShapley Value Attribution