SHAP explanations are widely used in high-stakes settings to justify decisions, yet they can differ substantially across repeated runs, even when the model, the input instance, and the prediction are held fixed. Prior work has documented disagreement between explanation methods; we show that substantial disagreement arises even within SHAP across reruns of the same estimator on the same trained model and instance. We call this phenomenon explanation multiplicity and develop an evaluation methodology for characterizing it under deployment-realistic computational budgets, combining a dual-seed protocol that compares model-induced and explainer-induced variability, a hierarchy of magnitude-based, rank-based, and set-based metrics, and randomized Dirichlet and Mallows null models that provide reference scales for observed disagreement. Across multiple datasets, models, and sampling strategies, we find that explanation multiplicity is pervasive and persists even for high-confidence predictions. The relative contribution of each source depends on the data regime and model: model-induced disagreement is generally greater on smaller datasets, while explainer-induced disagreement is greater on larger datasets. Commonly used L2 distance can understate this instability, while rank-based metrics reveal substantial changes in top-ranked features, including the leading feature. Improved sampling methods such as CTE do not eliminate rank-level multiplicity, and K-Means reduces run-to-run variation while its compressed-background explanations can diverge from the empirical-distribution reference. Practitioners should treat single-run SHAP outputs as realizations of a distribution rather than as authoritative artifacts.
Figures & tables
Figure 1 : Feature-based explanations for a high-confidence adverse credit prediction vary substantially across reruns of the same explanation pipeline. Top-ranked features differ with no overlap at top- 3 . A and B denote two runs of the same estimator. Example based on lending ( German Credit ), with SHAP for an FT-Transformer classifier.
Concept
What varies
What is fixed
Level
Reference
Model multiplicity
Trained model
Dataset, task
Model class
[ 13 , 22 ]
Inter-method disagreement
Method
Model, instance
Explainer
[ 8 ]
Explanation multiplicity
Pipeline run
Query Q
Query
This work
Table 1 : Comparison of explanation multiplicity to related concepts.
Figure 2 : Explanation pipeline. Randomness can enter at (a) model construction (training and selection) and (b) background sampling (how absent features are instantiated). Source (b) is specific to the explainer; different estimators handle it differently.
Figure 3 : Model vs. explainer multiplicity (dissection). Kernel SHAP with Random background sampling. Pair-distance distributions by source ( ℓ2 , Top-3 Jaccard, 1-RBO); Jaccard uses four-value frequency bars. Green: explainer seed varied with model fixed; orange: model seed varied with explainer fixed. Shaded bands denote randomized baseline ranges. Model-induced disagreement is generally greater on smaller datasets, with model-dependent exceptions; explainer-induced disagreement is greater on larger datasets. Full results appear in Figure 15 .
Figure 4 : Metric divergence for Kernel SHAP on GMSC . Percentage of run-pair–instance comparisons exceeding the randomized baseline, by method and model (Eq. ( 7 )). Each cell uses 10N comparisons. Low ℓ2 exceedance coexists with Top-3 membership churn and RBO disagreement.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Features
Total
Train
Test
Positive Rate
German Credit
16
988
790
198
29.9%
Diabetes
8
392
313
79
33.2%
ACS Income
8
45,960
36,768
9,192
43.7%
GMSC
10
45,000
36,000
9,000
6.7%
Appendix
Table 2 : Dataset statistics and the evaluated train/test fold from a five-fold stratified partition.
Model
Fixed Parameters
Grid Search Parameters (Values)
Decision Tree
-
max_depth ∈{3,5,None}
min_samples_leaf ∈{1,5,10}
Random Forest
n_estimators =200
max_depth ∈{None,7,15}
min_samples_leaf ∈{1,5}
XGBoost
n_estimators =200
max_depth ∈{3,5}
tree_method =‘hist’
learning_rate ∈{0.05,0.1}
Appendix
Table 3 : Hyperparameter Search Space. Fixed parameters are set to ensure fair comparison or convergence, while grid parameters are optimized using a validation set.
Dataset
Method
N
Any-seed
Pairwise
German Credit
Random
198
8.1%
4.4%
CTE
198
10.1%
4.9%
Diabetes
Random
79
11.4%
5.9%
CTE
79
15.2%
7.3%
GMSC
Random
9,000
67.7%
42.9%
CTE
9,000
62.0%
32.0%
Appendix
Table 4 : Top-1 churn for FT-Transformer with K=100 . Any-seed churn uses N queries, whereas pairwise churn uses all 10N run-pair–instance comparisons. Both are event frequencies without a randomized threshold.
Figure 5 : Rank positions of consensus features. For each query, the five-run mean absolute SHAP values select the first three consensus features. Each row records the rank of one such feature in the five contributing runs, pooling 5N observations; row frequencies sum to 100% before rounding. Columns include every possible rank. Entries are rounded percentages; <1 marks positive values that round to zero. Panels use FT-Transformer, K=100 , and a shared color scale. This within-sample rank summary is distinct from the Top-1 event frequencies in Table 4 .
Figure 6 : Sensitivity to Top- k depth. Each cell is the percentage of 10N run-pair–instance Jaccard distances strictly above the dataset- and k -specific Mallows baseline, with K=100 . No pair distances are averaged before thresholding. All four panels share the same color scale; the k=3 panel matches the main analysis. In these settings, the threshold is below the smallest positive Jaccard distance, so exceedance equals the frequency of Top- k set changes. The panels need not increase monotonically with k .
Figure 7 : Background size and Top-3 exceedance: FT-Transformer. Each point reports the percentage of 10N run-pair–instance distances strictly exceeding the dataset’s fixed baseline. The horizontal axis is logarithmic and marks the actual evaluated sizes, including the full-training endpoints K=790 (German Credit) and K=313 (Diabetes). The K=100 points reuse the main inputs. Lines connect observed sizes without smoothing; they do not imply monotonicity or isolate background sampling from coalition variability.
Figure 8 : Background size and Top-3 exceedance: MLP. The protocol, shared 0 – 100% vertical scale, and strict 10N aggregation match Figure 7 . Baselines are fixed across K within each dataset. A larger background does not consistently reduce disagreement at every intermediate size.
Model
Estimator
Any-seed Top-1
Pairwise Top-1
Identical Top-3
FTT
Conditional
10.1%
5.2%
85.7%
Marginal
13.9%
7.1%
73.8%
MLP
Conditional
12.7%
5.8%
82.2%
Marginal
16.5%
9.0%
80.1%
Appendix
Table 5 : Conditional and marginal Shapley estimation on Diabetes. Top-1 any-seed churn is the percentage of N=79 queries whose leading feature differs across at least two runs. Pairwise Top-1 churn and identical Top-3 sets use all 10N=790 run-pair–instance comparisons. These are raw event frequencies, without a randomized threshold.
Figure 9 : Pairwise disagreement on Arrhythmia. Left: the two runs select different leading features. Right: their Top-3 feature sets differ. Both panels use all ten run pairs for each of the 84 test queries ( 840 comparisons per cell) and the same 0 – 100% color scale. These are raw event frequencies, not baseline exceedance rates. Explanations have ten nonzero features out of 278; the low-dimensional Mallows calibration is not applied here.
Model
Method
Any-seed
Pairwise
Support
Selection
Rank
MLP
Random
54.8%
37.9%
.339
77.7%
86.3%
MLP
CTE
52.4%
34.6%
.358
78.4%
88.0%
FTT
Random
54.8%
36.0%
.335
70.9%
87.2%
FTT
CTE
52.4%
34.0%
.362
64.2%
89.4%
Appendix
Table 6 : Results on the Arrhythmia dataset. Top-1 any-seed uses N instances; pairwise uses 10N run-pair–instance comparisons. Support agreement is mean Jaccard similarity of nonzero supports over those comparisons. Selection/rank churn use only Top-3-disagreeing observations as their denominator and can co-occur.
Feature
Median
Attribution range
Rank range
Top-3 frequency
savings_status
0.055
[0.028, 0.065]
1–4
3/5
checking_status
0.050
[0.025, 0.065]
1–5
3/5
credit_amount
0.042
[0.018, 0.088]
1–6
3/5
credit_history
0.038
[0.005, 0.047]
3–11
3/5
duration
0.020
[0.015, 0.070]
2–5
2/5
existing_credits
0.004
[0.004, 0.084]
2–11
1/5
Appendix
Table 7 : Five-run explanation report for instance #187. All listed features are in the unstable margin; none meets the illustrative four-of-five Top-3 inclusion rule.
Dataset
Model
K=25
K=50
Max., K≥100
German Credit
MLP
0.00
0.00
0.00
German Credit
FTT
0.00–1.06
0.00–0.20
0.20
Diabetes
MLP
0.00
0.00
0.00
Diabetes
FTT
0.63–24.30
0.00–1.65
0.00
ACS Income
MLP
0.00–25.44
0.00–0.25
0.00
ACS Income
FTT
0.00–31.73
0.00–0.61
0.00
Appendix
Table 8 : Background size and ℓ2 baseline exceedance (%). At K=25 and K=50 , entries give the minimum–maximum across Random, K-Means, PredStrat, and CTE; a single value denotes an equal rate across methods. The last column is the maximum across all four methods and all evaluated K≥100 for that dataset/model. Each underlying rate uses strict thresholding over 10N run-pair–instance comparisons. Ranges are descriptive, not confidence intervals; displayed zeros represent exactly zero exceedances.
Figure 10 : Background size and 1− RBO exceedance: FT-Transformer. Each observed size contributes 10N run-pair–instance distances, with strict thresholding before aggregation. The dataset-specific threshold is fixed across sizes. RBO uses the full feature ranking, p=1−1/d , and the extrapolated final term. Actual full-training endpoints and canonical K=100 reuse follow Figure 7 .
Figure 11 : Background size and 1− RBO exceedance: MLP. Each observed size contributes 10N run-pair–instance distances, with strict thresholding before aggregation. The dataset-specific threshold is fixed across sizes. RBO uses the full feature ranking, p=1−1/d , and the extrapolated final term. Actual full-training endpoints and canonical K=100 reuse follow Figure 7 .
Figure 12 : Pairwise baseline exceedance for Kernel SHAP. For each dataset, model, and method, we report the percentage of 10N run-pair–instance distances strictly above the baseline. Sub-panels show results for (a) ℓ2 distance, (b) Top-3 Jaccard distance, and (c) 1− RBO. In the displayed matrices, ℓ2 exceedance rounds to 0% while rank-based metrics surface substantial multiplicity, particularly for GMSC and neural architectures.
Figure 13 : Pairwise baseline exceedance for Permutation SHAP. Same 10N run-pair–instance exceedance protocol as Figure 12 , using Permutation SHAP. Low ℓ2 exceedance coexists with rank-level disagreement. Internal coalition and background randomness are not separately identified by this comparison.
Figure 14 : Overall sensitivity across datasets. Run-pair–instance distance distributions for Kernel SHAP with Random background sampling on German Credit , Diabetes , ACS Income , and Give Me Some Credit , measured by ℓ2 , Top- k Jaccard, and 1-RBO; raw Jaccard uses four-value frequency bars. Model-seed and explainer-seed collections are pooled, each contributing 10N distances. Shaded bands indicate randomized baseline ranges. FT-Transformer and MLP exhibit Jaccard multiplicity approaching the baseline, whereas ℓ2 remains concentrated near zero across models. On Give Me Some Credit and German Credit , the four Jaccard values carry mass both at zero and at positive membership-change distances.
Figure 15 : Model vs. explainer multiplicity (dissection). Run-pair–instance distance distributions ( 10N per source) for Kernel SHAP with Random background sampling across four datasets (rows) and three metrics (columns: ℓ2 , Top-3 Jaccard, and 1-RBO). Jaccard panels use four-value frequency bars.
Figure 16 : Explainer-seed feature-wise multiplicity for FT-Transformer Model. For Kernel SHAP, the six features with the largest mean absolute pairwise attribution changes are shown, pooling within-method comparisons across Random, K-Means, PredStrat, and CTE. Bar length averages these changes over queries and run pairs; labels give mean absolute SHAP values. Features are ordered by sensitivity, not importance.
Figure 17 : Kernel SHAP vs. Permutation SHAP. Distributions of 10N run-pair–instance distances across five explainer seeds with Random background sampling, with raw Jaccard shown as frequency bars, comparing Kernel SHAP (blue) and Permutation SHAP (red) across four datasets, five model classes, and three metrics ( ℓ2 , Top-3 Jaccard, 1− RBO). Shaded bands denote randomized baselines. The two estimators show similar qualitative patterns; this does not establish invariance to all implementation choices or isolate background from coalition randomness.
Figure 18 : Comparing sampling methods. Run-pair–instance distance distributions ( 10N per configuration), using frequency bars for raw Jaccard, grouped by background sampling strategy (Random, K-Means, Pred-Strat, CTE), shown for Kernel SHAP (blue) and Permutation SHAP (red) across four datasets and three metrics for FT-Transformer. Shaded bands denote randomized baselines. CTE exhibits intra-method sensitivity comparable to Random sampling, while K-Means produces tightly concentrated distributions whose stability does not establish reference agreement (see Figures 19 – 21 ).
Figure 19 : Mean distance to the empirical-distribution reference. Each cell averages 5N run–instance distances to the full-training-background, same-budget reference. With complete equal-sized runs, averaging first within each instance gives the same scalar mean. Panels show (a) ℓ2 , (b) Top-3 Jaccard, and (c) 1− RBO. The reported comparisons cover Diabetes and GMSC. These are reference divergences, not universal errors against exact values. Neural comparisons have the limitation in Appendix D.1 .
Figure 20 : Baseline exceedance in reference comparisons. Each cell reports the percentage of 5N run–instance distances to the reference strictly above the threshold; no distance averaging precedes thresholding. Panels show (a) ℓ2 , (b) Top-3 Jaccard, and (c) 1− RBO, using the same-budget empirical-distribution reference. The reported comparisons cover Diabetes and GMSC; neural comparisons have the limitation in Appendix D.1 .
Figure 21 : Distribution of reference distances for FT-Transformer. For Diabetes and GMSC, each distribution pools 5N run–instance distances to the same-budget empirical-distribution reference, without averaging within instances. The ℓ2 /RBO panels use shaded reference ranges; Jaccard shows four-value frequencies. Neural reference comparisons have the limitation in Appendix D.1 .
Figure 22 : Confidence vs. stability. Raw run-pair–instance distances for Kernel SHAP with Random background sampling within certain and uncertain prediction groups, with the model fixed and explainer seed varied; Top-3 Jaccard uses four-value frequency bars. A group of Ng instances contributes all 10Ng distances. The available prediction probabilities cover 19 of 20 dataset–model configurations, excluding Diabetes/FTT; the Diabetes/MLP certain group is empty. Shaded bands denote randomized baselines. High-confidence predictions can still exhibit substantial instability.
Department of Computer Science, Oslo Metropolitan University, Oslo, Norway · Department of Holistic Systems, SimulaMet, Oslo, Norway · Department of Plastic and Reconstructive Surgery, Oslo University Hospital, Oslo, Norway