A Hierarchy of Entropy-Shapley Games for Multivariate Predictive Uncertainty
Authors: Niklas Koenen, Claudia Battistin, Jeriek Van den Abeele, Martin Jullum
Organizations: Leibniz Institute for Prevention Research and Epidemiology – BIPS University of Bremen, Germany · Simula Research Laboratory Oslo, Norway · Telenor Research & Innovation Fornebu, Norway · Norwegian Computing Center Oslo, Norway
Modern probabilistic machine learning models increasingly produce multivariate outputs with complex dependence structure, from multi-step time-series forecasts to sample path predictions. Understanding which input features drive the predictive uncertainty is important for risk-aware decisions, model diagnostics, and deciding whether the uncertainty should be mitigated or hedged against. This attribution problem requires a choice of how dependencies between output components are treated. Existing approaches reduce the output to a scalar through aggregation or projection before attribution, thereby obscuring whether features affect marginal uncertainty, dependence structure, or both, while component-wise analyses can miss dependence effects entirely. We close this gap by introducing a hierarchy of three entropy-based Shapley games that make this output-side choice explicit for any ordered multivariate outcome, ranging from per-component marginal entropy to fully joint entropy. The hierarchy isolates a cross-component attribution term that captures how each feature shifts the dependence between output components, a quantity invisible to component-wise methods. We establish a chain-rule decomposition of the joint attribution and characterize the cross-component term through conditional total correlation, providing both closed-form and sample-based estimators. Finally, we demonstrate how the framework captures differences in learned joint structure across probabilistic models from distributional regression to a zero-shot time series foundation model.
Figures & tables
Figure 1 : Entropy-Shapley hierarchy on the synthetic Gaussian DGP for an instance with high correlation. Bars connect equal quantities across panels, visualizing Prop. 1 and Prop. 2 .
Figure 2 : Entropy-Shapley hierarchy on NGBoost for two Bike Sharing instances (top: workday; bottom: weekend). Left: predicted mean and uncertainty bands. Center: Level 1 and Level 2 attributions (blue: <0 , red: >0 ). Right: Level 3 and cross-component (hatched) attributions.
Figure 3 : Cross-component share on 100 forecast origins for DeepAR and Chronos on the electricity dataset, estimated via the parametric (DeepAR only), Gaussian copula, and kNN.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Feature
Description
Temp 6am
06:00 normalized temperature
Hum 6am
06:00 normalized humidity
Wind 6am
06:00 normalized wind speed
Week Temp
7-day rolling mean temperature, lagged by 1 day
Prev Count
yesterday’s mean rentals per hour (scaled by 1/100 )
Workday
binary indicator (1 if working day)
Appendix
Table D.1 : Feature set used for the Bike Sharing experiment ( p=8 ).
Figure D.2 : Summary of the feature values (including the population means) and the predictive distribution of the two instances to be explained. The heatmaps show the correlation of the predictive distribution, i.e., of Σ(x) for the two instances.
Figure D.3 : Local SHAP vs. Entropy-Shapley attributions for the workday and weekend test instances (left, centre) and global aggregation over the full test set (right). SHAP attributes the predicted mean μ^t(x) in rentals, while Level 1 ( ϕH(t) ) and Level 2 ( ϕH(t∣<t) ) attribute the marginal and sequentially conditional predictive entropy in nats, respectively. The right panel compares the mean of ∣ϕHjoint∣ across the test set with the corresponding mean absolute SHAP values, each min-max normalized within its method. This highlights how feature effects on the mean and uncertainty are naturally decoupled: for instance, Workday strongly drives point prediction but not the joint (daily) uncertainty, whereas the opposite is true for Hum 6am and Temp 6am .
Figure D.4 : Sample-based estimation against the parametric model-based Level 2 reference vH(t∣<t)(S,x) , for DeepAR with Gaussian, Student- t , and log-normal likelihoods. Top: MAE on the value function. Bottom: MAE on the resulting Shapley values.
Originating in game theory, Shapley values are widely used for explaining a machine learning model's prediction by quantifying the contribution of each feature's value to the prediction. This requires a scalar prediction as in binary classification, whereas a multiclass probabilistic prediction is a discrete probability distribution, living on a multidimensional simplex. In such a multiclass setting the Shapley values are typically computed separately on each class in a one-vs-rest manner, ignoring the compositional nature of the output distribution. In this paper, we introduce Shapley compositions as a well-founded way to properly explain a multiclass probabilistic prediction, using the Aitchison geometry from compositional data analysis. We prove that the Shapley composition is the unique quantity satisfying linearity, symmetry and efficiency on the Aitchison simplex, extending the corresponding axiomatic properties of the standard Shapley value. We demonstrate this proper multiclass treatment in a range of scenarios.
We address the problem of explainability in machine learning models through feature attribution methods. In particular, we consider a variant of Shapley values known as Asymmetric Shapley Values (ASV), which enables the incorporation of causal knowledge into model-agnostic explanations through the use of a causal graph. We show that in certain contexts in which the computation of SHAP is #P-hard, the exact computation of ASV can be done in polynomial time. To extend this algorithmic result, we introduce a notion of equivalence classes over the topological orderings of the underlying causal graph, which is useful to reduce the time to compute ASV. In particular, we present a polynomial-time algorithm (in the number of equivalence classes) to compute it whenever the causal graph is a rooted directed tree. Finally, we develop an algorithm for approximating ASV in arbitrary causal DAGs which relies on a procedure to sample topological orderings uniformly at random. To implement this sampling mechanism we leverage known algorithms as well as simpler alternatives. Our experimental results demonstrate the practical viability of the proposed approach in realistic causal structures.
Ezequiel Companeetz, Santiago Cifuentes, Sergio Abriola
Departamento de Computación, Facultad de Ciencias Exactas y Naturales, UBA · Instituto de Ciencias de la Computación (ICC), CONICET
Machine learning pipelines commonly flatten relational data into single-table representations, discarding structural constraints. Widely used Shapley value-based feature attributions then rely on feature independence, evaluating the model on combinations that could never arise in the underlying data, producing misleading explanations. We propose RelShap, a framework that incorporates relational constraints and data provenance into Shapley value computation, restricting both background data and coalition evaluation to relationally valid configurations. The framework is estimator-agnostic and composes with Kernel SHAP, Monte Carlo, and Leverage SHAP without altering their sampling or weighting properties. Functional dependencies further induce equivalence classes over feature coalitions, which RelShap exploits to reduce runtime without changing Shapley values; we provide a combinatorial characterization of the expected speedup. Experiments across multiple datasets, models, and estimators show that RelShap produces explanations that are more faithful to the data-generating process, correctly identifying the dominant feature in controlled settings where existing methods, including Conditional SHAP and ManifoldShap, do not. Our code is available at: https://github.com/duneag2/relshap.
Seungeun Lee, Joao Fonseca, Julia Stoyanovich
New York University, New York, USA · INESC-ID, Lisbon, Portugal