MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees
Organizations: Department of Informatics, University of Oslo Oslo, Norway · Department of Electronic Systems, Aalborg University Aalborg, Denmark
Abstract
Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive information can be distributed across correlated predictors. Existing methods such as SHAP, LIME, HSIC, MI/CMI, and SAGE may therefore produce unstable rankings under multicollinearity or near-duplicate predictors. We propose the Mutual Correlation Impact Ratio Method (MCIR-M), a dependence-aware global feature-importance approach that quantifies the unique predictive information contributed by each feature beyond a selected dependence neighbourhood. MCIR-M introduces the Mutual Correlation Impact Ratio (MCIR), which conditions each feature on strongly dependent neighbours and computes a normalized ratio of conditional to block-level information. The population score lies in [0,1] and equals zero under exact conditional redundancy. We also introduce a lightweight estimation procedure that computes MCIR using a fraction of the available data and evaluates agreement with full-data explanations. Across controlled synthetic redundancy experiments and the UCI HAR benchmark, MCIR shows dependence-aware ranking behaviour, with its clearest advantage under injected near-duplicate predictors. Comparisons with independent and conditional SHAP, SAGE, HSIC, MI-based scores, and CIR-family baselines are mixed across real-data criteria. Reduced explanation samples lower computational burden in the evaluated configurations, while agreement with full-data explanations is assessed separately through ranking, head-set, and faithfulness diagnostics. Overall, MCIR-M provides a practical dependence-aware diagnostic for global explanation under strong feature dependence.
Figures & tables
| Data regime | Estimator | Scope and conditions |
|---|---|---|
| Continuous | Gaussian–copula after rank Gaussianization, | Primary estimator for continuous variables under a Gaussian-copula approximation. Rank-based invariance holds for strictly monotone transformations, with a fixed rule for ties. |
| Continuous, nonparametric | NN/KSG MI and conditional extensions | Uses local neighbour distances without a Gaussian-copula assumption. Applied in low- or moderate-dimensional conditioning blocks; accuracy may deteriorate as increases. |
| Discrete | Plug-in frequency estimator | Computes MI and CMI from empirical probability masses. Requires adequate cell counts; sparse contingency tables may introduce finite-sample bias. |
| Mixed continuous–discrete | Not evaluated empirically | The population definition remains valid, but this study does not implement or report a general mixed-type MI/CMI estimator. Application to such blocks requires a fully specified mixed-type estimator and its corresponding regularity assumptions. |
| High-dimensional conditioning | Dimension diagnostic | No uniform estimator guarantee is assumed. Results are evaluated as a function of , and large conditioning blocks are reported as an estimator-limitation regime. |
| Category | Specification |
|---|---|
| Software environment | Python 3.10; NumPy 1.26; SciPy 1.11; scikit-learn 1.3; pandas 2.1; PyTorch; Matplotlib; and Seaborn. Exact package versions are pinned in the repository environment file. |
| Execution environment | Linux-based Google Colab Pro sessions. Deep-model training uses the GPU assigned to the recorded session; attribution computations run on CPU. The exact CPU, GPU, memory, and runtime configuration is recorded with each experiment log. |
| Randomness control | Seed 7 for train/test partitions and seed 42 for subsampling, bootstrap, estimator, NumPy, scikit-learn, and PyTorch randomness. Deterministic operations are used where supported. |
| Gaussian–copula estimator | Rank transformation followed by empirical-CDF Gaussianization , deterministic tie handling, and ridge-regularized covariance estimation. Multivariate MI and CMI are computed through covariance determinant or Schur complement formulas. Ridge values are reported per experiment in Table 6 . |
| NN estimator | KSG-type MI and conditional extensions for continuous, low- or moderate-dimensional blocks; with Euclidean distance. Results are reported as a function of because estimator behaviour may degrade with conditioning dimension. |
| Discrete estimator | Plug-in MI and CMI estimation from empirical probability masses. Finite-sample corrections and minimum cell-count requirements are reported with the corresponding experiment. |
| Kendall | Jaccard@10 | Driver Recall@10 | ||
|---|---|---|---|---|
| 500 | 1 | 0.391 | 0.514 | 0.538 |
| 1000 | 1 | 0.445 | 0.583 | 0.638 |
| 2000 | 1 | 0.514 | 0.656 | 0.650 |
| MCIR | MI | PCIR | ||||
|---|---|---|---|---|---|---|
| 0 | 0.0000 | 0.0105 | 0.0000 | 0.0103 | 0.0000 | 0.0015 |
| 5 | 0.4667 | 0.4473 | 0.0000 | 0.4495 | 1.0000 | 0.2178 |
| 10 | 0.4667 | 0.5261 | 1.3333 | 0.5288 | 1.0000 | 0.2499 |
| 20 | 0.4667 | 0.5785 | 1.2667 | 0.5789 | 1.0000 | 0.2747 |
| 40 | 0.5333 | 0.6095 | 1.3333 | 0.6077 | 1.0000 | 0.2910 |
| Setting | Training rows | Accuracy | Macro-F1 | Jaccard@30 | F1 ratio | |
|---|---|---|---|---|---|---|
| Full training | 7,352 | 0.930 | 0.928 | 0.000 | 1.000 | 1.000 |
| Reduced training | 3,676 | 0.917 | 0.914 | 7.174 | 0.622 | 0.985 |
| Metric | Mean | 2.5% | 97.5% |
|---|---|---|---|
| Kendall | 0.58 | 0.51 | 0.64 |
| Spearman | 0.73 | 0.67 | 0.78 |
| Jaccard@10 | 0.40 | 0.30 | 0.50 |
| Method | Core Recall@4 | Original mass | Proxy mass | -drop AUC | Runtime (s) |
|---|---|---|---|---|---|
| MCIR | |||||
| LOCO | |||||
| Region-weighted LOCO | |||||
| Shapley-effects plug-in |
| Method | HouseEnergy-Sim | UCI HAR | ||
|---|---|---|---|---|
| Time (s) | Setting | Time (s) | Setting | |
| MCIR (copula) | 0.4487 | rank Gaussianization | 451.412 | rank Gaussianization |
| MCIR ( NN) | 4.3769 | 139.738 | ||
| HSIC (RBF) | 7.0033 | median bandwidth | 207.823 | median bandwidth |
| Method | Deletion AUC | Accuracy after top-128 deletion |
|---|---|---|
| MCIR | 0.887 | |
| PCIR | 0.912 | |
| HSIC | 0.938 | |
| MI | 0.951 |
| Area | Test | J@8 | Group-J@8 | |||
|---|---|---|---|---|---|---|
| NO1 | 0.9924 | 0.666 | 0.508 | 0.639 | 0.623 | 64.57 |
| NO2 | 0.9919 | 0.709 | 0.558 | 0.799 | 0.814 | 37.04 |
| NO3 | 0.9924 | 0.659 | 0.507 | 0.671 | 0.757 | 22.30 |
| NO4 | 0.9937 | 0.640 | 0.478 | 0.560 | 0.617 | 14.40 |
| NO5 | 0.9903 | 0.618 | 0.471 | 0.600 | 0.713 | 14.13 |
| MCIR versus TreeSHAP | Cumulative -Drop AUC | |||||||
|---|---|---|---|---|---|---|---|---|
| Area | J@8 | Group-J@8 | MCIR | TreeSHAP | MI | HSIC | ||
| NO1 | 0.377 | 0.254 | 0.333 | 0.500 | 1.695 | 1.798 | 1.787 | 1.788 |
| NO2 | 0.552 | 0.397 | 0.600 | 0.667 | 1.276 | 1.388 | 1.401 | 1.407 |
| NO3 | 0.257 | 0.169 | 0.333 | 0.429 | 1.881 | 2.017 | 2.030 | 2.025 |
| NO4 | 0.273 | 0.228 | 0.455 | 0.429 | 2.747 | 3.188 | 3.211 | 3.127 |
| NO5 | 0.247 | 0.169 | 0.333 | 0.429 | 1.418 | 1.591 | 1.582 | 1.581 |
| Rank | NO1 | NO2 | NO3 | NO4 | NO5 |
|---|---|---|---|---|---|
| 1 | sin_hour | sin_hour | relative_humidity_2m | load_lag_1h | cos_hour |
| 2 | cos_hour | cos_hour | load_lag_1h | wind_speed_10m | load_lag_1h |
| 3 | load_lag_1h | load_lag_1h | cos_hour | cos_hour | relative_humidity_2m |
| 4 | load_roll_mean_24h | relative_humidity_2m | load_roll_mean_24h | load_roll_mean_24h | load_roll_std_6h |
| 5 | load_lag_24h | load_lag_24h | precipitation | load_roll_std_6h | load_roll_mean_24h |
| 6 | wind_speed_10m | load_roll_mean_24h | load_roll_std_6h | sin_dow | load_lag_24h |
| Method | Minimum (s) | Maximum (s) |
|---|---|---|
| Marginal MI | 0.0559 | 0.0675 |
| MCIR | 0.0810 | 0.0983 |
| HSIC | 0.1108 | 0.1801 |
| TreeSHAP | 0.2280 | 0.2817 |
| Property | Formal result | Evidence and scope |
|---|---|---|
| Boundedness and endpoints | Equations 26 and 27 ; Eq. equation 28 | Population guarantee. , with the endpoints characterized by the corresponding information quantities. Empirical clipping in Equation 37 enforces this range and is therefore not independent empirical validation. |
| Exact conditional-redundancy collapse | Equation 29 | Theory with sensitivity evidence. Measurability of with respect to implies . The noisy-duplicate results in Figures 8 and 8 and Tables 26 , 16 and 17 assess finite-sample sensitivity, not exact deterministic collapse. |
| Weak-dependence reduction | Proposition 2 | Theoretical result. MCIR approaches a normalized marginal-information score under vanishing conditioning effects and denominator separation. No controlled zero-dependence experiment or ordering equivalence with PCIR is claimed. |
| Population reparameterization invariance | Proposition 1 | Population guarantee. MCIR is invariant under bimeasurable bijections of , , and . The separate rank–Gaussianized estimator result is given in Proposition 7 ; no transformation experiment is reported. |
| Consistency and rank stability | Equations 38 and 1 | Conditional guarantee; indirect evidence. The result requires estimator-error, denominator-separation, and population-margin conditions. Subsampling results in Tables 11 and 10 use a full-data empirical reference rather than the unavailable population ranking. The conditioning-dimension stress test in Figures 13 and 31 additionally shows that empirical score error and ranking agreement can diverge as increases. |
| Estimator switching | Equation 51 | Empirical heuristic. Fixed-estimator sensitivity is reported in Figures 11 and 13 , with repeated-subsampling results in Appendix Table 37 . Performance depends on and ; no oracle inequality is claimed. |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| OrigMass | FamilyShare | |||||
|---|---|---|---|---|---|---|
| MCIR | TreeSHAP | PFI | MCIR | TreeSHAP | PFI | |
| 1 | 0.908 | 0.784 | 0.834 | |||
| 2 | 0.904 | 0.687 | 0.776 | |||
| 4 | 0.887 | 0.503 | 0.651 | |||
| 8 | 0.852 | 0.150 | 0.387 | |||
| MCIR | TreeSHAP | PFI | |
|---|---|---|---|
| 0 | 0.0000 | 0.0000 | 0.0000 |
| 1 | 0.0000 | 0.6667 | 0.2500 |
| 2 | 0.0000 | 0.6667 | 0.6667 |
| 4 | 0.0000 | 0.6667 | 0.6667 |
| 8 | 0.0000 | 0.6667 | 0.6667 |
| 16 | 0.0000 | 0.6667 | 1.3333 |
| MCIR | TreeSHAP | PFI | |
|---|---|---|---|
| 1 | |||
| 2 | |||
| 4 | |||
| 8 |
| Method | Time (s) | Relative time |
|---|---|---|
| MCIR (lightweight) | 0.0341 | |
| TreeSHAP | 0.9114 | |
| PFI (5 repeats) | 0.2496 |
| Method A | Method B | Jaccard@4 | ||
|---|---|---|---|---|
| MCIR | LOCO | 0.436 | 0.568 | 0.560 |
| MCIR | Region-weighted LOCO | 0.449 | 0.594 | 0.667 |
| MCIR | Shapley-effects plug-in | 0.303 | ||
| LOCO | Region-weighted LOCO | 0.764 | 0.865 | 0.720 |
| LOCO | Shapley-effects plug-in | 0.205 | ||
| Region-weighted LOCO | Shapley-effects plug-in | 0.195 |
| Estimator | Score MAE | J@4 | Negative components | Runtime (s) | |||
|---|---|---|---|---|---|---|---|
| 500 | 1 | copula | 0.163 | 0.516 | 0.295 | 0.000 | 0.08 |
| 500 | 1 | knn | 0.204 | 0.136 | 0.268 | 0.365 | 0.26 |
| 500 | 2 | copula | 0.101 | 0.441 | 0.533 | 0.000 | 0.10 |
| 500 | 2 | knn | 0.149 | 0.134 | 0.281 | 0.380 | 0.33 |
| 500 | 3 | copula | 0.067 | 0.394 | 0.433 | 0.000 | 0.12 |
| 500 | 3 | knn | 0.115 | 0.118 | 0.295 | 0.333 | 0.37 |
| Method | 20% | 40% | 60% |
|---|---|---|---|
| Exact Jaccard@ | |||
| MCIR | 0.585 | 0.624 | 0.627 |
| PCIR | 0.673 | 0.710 | 0.767 |
| SHAP | 0.644 | 0.698 | 0.747 |
| HSIC | 0.674 | 0.711 | 0.787 |
| Group-Jaccard@ | |||
| MCIR | PCIR | SHAP | HSIC | |
|---|---|---|---|---|
| 0.60 | 0.6405 | 0.5036 | 0.9200 | 0.4284 |
| 0.70 | 0.6261 | 0.4648 | 0.9667 | 0.4334 |
| 0.75 | 0.5792 | 0.4204 | 0.9741 | 0.4582 |
| 0.80 | 0.5946 | 0.4720 | 0.9788 | 0.5201 |
| 0.85 | 0.5821 | 0.4321 | 0.9662 | 0.4994 |
| 0.90 | 0.5732 | 0.4372 | 0.9490 | 0.5006 |
| GT Recall@ | Runtime | |||
|---|---|---|---|---|
| Copula | NN | Copula | NN | |
| 1 | 0.4615 | 0.6154 | 0.0794 | 56.0264 |
| 2 | 0.6154 | 0.7692 | 0.0474 | 13.6651 |
| 3 | 0.5385 | 0.7692 | 0.0408 | 15.9895 |
| 5 | 0.6923 | 0.7692 | 0.0441 | 22.2782 |
| 200 | 4 (13.3)/26 (86.7) | 0 (0.0)/30 (100.0) | 0 (0.0)/30 (100.0) | 0 (0.0)/30 (100.0) |
|---|---|---|---|---|
| 500 | 26 (86.7)/4 (13.3) | 0 (0.0)/30 (100.0) | 0 (0.0)/30 (100.0) | 0 (0.0)/30 (100.0) |
| 1000 | 30 (100.0)/0 (0.0) | 0 (0.0)/30 (100.0) | 0 (0.0)/30 (100.0) | 0 (0.0)/30 (100.0) |
| 2000 | 30 (100.0)/0 (0.0) | 4 (13.3)/26 (86.7) | 0 (0.0)/30 (100.0) | 0 (0.0)/30 (100.0) |
| 3000 | 30 (100.0)/0 (0.0) | 20 (66.7)/10 (33.3) | 0 (0.0)/30 (100.0) | 0 (0.0)/30 (100.0) |
| Features | Screening | Calibration | Final MCIR | Total |
|---|---|---|---|---|
| 20 | 0.001 | 0.179 | 0.004 | 0.184 |
| 50 | 0.002 | 0.418 | 0.012 | 0.432 |
| 100 | 0.002 | 0.869 | 0.029 | 0.900 |
| 200 | 0.003 | 1.939 | 0.069 | 2.011 |
| 400 | 0.005 | 4.006 | 0.116 | 4.127 |
| 561 | 0.007 | 5.566 | 0.188 | 5.762 |
| Observation | Interpretation |
|---|---|
| Pairwise screening omits relevant structure | The association between , , and is mediated by their product, which is weakly visible to pairwise linear correlation. |
| Low all-relevant recall for correlation- MCIR | The limitation arises from constructing , not solely from the bounded MCIR normalization. |
| High recall for MI, HSIC, and RF-PFI | These methods recover marginal or predictive relevance but do not estimate the same unique conditional contribution as MCIR. |
| Reference- is not deployable | It is supplied manually only to show how the ranking changes when interaction-informed neighbourhoods are imposed. |