Explainability of Complex AI Models with Correlation Impact Ratio
Authors: Poushali Sengupta, Rabindra Khadka, Sabita Maharjan, Frank Eliassen, Yan Zhang, Shashi Raj Pandey, Pedro G. Lind, Anis Yazidi
Organizations: Institute of Informatics, University of Oslo, Oslo, Norway · Department of Computer Science, Oslo Metropolitan University, Oslo, Norway · Aalborg University, Denmark
Complex AI systems make better predictions but often lack transparency, limiting trustworthiness, interpretability, and safe deployment. Common post hoc AI explainers, such as LIME, SHAP, HSIC, and SAGE, are model agnostic but are too restricted in one significant regard: they tend to misrank correlated features and require costly perturbations, which do not scale to high dimensional data. We introduce ExCIR (Explainability through Correlation Impact Ratio), a theoretically grounded, simple, and reliable metric for explaining the contribution of input features to model outputs, which remains stable and consistent under noise and sampling variations. We demonstrate that ExCIR captures dependencies arising from correlated features through a lightweight single pass formulation. Experimental evaluations on diverse datasets, including EEG, synthetic vehicular data, Digits, and Cats-Dogs, validate the effectiveness and stability of ExCIR across domains, achieving more interpretable feature explanations than existing methods while remaining computationally efficient. To this end, we further extend ExCIR with an information theoretic foundation that unifies the correlation ratio with Canonical Correlation Analysis under mutual information bounds, enabling multi output and class conditioned explainability at scale.
Figures & tables
Aspect
Status quo (SOTA)
ExCIR (ours)
Computation
Shapley family: exact O(2k) (number of features k ). KernelSHAP: O(mk) model calls ( m perturbation samples). TMC-SHAP: O(mk) (Monte Carlo paths). TreeSHAP: O(TLk) (trees T , max depth L ). GradientSHAP/IG/DeepSHAP: O(mk) backprop passes. HSIC/MI estimators: typically O(n2k) (pairwise kernels).
Local/perturbation-driven; global order unstable under correlation and noise.
Performance-aligned global ranking; higher top- k sufficiency with compact subsets; correlation-aware.
Deployment
Perturbation-heavy pipelines; repeated model evaluations; full data required.
Lightweight-transfer; single-pass, low-memory; preserves ranking under subsampling (20-40% data).
Calibration
Unbounded scores; difficult cross-run comparison; sensitive to retraining.
Bounded CIR ∈[0,1] with MI-linked upper bound; comparable across datasets or models; stable under sampling & feature noise.
TABLE II : ExCIR novelties vs. SOTA in one view.
Fig. 1 : CIR geometry. We center fi and y at the mid-mean mi=21(f^i+y^) . The alignment (numerator) uses symmetric offsets ∣f^i−mi∣ and ∣y^−mi∣ ; the scatter (denominator) aggregates sample deviations around the same pivot mi .
Fig. 2 : Linear vs. nonlinear dependence scores.
Fig. 3 : Inception module used in CAU-EEG backbone.
TABLE V : Top-8 ranked features per method. Left→right = higher→lower importance.
(A) Accuracy–cost
Method / Model
Kind
Time (s)
Acc
Drop
ρs
Top-10
GBM (baseline predictor)
fit
0.701
0.000
ExCIR–LW (20%)
explain
0.005
0.701
0.000
0.94
0.98
ExCIR–LW (30%)
explain
0.007
0.701
0.000
0.95
1.00
ExCIR–LW (50%)
explain
0.008
0.701
0.000
0.96
1.00
LIME + TinyGBM (20 × 2)
fit
1.70
0.698
0.003
0.54
0.50
TABLE VI : Head-to-head accuracy–cost and comparator budget sensitivity (Vehicular).
Fig. 5 : ExCIR uncertainty and agreement under bootstrapping (vehicular). (Left) 95% CI of ExCIR scores (top features; B=100 ) (Right) Top-set overlap across bootstraps (vertical guide at k=8 ) .
Fig. 6 : (a) Faithfulness: higher AOPC, lower deletion; (b) remix-invariance via Kendall- τ under input remixes; (c) runtime scaling on the accepted LW environment.
Per-feature ExCIR (validation)
Block (Group) ExCIR
Rank
Feature
Group
CIR
Group (rank)
GroupCIR
1
brake
Control
0.127
Control (1)
0.428
2
tire_rr
Tires
0.119
Environment (2)
0.226
3
rpm
Powertrain
0.119
Dynamics (3)
0.207
4
road_grade
Environment
0.118
Powertrain (4)
0.177
5
maf
Powertrain
0.118
Tires (5)
0.143
TABLE IX : Per-feature and Block CIR (vehicular).
Fig. 7 : Cats–Dogs saliency maps generated by CC-CIR., Left: Val Image, Middle: ExCIR Map, Right: Overlay.
Fig. 8 : Per-class ExCIR scores on digits 1 (left), 8 (middle), and 9 (right).
Metric
Value
Interpretation
Validation Accuracy
97.6%
Base classifier performance
Test Accuracy
96.1%
Generalization check
Kendall– τ (after remix)
0.91
Rank invariance under Y′M
Top–8 overlap
0.88
Leader preservation
Top–10 overlap
0.85
Cross-class consistency
Relative Runtime
1.08×
Over scalar ExCIR
TABLE XI : Summary of CC-CIR multi-output results.
Dataset
Model
Accuracy
Method
Time (s)
CIFAR-10
ResNet-18
95.15±0.21 %
ExCIR
0.16
BlockCIR
0.27
MI
3.14
SHAP
30.12
LIME
1.34
20 Newsgroups
Logistic Reg.
89.3%
ExCIR
0.59
TABLE XV : Same-model predictive performance and attribution runtime. Accuracy is reported once because all methods explain the same predictor.
Appendix figures & tables59 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Value / Computation
n′
5
fi
[1,2,2,3,4]
y′
[0.8,1.1,0.9,1.3,1.5]
Means
f^i=2.4,y^′=1.12
Mid-mean
mi=(f^i+y^′)/2=1.76
Numerator
n′[(f^i−mi)2+(y^′−mi)2]=4.096
Appendix
TABLE S2 : Toy CIR example.
Fig. S1 : BlockCIR: correlated features are grouped by domain and summarized to zB , which is compared with the model output y′ . Each CIR(zB,y′) quantifies the group’s overall contribution, yielding interpretable, non-redundant attributions.
Fig. S2 : Multi-output ExCIR via CCA. Inputs and outputs are projected onto canonical directions z=X′w⋆ and s=Y′u⋆ , then CIR(z,s) is computed on the canonical pair. For vector outputs, use s=Y′u⋆ ; for class-conditioned explanations, use s=Y′wc for class c . Scores are invariant to well-conditioned linear reparameterizations of Y′ .
Fig. S3 : Linear case: CCA and ExCIR produce nearly identical per-feature scores.
Fig. S5 : Nonlinear pattern captured by ExCIR but not CCA—the sinusoidal driver x0 .
Fig. S6 : Nonlinear pattern captured by ExCIR but not CCA—the quadratic driver x1 .
Fig. S7 : Nonlinear pattern captured by ExCIR but not CCA—the step-function driver x2 .
Fig. S8 : 3D feature-space for the sinusoid driver ( x0 ). ExCIR’s nonlinear map reveals a helical structure enabling R2=0.21 .
Fig. S9 : 3D feature-space for the quadratic driver ( x1 ). ExCIR straightens the parabola, enabling R2=0.65 .
Metric
Value
Spearman ρ (CCA vs ExCIR)
0.979
p -value
3.09×10−8
Appendix
TABLE S3 : Linear regime ranking agreement.
Method
@3
@5
@8
CCA ( ∣corr∣ )
0.33
0.20
0.25
ExCIR (feature-space)
0.67
0.60
0.38
Appendix
TABLE S4 : Nonlinear regime: Precision@k for true nonlinear drivers.
Driver
CCA ∣r∣
ExCIR R2
x0 (sinusoid)
0.018
0.211
x1 (quadratic)
0.041
0.647
x2 (step)
0.189
0.038
Appendix
TABLE S5 : Input-space CCA vs feature-space ExCIR for nonlinear drivers.
Fig. S10 : Methodology Schema We compute a bounded Correlation Impact Ratio (CIR) per feature to quantify co-movement with predictions. To keep explanations with less computational cost, we train a lightweight model on a distributionally similar subset and validate transfer so CIR agrees with the full model. The resulting ranking is used for top- k retraining, monitoring, and audits.
Fig. S11 : (a) shows how projection and embedding distances converge, indicating that the lightweight model’s output aligns as a rotated and translated version of the original model’s output, preserving accuracy, (b) compares kernel density estimates for output distributions in both models, showing nearly identical positions with minor alignment shifts indicated by differences in kurtosis, (c) compares the final output distributions of the lightweight and original models, demonstrating that the lightweight model mirrors the original’s output, maintaining similar accuracy.
Method
No. of Features
Accuracy (%)
SHAP-Ranked Features
6
56.2
ExCIR/ExCIR-LW Ranked Features
6
62.7
SHAP-Ranked Features
8
56.23
ExCIR/ExCIR-LW Ranked Features
8
65.1
Appendix
TABLE S7 : Comparison of predictive accuracy between SHAP and ExCIR/ExCIR-LW-ranked features.
Fig. S12 : Distribution of L2 distances between baseline and perturbed ExCIR profiles for CAU–EEG ( M=100 trials).
Fig. S13 : Top- k sufficiency. Test accuracy when keeping only the k highest-ranked features per method.
Fig. S14 : Top- k sufficiency. Test accuracy when keeping only the k highest-ranked features per method.
Fig. S15 : Necessity curves. Test accuracy after removing the top- m features.
Fig. S16 : Noise robustness. ExCIR agreement with its own baseline under small i.i.d. noise.
Fig. S18 : Experiment 5 (agreement–cost). Pareto scatter of wall–clock time vs. CIR rank correlation (full vs. lightweight). Marker size encodes top–8 overlap. The operating point at f=0.20 attains perfect agreement at minimal cost.
Property
SHAP
LIME
ExCIR
Requires model gradients
✗
✗
✓
Requires perturbation/sampling
✓
✓
✗
Observation-only support
✗
✗
✓
Runs on edge devices
✗
✗
✓
Constant memory per feature
✗
✗
✓
Bounded attribution score
✗
✗
✓
Appendix
TABLE S8 : System Readiness Comparison: ExCIR vs. SHAP and LIME
Fig. S19 : Experiment 6 (runtime scaling). End–to–end time grows with the fraction of rows kept. At f=0.20 we already match the full CIR ranking (see Fig. 30(a) ) with a much smaller cost.
Fig. S20 : Calibration and threshold stability. Left: reliability diagram. Right: accuracy vs. decision threshold.
Fig. S21 : Drift sensitivity. Change in ExCIR under simulated drift (positive bars indicate increased importance).
TABLE S12: Top-8 ranked features per method.CAU–EEGLeft→right = higher→lower importance. CIR(LW) aligns with CIR(full), preserving lightweight consistency.
Figure 49
Fig. S27 : Necessity and randomization sanity checks for ExCIR. ( a ) Feature–removal (“necessity”) curves show how test accuracy decreases as the top m ranked features are progressively removed and the model retrained. ExCIR exhibits the steepest accuracy drop, confirming that its highest-ranked features are the ones the model relies on most strongly. ( b ) Randomization sanity test evaluates rank stability under label shuffling and model re-initialization. As expected, correct recomputation on perturbed models drives rank correlation toward 0, restoring sanity; the earlier flat result was traced to reused baseline predictions (see text for details).
Fig. S28 : Noise robustness and correlated-block recovery in ExCIR. ( a ) Rank-stability analysis under additive feature noise shows that ExCIR maintains high Spearman correlation and nearly perfect Top-10 overlap with the noise-free baseline, demonstrating robustness to small perturbations at evaluation time. ( b ) Synthetic correlated-blocks experiment verifies that ExCIR correctly recovers the underlying block hierarchy (B1 > B2 > B3), highlighting its ability to identify dominant correlated groups and preserve meaningful ordering among them.
Fig. S29 : Agreement under growing within-group correlation: ExCIR vs SHAP/LIME (Spearman rank correlation).
Fig. S30 : Agreement–cost trade-off and runtime scaling in lightweight ExCIR. ( a ) Agreement–cost sweep showing Spearman rank correlation and Top- k overlap between ExCIR-LW and full ExCIR across varying lightweight fractions. Even at 20–30% of the validation rows, rank agreement exceeds 0.9 with minimal compute time. ( b ) Runtime scaling curve illustrates that execution time grows linearly with lightweight fraction f , confirming the sub-linear trade-off between fidelity and cost for the vehicular study.
Fig. S31 : Model calibration and threshold stability for ExCIR. ( a ) Calibration curve on the test split shows predicted probabilities closely following the diagonal, indicating well-calibrated model confidence. ( b ) Accuracy–vs–threshold plot demonstrates a broad, smooth optimum, ensuring that ExCIR explanations remain reliable across a range of decision thresholds.
Fig. S32 : Runtime scaling and drift sensitivity of ExCIR. ( a ) ExCIR response to a controlled distributional drift shows the most affected feature groups (e.g., tires, grade, and powertrain load) becoming more salient, confirming interpretability under data shifts. ( b , d ) ExCIR runtime scales linearly with both the number of features d and samples n , consistent with its single-pass closed-form computation. ( c , e ) In contrast, SHAP runtimes remain nearly flat in n but increase sharply with d , as shown for a fixed sampling budget of ∼ 800 point. Together these results demonstrate ExCIR’s efficient scaling with dataset size and its stability under moderate distributional drift.
Fig. S33 : Multi-class extension, whitening, and spurious correlation behavior in ExCIR. ( a ) Multi-class ExCIR shows class-wise CIR values aggregated across categories, where the same dominant features recur with modest variation, confirming the robustness of the multi-output formulation. ( b ) Whitening within correlated feature blocks enhances ExCIR separability, yielding cleaner within-group contrast and improved interpretability. ( c ) Spurious correlation test compares environments A and B: after residualization, ExCIR correctly demotes the spurious driver and restores expected directional behavior, highlighting reliability under confounding and distributional shifts.
Fig. S34 : Uncertainty quantification and counterfactual sanity checks for ExCIR. ( a ) Non-parametric bootstrap confidence intervals show narrow 95% bands for the leading features, indicating strong stability and low variance in ExCIR rankings across resamples. ( b ) Counterfactual sanity curves illustrate model responses under small, realistic perturbations: increasing speed markedly raises predicted risk, increased brake pressure has a mild effect, and higher tire pressure reduces risk—confirming that ExCIR’s global attributions align with domain intuition.
Fig. S35 : Multi-output and image-patch ExCIR evaluations (2×2 panel). ( a ) Directional sensitivity along canonical perturbations shows near-linear response for salient pixels. ( b ) Patch-level ExCIR maps reveal spatially coherent relevance across classes. ( c ) Calibration robustness under remixing confirms ranking stability across softmax perturbations. ( d ) CIR distribution placeholder for illustrating variation in joint influence across patches or outputs.
Fig. S36 : Top-10 Jaccard Overlap between class-wise ExCIR rankings on Digits. Diagonal dominance indicates intra-class consistency, while off-diagonal values reflect cross-class explanation divergence.
Fig. S37 : Top-10 Jaccard Overlap between class-wise ExCIR rankings on Digits. Diagonal dominance indicates intra-class consistency, while off-diagonal values reflect cross-class explanation divergence.
Fig. S38 : AOPC curves and class-level ExCIR heatmap. Panel (a) shows insertion/deletion behavior under ExCIR rankings; panel (b) visualizes global importance for the “dog” class.
Fig. S39 : Left: a validation image. Middle: the same global ExCIR map from Fig. 38(b) . Right: overlay. This overlay is illustrative: the map is global (average over many images), not an instance-specific saliency.
Condition
Sufficiency ↑
Top-8 overlap ↑
Random labels
0.10
0.13
Random features
0.12
0.15
Constant model
0.11
0.12
Trained (baseline)
0.71
1.00
Appendix
TABLE S18 : Negative controls (Vehicular, val).
Fig. S40 : Performance summary of ExCIR across eight evaluation dimensions (Q1–Q8). Each bar represents the normalized score (1–5 scale) for a specific evaluation criterion—fidelity, validity, sufficiency, robustness, group dynamics, efficiency, multi-output stability, and significance. All scores ≥4.7/5.0 indicate strong, stable performance across datasets and evaluation settings.
Aspect
Status quo (SOTA)
ExCIR (ours)
Computation
Sampling/perturbation-heavy; cost grows with k (e.g., SHAP ∼2k )
Closed-form, observation-only; one-time O(n3) then O(n) per feature; independent of k
Ranking, sufficiency
Local-slope emphasis; unclear/unstable global order
Performance-aligned ranking; higher top- k sufficiency (compact subsets)
similar lightweight environment keeps all features, preserves ranking/accuracy.
Calibration
Unbounded, hard to compare across runs
Bounded CIR ∈[0,1] with sensitivity link; comparable across datasets/models/time
Appendix
TABLE S21 : ExCIR novelties vs. SOTA in one view.
Feature
True Role
MI Rank
ExCIR Rank
X1
Predictive
1
1
X2
Redundant
2
3
X3
Noise
3
2
Appendix
TABLE S26 : Feature-importance under redundancy (synthetic experiment).
Item
Setting
Data split
45,000 training / 5,000 validation / 10,000 official test
Architecture
CIFAR-adapted ResNet-18
Input stem
3×3 convolution, stride 1; no initial max-pooling
Normalization
CIFAR-10 channel means and standard deviations
Training augmentation
Random crop with padding 4; horizontal flip
Optimizer
SGD with Nesterov momentum
Appendix
TABLE S27 : Complete CIFAR-10 ResNet-18 training recipe.
Seed
Selected epoch
Validation accuracy (%)
Test accuracy (%)
7
194
95.36
95.13
17
186
95.16
95.02
27
189
94.94
95.48
Mean ± SD
–
95.15±0.21
95.21±0.24
Appendix
TABLE S28 : Clean CIFAR-10 predictive performance across independent seeds. The best epoch is selected using validation accuracy; test accuracy is then measured on the official held-out test partition.
σ
Test accuracy (%)
Spearman vs. clean
Top-8 Jaccard
0.00
95.13
1.00
1.00
0.01
95.01
1.00
1.00
0.03
94.67
1.00
1.00
0.05
93.97
1.00
1.00
0.10
89.85
1.00
1.00
Appendix
TABLE S29 : CIFAR-10 predictive accuracy and ExCIR ranking stability under evaluation-only Gaussian noise.
Fig. S41 : Frozen-model CIFAR-10 robustness. The left axis reports test accuracy; the right axis reports ExCIR rank stability relative to clean inputs. Gaussian noise is added in normalized input space only during evaluation. Predictive accuracy decreases with noise intensity, while Spearman correlation and Top-8 Jaccard similarity remain 1.00 over the tested range.
Fig. S42 : Fixed-seed qualitative CIFAR-10 examples. Five correctly classified truck images (left) and their projected class-conditioned ExCIR patch-score distributions (right) are shown together with predicted confidence. The examples were selected from correctly classified test instances before inspecting their maps.
Fig. S43 : Token-level ExCIR visualization for a correctly classified rec.sport.baseball test document. Bars show the ten largest local ExCIR scores. The most influential observed terms include team , Baltimore , fan , and Orioles .
Fig. S44 : Runtime comparison on CIFAR-10 with ResNet-18. ExCIR is the fastest method in the non-linear vision setting.
Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive information can be distributed across correlated predictors. Existing methods such as SHAP, LIME, HSIC, MI/CMI, and SAGE may therefore produce unstable rankings under multicollinearity or near-duplicate predictors. We propose the Mutual Correlation Impact Ratio Method (MCIR-M), a dependence-aware global feature-importance approach that quantifies the unique predictive information contributed by each feature beyond a selected dependence neighbourhood. MCIR-M introduces the Mutual Correlation Impact Ratio (MCIR), which conditions each feature on strongly dependent neighbours and computes a normalized ratio of conditional to block-level information. The population score lies in [0,1] and equals zero under exact conditional redundancy. We also introduce a lightweight estimation procedure that computes MCIR using a fraction of the available data and evaluates agreement with full-data explanations. Across controlled synthetic redundancy experiments and the UCI HAR benchmark, MCIR shows dependence-aware ranking behaviour, with its clearest advantage under injected near-duplicate predictors. Comparisons with independent and conditional SHAP, SAGE, HSIC, MI-based scores, and CIR-family baselines are mixed across real-data criteria. Reduced explanation samples lower computational burden in the evaluated configurations, while agreement with full-data explanations is assessed separately through ranking, head-set, and faithfulness diagnostics. Overall, MCIR-M provides a practical dependence-aware diagnostic for global explanation under strong feature dependence.
Poushali Sengupta, Sabita Maharjan, Frank Eliassen +2
Department of Informatics, University of Oslo Oslo, Norway · Department of Electronic Systems, Aalborg University Aalborg, Denmark
The rapid adoption of deep learning models in high-risk domains has intensified the need for trustworthy Explainable Artificial Intelligence (XAI). However, objectively evaluating explanation fidelity and aligning XAI metrics with human-centered understanding remain critical open challenges. In this work, we propose a model-agnostic metric, the EPC score, which is an extension of the Explainability-Performance Coefficient (EPC), that quantifies explanation quality by explicitly balancing the trade-off between feature selection sparsity and preserved model performance. Through an empirical validation across tabular, text, and image modalities, we show that the EPC score effectively uncovers operational dependencies among network activations, data dimensionality, and explainer performance. Furthermore, we validate the EPC score against independent human-based explanations, proving that higher EPC scores strongly align with human lexical sentiment judgments and spatial visual annotations.
Christian Oliva, Luis F. Lago-Fernández
Grupo de Neurocomputaci´on Biol´ogica, Departamento de Ingenier´ıa Inform´atica, Escuela Polit´ecnica Superior, Universidad Aut´onoma de Madrid, Spain
This paper investigates a unexplored yet impactful vulnerability in AI explainability used in intrusion detection (IDS): multicollinearity-induced instability. Despite extensive reliance on post-hoc explainability tools such as SHAP or LIME, the impact of correlated features on explanation robustness is not evaluated. We introduce a formal theorem stating that multicollinearity inflates attribution variance. This demonstrates that explanations and feature importances are non-identifiable under multicollinearity. A suite of comprehensive experiments validates the theorem on a representative benchmark dataset, UNSW-NB15. Four widely used families of models are evaluated, including linear, tree-based, kernel, and neural, across full and pruned feature sets based on VIF and correlation thresholding. We propose the novel metric of Explanability Fragility Score and two novel methods to mitigate it with variable integration complexity. CAA-Filtering focuses on stabilising explanations by grouping attributions of trained models. SHARP is a novel training-time regularisation framework that penalises attribution instability, enabling controllable and monotonic improvement of explainability stability. The findings support stable predictive performance, using Kendall's τ to quantify instability across bootstrapped explanations. This work has direct implications for the trustworthiness and reproducibility of XAI in security-critical contexts, and motivates incorporating multicollinearity mitigations into the IDS pipelines, providing a set of guidelines for practitioners.
Ioannis J. Vourganas, Anna Lito Michala
Netrity Ltd, Glasgow, UK · University of Glasgow, Glasgow, UK