Revisiting Explainable AI through Model-Independent Concept Dictionaries
Authors: Thomas Schnake, Doreen Schöppenthau, Alexander Meyer, Jacques Corbeil, Klaus-Robert Müller, Grégoire Montavon
Organizations: Department of Chemistry, Chemical Physics Theory Group, University of Toronto, Toronto, Canada · Vector Institute for Artificial Intelligence, Toronto, Canada · Institute for AI in Medicine (IKIM), Charité – Universitätsmedizin Berlin, Germany · Department of molecular medicine, Université Laval, Québec, QC, Canada · Mila – Quebec Artificial Intelligence Institute, Canada · Machine Learning Group, Technische Universität Berlin, Germany · Department of Artificial Intelligence, Korea University, Seoul, Korea · Max Planck Institute for Informatics, Saarbrücken, Germany
Modern applications of AI rely on increasingly complex models. Explainable AI (XAI) has emerged as a set of techniques aimed at improving model transparency. However, existing XAI methods typically assume input features to be inherently interpretable, or they rely on intermediate internal abstractions that are difficult to characterize and highly architecture-specific, hindering consistent use across models. To address these limitations, we propose DictXAI, a method that defines concepts directly in the input domain via a dictionary---a large, potentially overcomplete set of predefined elements, each carrying an interpretable meaning. Technically, DictXAI first computes a sparse code of the input and then attributes the model's prediction to the associated dictionary elements. We demonstrate the actionable nature of DictXAI explanations, showing that they can attribute AI malfunctions (e.g., Clever Hans effects) directly to identifiable artifact patterns in the data, while fostering human-AI alignment on intricate biomedical signals. We further demonstrate our method's ability to operate across a wide variety of dictionaries, including learned image bases, analytically defined waveforms for electrocardiography, and experimentally acquired dictionary elements. Overall, our results show that DictXAI provides more interpretable, actionable, and architecture-agnostic insights than classical XAI or existing concept-based approaches.
Figures & tables
Figure 1: Overview of DictXAI. Left (Classical XAI): Standard workflows perform prediction and subsequently explain model decisions by computing attribution heatmaps directly on raw input pixels. Right (DictXAI, ours): Our approach introduces a predefined dictionary of concepts (here, multi-scale oriented Gabor filters). First, sources are inferred via sparse coding of the input ①. The input is then reconstructed via the decoder and passed to the classifier to predict ②, after which explanations are propagated back through the classifier and the decoder to the dictionary coefficients ③. This produces sparse explanations (where crosses denote zero attribution due to sparsity and red dot sizes indicate attribution magnitude) that are semantically interpretable along explicit axes (e.g., angle and scale) while preserving foundational XAI desiderata such as conservation and continuity.
Figure 2: Dictionary designs and concept attribution profiles on MNIST. A Subset of atoms of intermediate L2 norm from a dictionary learned on the MNIST training set via sparse dictionary learning ( K=800 ; MSE=0.017 on 1000 test images; λ=10−3 ). B Parameter sweeps from an analytically defined Gabor dictionary ( K≈10000 ; MSE=0.032 on the same 1000 test images; λ=10−3 ). Each row isolates variations along a single parameter while holding others fixed: orientation θ (top), envelope scale σ (middle), and vertical center y0 (bottom). C Class-aggregated profiles for the Gabor dictionary across test digits (classes 5–9, 30 samples each). Panels display Hinton diagrams where marker size denotes magnitude and color denotes sign (red: positive, blue: negative). Top row: mean sparse coding coefficients; bottom row: DictXAI explanation (mean relevance scores). Values are plotted across the flattened 5×5 spatial grid ( y -axis, indices 0–24, row-major) and orientation θ∈[0,π] ( x -axis), with the scale parameter σ averaged out.
Figure 3: Horizontal-stripe variant, class “6”. Upper row: the original input, and the same sample after the oracle removes the top- k most artifactual elements encoded by each method. Elements are ranked by their overlap Sj with the watermark mask over a fixed 500 -image training subset. To keep the comparison fair across differently sized bases, k is fixed at the grid point closest to 5% of each method’s total concept count (between 4.1 and 5.7% ). Lower row: pixel-wise LRP- γ heatmaps of the base model and of models retrained on the resulting pruned data, each averaged over five seeds. The color scale for each panel is normalized to that of the base model on the unpruned, contaminated data.
Figure 4: Remove-and-retrain (ROAR) curves for the three watermark variants: overall accuracy on the decorrelated test set as a function of the fraction of concepts removed before fine-tuning. Lines and shading represent mean and std over five seeds, and stars mark maxima. For each XAI method, extracted concept elements are inspected pixel-wise, ranked by their overlap with the watermark ( Sj ), and removed in this order. The three horizontal references are fine-tuning on watermark-free data (dashed, upper bound), fine-tuning without any removal (dash-dotted, the k=0 control), and the original model (dotted).
Figure 5: Explaining the prediction of QRS segment width in an ECG signal. A : Zoomed-in excerpt of the 10-second Lead I recording for two samples from the MIMIC-IV-ECG dataset, one from each class. B : Explanation of the trained ML model, showing amplitude-driven explanations for the LRP baseline and DictXAI’s significantly more discriminative explanations attributing class evidence to distinct AMS widths. C : Verification that the ML model’s strategy of relying mainly on class 2 features lacks robustness to noise. Adding moderate noise ( std≤0.3 ) degrades feature detection, causing the model prediction difference ( y2−y1 ) to revert toward the Class 1 baseline.
Figure 6: Comparison between classical spectrum-based explanations (top) and DictXAI microbe-based explanations (bottom) of spectral similarity, computed on raw spectra (left) versus band-pass filtered spectra (right). Top: Traces depict MALDI-TOF mass spectra for two co-cultures ( x,x′ ), with red bars highlighting spectral regions contributing positively to predicted similarity. Bottom: DictXAI bipartite relevance graphs linking dictionary coefficients from the two mixtures. Taxa in black and gray denote species present and absent from the respective mixture; edge thickness denotes the attributed concept-level relevance. Filtering isolates true biological matches (e.g., C. tertium , DH5a-K12 ) from broad spectral background.
Explainable AI (XAI) techniques are increasingly important for the validation and responsible use of modern deep learning models, but are difficult to evaluate due to the lack of good ground-truth to compare against. We propose a framework that serves as a quantifiable metric for the quality of XAI methods, based on continuous input perturbation. Our metric formally considers the sufficiency and necessity of the attributed information to the model's decision-making, and we illustrate a range of cases where it aligns better with human intuitions of explanation quality than do existing metrics. To exploit the properties of this metric, we also propose a novel XAI method, considering the case where we fine-tune a model using a differentiable approximation of the metric as a supervision signal. The result is an adapter module that can be trained on top of any black-box model to output causal explanations of the model's decision process, without degrading model performance. We show that the explanations generated by this method outperform those of competing XAI techniques according to a number of quantifiable metrics.
Amritpal Singh, Andrey Barsky, Mohamed Ali Souibgui +2
Computer Vision Center, Barcelona, Spain · Autonomous University of Barcelona, Spain
Complex AI systems make better predictions but often lack transparency, limiting trustworthiness, interpretability, and safe deployment. Common post hoc AI explainers, such as LIME, SHAP, HSIC, and SAGE, are model agnostic but are too restricted in one significant regard: they tend to misrank correlated features and require costly perturbations, which do not scale to high dimensional data. We introduce ExCIR (Explainability through Correlation Impact Ratio), a theoretically grounded, simple, and reliable metric for explaining the contribution of input features to model outputs, which remains stable and consistent under noise and sampling variations. We demonstrate that ExCIR captures dependencies arising from correlated features through a lightweight single pass formulation. Experimental evaluations on diverse datasets, including EEG, synthetic vehicular data, Digits, and Cats-Dogs, validate the effectiveness and stability of ExCIR across domains, achieving more interpretable feature explanations than existing methods while remaining computationally efficient. To this end, we further extend ExCIR with an information theoretic foundation that unifies the correlation ratio with Canonical Correlation Analysis under mutual information bounds, enabling multi output and class conditioned explainability at scale.
Institute of Informatics, University of Oslo, Oslo, Norway · Department of Computer Science, Oslo Metropolitan University, Oslo, Norway · Aalborg University, Denmark
Mechanistic interpretability aims to explain a model's behavior by identifying causally responsible internal structures. Dictionary-based explainers such as sparse autoencoders and transcoders are a primary tool, but their faithfulness under out-of-distribution (OOD) shift has received little systematic attention. We show that distribution shift rotates the subspace that the model actively uses, misaligning the explainer's dictionary trained on in-distribution (ID) activations. We formalize this misalignment as the faithfulness gap, a geometric distance between the ID dictionary and the OOD-active subspace, and show that it controls OOD faithfulness degradation. To reduce this gap, we propose the Geometry-Adaptive Explainer (GAE), which realigns the explainer's dictionary with the OOD-active subspace while preserving the original feature structure. This requires only unlabeled OOD activations and no gradient updates. We prove that GAE improves over the unadapted ID explainer, with excess loss bounded quadratically by the second-moment shift. Empirically, GAE even matches or surpasses all training-based baselines in causal faithfulness across multiple models and OOD settings.