Concept-based explanations describe neural network predictions through human-understandable properties of inputs called concepts. The field encompasses approaches that differ in how they define and represent concepts and connect them to model predictions. We introduce a theoretical framework that describes these approaches in a common mathematical language and supports a shared analysis of their properties. For concept discovery, which identifies concepts automatically within a latent space of a trained model, we employ a concept autoencoder view. An encoder extracts concept representations from the model's latent space, and a decoder uses them to reconstruct the original latent representation. The autoencoder's reconstruction error measures how accurately its decoder recovers the original latent representation. We revisit model completeness: how well the concepts can reproduce the model's outputs. We show that model incompleteness of the concepts can be bounded by the autoencoder's reconstruction error. The autoencoder view also provides a common way to define individual concept attributions, which measure each concept's contribution to a prediction. We establish when these attributions sum to the model's prediction, and bound the discrepancy otherwise, thus providing attribution completeness guarantees.
Figures & tables
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Z
Vi
πi(z)
σi
D
γ affine
Explainable by design
CBM, probabilities
[0,1]k
[0,1]
zi
\id
\id
if g
CBM, logits
Rk
R
zi
s
\id
if g
PCBM, LF-CBM
Rk
R
zi
–
\id
yes
CEM
∏iR2m
R2m
zi
s(⟨w,v⟩+β)
\id
no
Concept detection
Appendix
Table 1: Instantiations of concept-based methods in our framework: CBMs ( Koh et al., 2020 ) , PCBM ( Yuksekgonul et al., 2023 ) , LF-CBM ( Oikarinen et al., 2023 ) , CEM ( Espinosa Zarlenga et al., 2022 ) , TCAV ( Kim et al., 2018 ) , CAR ( Crabbé & Schaar, 2022 ) , ACE ( Ghorbani et al., 2019 ) , CRAFT ( Fel et al., 2023b ) , ICE ( Zhang et al., 2021 ) , SAE ( Bricken et al., 2023 ) , and MCD ( Vielhaben et al., 2023 ) . Columns give the latent space Z , the concept representation space Vi , the probe πi , the activation function σi , the decoder D , and whether the decoded head γ=g∘D is affine; “if g ” means that γ is affine whenever g is. A dash marks a component the method does not specify. s is the logistic sigmoid, U is a dictionary with concept directions ui in its rows, NMF encoders compute nonnegative least-squares coefficients with U fixed, and Qi projects channels onto the i -th MCD subspace, with the residual subspace included as a concept.
Boundary
Method
k
ρ\fiderror
α
κ
ρ^\modelerror
5
0.106
0.270
0.394
0.967
PCA
25
0.103
0.266
0.386
0.927
50
0.099
0.261
0.381
0.939
5
0.106
0.269
0.395
0.975
NMF
25
0.102
0.264
0.389
0.960
50
0.101
0.263
0.385
0.964
Appendix
Table 2: Tightness of the fidelity and model completeness bounds, mean over three seeds; the standard deviation over seeds is at most 0.012 for ρ^\modelerror and at most 0.002 for the other ratios. ρ\fiderror=ακ is the tightness of , with alignment α and spatial coherence κ from . ρ^\modelerror upper-bounds the tightness of .
\adderror(A,φ)
ρ\atterror×103
Boundary
k
Bound
Ins.
Occ.
G × I
Ins.
Occ.
G × I
5
2281
1.31
1.39
0.10
3.0
3.0
2.9
layer4
25
2381
2.34
2.92
0.06
2.9
1.9
2.4
50
2749
3.49
4.20
0.06
2.7
1.4
2.0
5
2593
0.78
0.87
0.25
3.0
2.8
2.9
penultimate
25
1909
1.70
3.32
0.13
3.5
2.2
2.9
Appendix
Table 3: Tightness of the attribution bound for the nonlinear autoencoder, mean over three seeds. Bound is \fiderror(A)+ME[∥C(x)∥24] with \fiderror on the explained scores, shared by insertion (Ins.), occlusion (Occ.), and gradient-times-input (G × I), and ρ\atterror is the tightness of the attribution bound obtained by combining with . The bound, and with it ρ\atterror , varies across seeds by up to 40% because M is estimated by sampling, whereas \adderror varies by at most 0.4 . For the four methods with affine decoders, \adderror<2⋅10−5 and ρ\atterror=1 for all three attribution functions.
Concept-based explanations offer a promising approach for explaining the predictions of deep neural networks in terms of high-level, human-understandable concepts. However, existing methods either do not establish a causal connection between the concepts and model predictions or are limited in expressivity and only able to infer causal explanations involving single concepts. At the same time, the parallel line of work on formal abductive and contrastive explanations computes the minimal set of input features causally relevant for model outcomes but only considers low-level features such as pixels. Merging these two threads, in this work, we propose the notion of concept-based abductive and contrastive explanations that capture the minimal sets of high-level concepts causally relevant for model outcomes. We then present a family of algorithms that enumerate all minimal explanations while using concept erasure procedures to establish causal relationships. By appropriately aggregating such explanations, we are not only able to understand model predictions on individual images but also on collections of images where the model exhibits a user-specified, common behavior. We evaluate our approach on multiple models, datasets, and behaviors, and demonstrate its effectiveness in computing helpful, user-friendly explanations.
Ronaldo Canizales, Divya Gopinath, Corina Păsăreanu +1
Colorado State University, Fort Collins CO, USA · KBR Inc., NASA Ames, Moffett Field CA, USA · Carnegie Mellon University, Pittsburgh PA, USA
Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are not automatically trustworthy. We introduce ConceptSMILE, a model-agnostic perturbation-based auditing framework for evaluating the reliability of concept-based explanations. Rather than replacing SMILE, ConceptSMILE extends its perturbation-based logic from feature- or region-level attribution to the auditing of human-understandable concept explanations. The framework perturbs input regions, measures concept-response shifts, applies locality weighting, and fits an XGBoost surrogate to approximate local concept behaviour. Reliability is assessed through attribution accuracy, surrogate fidelity, faithfulness, stability, and consistency. We evaluate ConceptSMILE on retinal fundus images by comparing MedSAM-derived visual concepts with VLM-based semantic concepts. Results show that reliability varies across concepts and pathways: MedSAM achieves stronger spatial attribution and the highest surrogate fidelity (R2=0.8503, Rw2=0.8465), while the VLM pathway shows stronger vessel faithfulness and stronger stability under selected artefact conditions. ConceptSMILE provides an independent audit layer for evaluating the trustworthiness of concept-based XAI.
School of Computer Science University of Hull Hull, United Kingdom · PhD Researcher (Engineering), WMG University of Warwick Coventry, United Kingdom · School of Computing Newcastle University Newcastle upon Tyne, United Kingdom
Techniques for concept extraction, such as sparse autoencoders and transcoders, aim to extract high-level symbolic concepts from low-level nonsymbolic representations. When these extracted concepts are used for downstream tasks such as model steering and unlearning, it is essential to understand their guarantees, or lack thereof. In this work, we present a unified theoretical framework for unsupervised concept extraction, in which we frame the task of concept extraction as identifying a generative model. We present a general meta-theorem for identifiability, which reduces the problem of establishing identifiability guarantees to the problem of characterizing the intersection of two sets. As we demonstrate on a range of widely-used approaches, this meta-theorem substantially simplifies the task of proving such guarantees, thus paving the way for the development of new, principled approaches for concept extraction.