When Are Concept Bottleneck Model Explanations Faithful and Compact?
Authors: Stefano Teso, Emanuele Marconato, Steve Azzolin, Antonio Vergari
Organizations: CIMeC & DISI University of Trento · Department of Mathematical Sciences University of Copenhagen · DISI University of Trento · School of Informatics University of Edinburgh
Concept bottleneck models (CBMs) are neural classifiers that allow to explain their decisions via high-level concepts, potentially enabling understanding, steering and debugging. However, their explanations are often derived heuristically. Building on formal explainability, we argue they should also be faithful, i.e., not misreport which concepts actually matter. We show that, for widespread CBM architectures, including recent VLM-based variants, faithful explanations must include all concepts in the bottleneck, compromising interpretability when this is large. This result applies to both heuristic and faithful-by-construction formal explanations. To encourage the existence of compact faithful explanations, we suggest i) modeling concepts probabilistically as binary or categorical random variables (rather than logits), and ii) employing per-concept training-time sparsification via group lasso (rather than regular elastic net). We also extend algorithms from formal explainability to CBMs, and show they outperform natural heuristics in terms of guarantees and explanation size. Overall, our work warns against naive interpretability claims and provides formal conditions and practical strategies for ensuring CBMs are as interpretable as advertised.
Figures & tables
Figure 1: Current heuristic explanation for CBMs can be unfaithful as shown for a CBM combining a neural backbone generating concept activations s∈S and a linear inference layer with per-class weights wy , with y∈[m] . Left: Standard heuristics compute explanations E by looking at the magnitude of activations, weights, or both, cf. Eq. 2 . Right: A heuristic explanation E that covers the k=2 top weights: it justifies two decisions on the ground that the same two activations attain the same values. Yet, the decisions differ: the CBM must be looking at activations outside of E to produce them. Hence, E is insufficient to fully justify either decision.
Act. S
Wei. W
Representative Models (Not Exhaustive)
(−∞,∞)b
Dense
CBM ( Koh et al., 2020 ) , ProbCBM ( Kim et al., 2023 ) , MCBM ( Almudévar et al., 2026 )
(−∞,∞)b
Elastic Net
LF-CBM ( Oikarinen et al., 2023 ) , VLG-CBM ( Srivastava et al., 2024 )
[α,∞)b,α≥0
Elastic Net
Selective CBM ( Schrodi et al., 2025 )
[−1,1]b
Dense
LaBo ( Yang et al., 2023 )
⊆[0,1]b
Dense
CB2M ( Steinmann et al., 2024 ) , GlanceNets ( Marconato et al., 2022 )
Table 1: Representative CBMs grouped by domain of concept activations S ( P1 ) and structure of the inference layer’s weights W ( P2 ). Constraining S and W can make it more difficult to construct counterexamples that flip the prediction, allowing (weak) AXp to be smaller. Few CBMs employ probabilistic activations and none group lasso.
Dataset
m
b
Cat.
Train
Val
Test
CelebA
2
39
0
162,770
19,867
19,962
Derm7pt
5
28
7
413
203
395
WBCAtt
5
31
11
6,169
1,030
3,099
CUB
200
312
28
5,394
600
5,794
Table 2: Datasets. Cat. is the # of categorical concepts.
F1C
F1Y
Dataset
prob
logit
prob
logit
CelebA
0.760
0.760
0.989
0.991
Derm7pt
0.676
0.670
0.653
0.666
WBCAtt
0.877
0.872
0.943
0.948
CUB
0.796
0.781
0.817
0.832
Table 3: Probabilities are competitive with logits in terms of F1 score.
Figure 2: Modeling concepts as probabilities shrinks AXp s on Derm7pt (left), WBCAtt (middle), and CUB (right). Top row: logits. Bottom row: probabilities. CelebA is left to Section D.2 .
F1C
F1Y
Dataset
dense
enet
glasso
CUB
0.796
0.817
0.759
0.756
Derm7pt
0.676
0.653
0.648
0.638
WBCAtt
0.877
0.943
0.916
0.927
LF-CBM
–
0.545
0.589
0.564
VLG-CBM
–
0.626
0.632
0.576
Table 4: glasso shrinks faithful explanations at less than 10%F1Y cost . VLM-CBMs have no F1C as they have no ground-truth annotations.
Derm7pt ( b=28 )
WBCAtt ( b=31 )
CUB ( b=312 )
Act
Wei
Prod
EA+
Act
Wei
Prod
EA+
Act
Wei
Prod
EA+
dense
34.1
47.3
35.1
29.6
34.0
57.6
35.6
26.3
16.6
28.6
18.5
15.4
enet
27.7
36.9
{\color[rgb]{1,0,0}36.9}
9.5
29.9
31.0
30.9
7.3
13.1
{\color[rgb]{1,0,0}87.0}
{\color[rgb]{1,0,0}87.0}
10.2
glasso
27.4
8.1
7.8
7.3
28.1
11.7
13.6
8.1
14.3
9.8
9.4
7.7
LF-CBM ( b=202 )
VLG-CBM ( b=419 )
LaBo ( b=10,000 )
Act
Wei
Prod
EA+
Act
Wei
Prod
EA+
Act
Wei
Prod
EA+
Table 5: glasso consistently yields smaller faithful explanations than enet . Explanation size is reported as a percentage ( i.e. , normalized by bottleneck size b ) for concision. Bold indicates best in column and red no shrinkage w.r.t. the dense baseline.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8
CBM
CBM + glasso
EA
EA+
EA
EA+
Dataset
b
Size
Calls
Size
Calls
Size
Calls
Size
Calls
CelebA
39
20.8 ± 4.6
39
12.6 ± 3.7
102
–
–
–
–
Derm7pt
28
17.1 ± 1.8
28
8.3 ± 2.0
53
3.1 ± 2.9
28
2.1 ± 0.8
33
WBCAtt
31
9.9 ± 3.6
31
8.2 ± 3.0
60
2.6 ± 0.9
31
2.4 ± 0.8
37
CUB
312
169.5 ± 28.5
312
46.3 ± 18.6
221
47.0 ± 9.9
312
23.5 ± 3.6
134
Appendix
Table 6: Comparison between EA and our EA+ on regular CBMs with prob. activations. Bold indicates a size improvement of at least one activation.
VLM-CBM
VLM-CBM + glasso
EA
EA+
EA
EA+
Dataset
b
Size
Calls
Size
Calls
Size
Calls
Size
Calls
LF-CBM
202
201.0 ± 0.8
202
200.9 ± 0.9
633
36.9 ± 5.8
202
35.3 ± 5.9
308
VLG-CBM
419
416.5 ± 1.5
419
416.3 ± 1.5
1286
78.0 ± 8.4
419
66.5 ± 11.3
593
Appendix
Table 7: Comparison between EA and our EA+ on VLM-CBMs. Bold indicates a size improvement of at least one activation.
CBM
CBM + glasso
\textscTopK∗
EA+
\textscTopK∗
EA+
Dataset
b
Calls
Time
Calls
Time
Calls
Time
Calls
Time
CelebA
39
6
0.001
102
0.007
–
–
–
–
Derm7pt
28
6
0.004
53
0.025
6
0.004
33
0.016
WBCAtt
31
6
0.005
60
0.040
6
0.006
37
0.026
CUB
312
9
0.704
221
16.607
9
0.704
134
10.667
Appendix
Table 8: Runtime cost of all algorithms for regular CBMs . Time is measured in seconds.
VLM-CBM
VLM-CBM + glasso
\textscTopK∗
EA+
\textscTopK∗
EA+
Dataset
b
Calls
Time
Calls
Time
Calls
Time
Calls
Time
LF-CBM
202
8
0.005
633
0.188
8
0.006
308
0.147
VLG-CBM
419
9
0.012
1286
1.095
9
0.020
593
1.215
Appendix
Table 9: Runtime cost of all algorithms for VLM-CBMs . Time is measured in seconds.
CBM
CBM + glasso
Dataset
Act
Wei
Prod
EA+
Act
Wei
Prod
EA+
CelebA
35.0 ± 3.8
15.7 ± 4.9
29.6 ± 5.1
12.6 ± 3.7
–
–
–
–
Derm7pt
9.5 ± 2.8
13.4 ± 5.0
9.7 ± 2.5
8.3 ± 2.0
7.7 ± 1.8
2.3 ± 0.7
2.2 ± 0.9
2.1 ± 0.8
WBCAtt
10.7 ± 4.3
17.5 ± 4.8
11.0 ± 4.9
8.2 ± 3.0
8.7 ± 3.2
3.5 ± 0.9
4.1 ± 1.3
2.4 ± 0.8
CUB
50.3 ± 19.0
84.7 ± 36.8
55.4 ± 21.8
46.3 ± 18.6
44.4 ± 14.6
30.0 ± 5.2
28.2 ± 4.3
23.5 ± 3.6
Appendix
Table 10: Comparison between explanation sizes across datasets for regular CBMs.
CBM
CBM + glasso
Dataset
Act
Wei
Prod
EA+
Act
Wei
Prod
EA+
LF-CBM
202.0 ± 0.2
201.8 ± 0.5
201.8 ± 0.5
200.9 ± 0.9
201.3 ± 1.3
39.0 ± 6.8
39.1 ± 6.8
35.3 ± 5.9
VLG-CBM
418.6 ± 0.6
418.1 ± 1.1
418.2 ± 1.1
416.3 ± 1.5
418.0 ± 2.1
71.6 ± 12.2
71.7 ± 12.2
66.5 ± 11.3
Appendix
Table 11: Comparison between explanation sizes for CUB using VLM-CBM vocabularies.
TopK
Dataset
Act
Wei
Prod
EA+
CelebA
0%
10%
1%
100%
Derm7pt
37%
2%
37%
100%
WBCAtt
9%
0%
12%
100%
CUB
5%
1%
5%
100%
Average
12.8%
3.5%
13.8%
100%
Appendix
Table 12: Whereas EA+ always computes AXp s, heuristics do not . Left : dense CBMs. Right : glasso -sparsified CBMs.
Figure 3: Modeling concepts as probabilities shrinks AXp s the most on all datasets. From the top: CelebA , Derm7pt , WBCAtt , CUB ). Left to right: logits, ReLU, probabilities.
Figure 4: enet (middle) works partially, glasso (below) works consistently.
Figure 5: Q3: For VLM-CBMs, glasso outperforms other sparsification strategies , as shown here with CUB . Each VLM-CBM employs its own LLM-generated vocabulary, with LaBo ’s sporting 10,000 activations.
Figure 6: Selective CBMs behave like ReLU CBMs and worse than sparsified CBMs in terms of AXp size, shown here for CelebA (top) and WBCAtt (bottom).
Explainability of deep learning algorithms is critical for computer-vision applications with high-stake decisions. Concept bottleneck models (CBM) have recently shown promising performance to provide explainable and accurate predictions for classification problems, based on a bottleneck of high-level concepts. Existing CBM methods rely on a linear aggregation of the concept scores to compute predictions. However, a large number of concepts is often used in this linear approach, which undermines explainability and favors information leakage. In general, the underlying relation between concepts and output logits is not linear. Therefore, we introduce Hoeffding Concept Bottleneck Models (HCBM), which build on the Hoeffding functional decomposition of gradient-boosted trees to provide non-linear and sparse aggregations of concept scores, and generate compact predictions using prime implicants. HCBM are proved to be robust to interconcept leakage, and outperform standard linear CBM in practice, as shown in extensive experiments. Beyond classification, HCBM can be adapted to object detection, and we focus on a challenging case with overhead images to show the high performance of HCBM in these settings.
Concept Bottleneck Models (CBMs) are interpretable-by-design neural networks that detect human-understandable concepts from the input and use them to generate predictions. By allowing users to inspect the concepts underlying a prediction and explore how predictions change under alternative concept configurations, CBMs have emerged as one of the most prominent approaches to supporting human-AI collaboration. However, user studies investigating their actual effectiveness as decision-support systems remain limited. We present two large-scale user studies (N participants = 705, N observations = 6,959) evaluating how concept-based explanations and user interventions on the model's concepts affect the performance of the human-AI team in two distinct binary classification tasks. Our results show that CBMs, and particularly their interactive component, can improve human-AI team accuracy relative to both unaided human performance and performance with non-interpretable AI support. However, these benefits emerge only under certain conditions: classification tasks perceived as difficult, easily identifiable concepts, and active interaction with the model. We also discuss how inaccurate concept detection may undermine users' trust in the model. Overall, this work provides practical guidance for the deployment of CBMs as effective decision-support tools.
Concept Bottleneck Models (CBMs) have become a popular approach to enable interpretability in neural networks by constraining classifier inputs to a set of human-understandable concepts. While effective, current models embed concepts in flat Euclidean space, treating them as independent, orthogonal dimensions. Concepts, however, are highly structured and organized in semantic hierarchies. To resolve this mismatch, we propose Hyperbolic Concept Bottleneck Models (HypCBM), a post-hoc framework that grounds the bottleneck in this structure by reformulating concept activation as asymmetric geometric containment in hyperbolic space. Rather than treating entailment cones as a pre-training penalty, we show they encode a natural test-time activation signal: the margin of inclusion within a concept's entailment cone yields sparse, hierarchy-aware activations without any additional supervision or learned modules. We further introduce an adaptive scaling law for hierarchically faithful interventions, propagating user corrections coherently through the concept tree. Empirically, HypCBM rivals post-hoc Euclidean models trained on 20× more data in sparse regimes required for human interpretability, with stronger hierarchical consistency and improved robustness to input corruptions.
Daniel Uyterlinde, Swasti Shreya Mishra, Pascal Mettes
Informatics Institute, University of Amsterdam, The Netherlands