A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees
Organizations: Masaryk University, Faculty of Informatics, Brno, Czechia
Abstract
Concept-based explanations describe neural network predictions through human-understandable properties of inputs called concepts. The field encompasses approaches that differ in how they define and represent concepts and connect them to model predictions. We introduce a theoretical framework that describes these approaches in a common mathematical language and supports a shared analysis of their properties. For concept discovery, which identifies concepts automatically within a latent space of a trained model, we employ a concept autoencoder view. An encoder extracts concept representations from the model's latent space, and a decoder uses them to reconstruct the original latent representation. The autoencoder's reconstruction error measures how accurately its decoder recovers the original latent representation. We revisit model completeness: how well the concepts can reproduce the model's outputs. We show that model incompleteness of the concepts can be bounded by the autoencoder's reconstruction error. The autoencoder view also provides a common way to define individual concept attributions, which measure each concept's contribution to a prediction. We establish when these attributions sum to the model's prediction, and bound the discrepancy otherwise, thus providing attribution completeness guarantees.
Figures & tables
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | affine | |||||
| Explainable by design | ||||||
| CBM, probabilities | if | |||||
| CBM, logits | if | |||||
| PCBM, LF-CBM | – | yes | ||||
| CEM | no | |||||
| Concept detection | ||||||
| Boundary | Method | |||||
|---|---|---|---|---|---|---|
| 5 | 0.106 | 0.270 | 0.394 | 0.967 | ||
| PCA | 25 | 0.103 | 0.266 | 0.386 | 0.927 | |
| 50 | 0.099 | 0.261 | 0.381 | 0.939 | ||
| 5 | 0.106 | 0.269 | 0.395 | 0.975 | ||
| NMF | 25 | 0.102 | 0.264 | 0.389 | 0.960 | |
| 50 | 0.101 | 0.263 | 0.385 | 0.964 |
| Boundary | Bound | Ins. | Occ. | G I | Ins. | Occ. | G I | |
|---|---|---|---|---|---|---|---|---|
| 5 | 2281 | 1.31 | 1.39 | 0.10 | 3.0 | 3.0 | 2.9 | |
| layer4 | 25 | 2381 | 2.34 | 2.92 | 0.06 | 2.9 | 1.9 | 2.4 |
| 50 | 2749 | 3.49 | 4.20 | 0.06 | 2.7 | 1.4 | 2.0 | |
| 5 | 2593 | 0.78 | 0.87 | 0.25 | 3.0 | 2.8 | 2.9 | |
| penultimate | 25 | 1909 | 1.70 | 3.32 | 0.13 | 3.5 | 2.2 | 2.9 |