ICON Decomposition: Auditing deep neural networks for shortcuts by decomposing layer-wise representations using concepts
Authors: Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer, Marc-Andre Schulz, Nys Tjade Siegel, Maximilian Dreyer, Frederik Pahde, Wojciech Samek, +2 more
Organizations: Hertie Institute for AI in Brain Health, University of Tübingen, Tübingen, Germany. · Department of Psychiatry and Neurosciences, Charité - Universitätsmedizin Berlin, Berlin, Germany. · Department of Psychology, Humboldt-Universität zu Berlin, Berlin, Germany. · Chair of Statistics, Humboldt-Universität zu Berlin, Berlin, Germany. · Tübingen AI Center, University of Tübingen, Tübingen, Germany. · German Center for Mental Health (DZPG), Tübingen, Germany. · Department of Artificial Intelligence, Fraunhofer Heinrich Hertz Institute, Berlin, Germany. · Department of Electrical Engineering and Computer Science, Technische Universität Berlin, Berlin, Germany.
Deep neural networks often exploit spurious associations, a failure known as shortcut learning. Before deployment, models should be audited for reliance on a set of concepts, such as acquisition artifacts or demographics. Current methods, such as linear probes and concept activation vectors, measure reliance by asking whether each concept, in isolation, is decodable from a layer. Their scores therefore reflect not only reliance but also correlations in the audit dataset. We introduce Independent Canonical cONcept (ICON) decomposition, which quantifies the share of a layer's variance each concept explains, conditional on all other concepts and the outcome. ICON scores are variance shares, comparable across layers and between continuous and categorical concepts. ICON also reports the share the set leaves unexplained. On simulated data, ICON recovers the true importance more accurately than seven baselines. On skin-cancer and neuroimaging models, ICON distinguishes learned shortcuts from correlated concepts, confirmed by retraining and out-of-distribution tests.