cs.SDSep 27, 2026

Uncovering shortcut learning in audio classifiers by discovering recurring concepts in temporal explanations

Authors: Cecilia Bolaños, Luciana Ferrer, Magdalena Fuentes

Organizations: Departamento de Computación, FCEyN, UBA · ICC, CONICET-UBA, Buenos Aires, Argentina · Music and Audio Research Lab, New York University, USA · Integrated Design & Media, New York University, USA

Abstract

Correlations between events in machine learning datasets may result in shortcut learning, where models learn to predict the target event based on the presence of a correlated event. When these correlations are spurious -- arising from data collection artifacts -- models are likely to perform poorly in practice. We propose a pipeline to uncover shortcut learning in audio classifiers by discovering recurring concepts in their temporal explanations. Specifically, we isolate audio segments that explain classifier decisions, caption them with an ensemble of Large Audio-Language Models, and use a Large Language Model to extract recurring concepts. The resulting concepts can be audited by humans to uncover potential shortcut learning. We evaluate our framework using datasets curated from AudioSet Strong, controlling for the presence or absence of spurious correlations. Results show that this approach reliably uncovers learned shortcuts, such as the model relying on the presence of "laughter" to predict "applause".

Figures & tables

Explore similar work

CardsList
  1. SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

    May 13, 2026Giries Abu Ayoub, Morad Tukan, Loay MualemFew-Shot LearningAudio Understanding

  2. Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers

    Jul 17, 2026Shilin Gao, Mark J. F. Gales, Kate M. KnillLinguisticsShortcut Learning

  3. APEX: Audio Prototype EXplanations for Classification Tasks

    May 11, 2026Piotr Kawa, Kornel Howil, Piotr Borycki +3Audio UnderstandingAudio Editing