cs.SDSep 27, 2026

Mask-Induced Displacement in Audio XAI via Logit Trajectory Decomposition

Authors: Nico García-Peguinho, David Kelly, Fabrizio Smeraldi, Anna Xambó Sedó

Organizations: School of Electronic Engineering and Computer Science, Queen Mary University of London · Department of Informatics, King's College London · Department of Informatics, King’s College London

Abstract

Perturbation-based XAI methods for audio classifiers often estimate feature importance by masking spectrogram regions and crediting output changes to the retained signal. Yet they typically assume the fill (the mask replacement) is negligible. We propose logit-space trajectory decomposition to examine this assumption. An on-axis component captures output along a line connecting the fully filled (occluded) spectrogram to the fully retained original; an off-axis component captures perpendicular displacement. We evaluate across three fills, three audio classifiers, and 1,100 AudioSet clips. We demonstrate that no fill is acoustically neutral under full occlusion: Zero fill activates silence and Gaussian Noise fill activates broadband noise. Under partial masking, 41-77% of output displacement is off-axis, with the direction of the residual stable across mask retention fraction and specific to each model-fill combination. Attribution methods are unevenly exposed to off-axis displacement through their sampling and weighting strategies, revealing apparatus-dependence. Where the off-axis residual is stable and low-dimensional, its structure affords mitigation.

Figures & tables

Explore similar work

May 11, 2026cs.SD

APEX: Audio Prototype EXplanations for Classification Tasks

Explainable AI (XAI) has achieved remarkable success in image classification, yet the audio domain lacks equally mature solutions. Current methods apply vision-based attribution techniques to spectrograms, overlooking fundamental differences between visual and acoustic signals. While prototype reasoning is promising, acoustic similarity remains multidimensional. We introduce APEX (Audio Prototype EXplanations), a post-hoc framework for interpreting pre-trained audio classifiers. Crucially, APEX requires no fine-tuning of the original backbone and strictly preserves output invariance. APEX disentangles explanations into four perspectives: Square-based prototypes to localize transient events, Time-based for temporal patterns, Frequency-based highlighting spectral bands, and Time-Frequency-based integrating both. This yields intuitive, example-based explanations that respect acoustic properties, providing greater semantic clarity than standard gradient-based methods.
Jun 12, 2026cs.SD

The Perceived Fragility of Explanations in Audio Models: Manipulation of Attribution with Unchanged Predictions

This paper investigates the fragility of post-hoc explanation methods in audio deepfake detection. While previous work on explanation manipulation focused on images using standard LpL_p metrics, we introduce a psychoacoustic framework that optimizes inaudible perturbations to decouple model attributions from final classifications. We evaluate this vulnerability across state-of-the-art architectures under strict prediction-preserving constraints. By evaluating the manipulation cost through domain-specific perceptual audio quality metrics alongside explanation alignment criteria, our framework demonstrates that an adversary can systematically distort automated explanation heatmaps while preserving the predicted deepfake label. Full code available at: https://github.com/cncPomper/Audio-XAI
Jun 12, 2026cs.SD

Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models

Transformer-based automatic speech recognition (ASR) models such as Whisper are highly accurate, but their predictions remain difficult to interpret. Existing explainable AI (XAI) methods often lack faithfulness and precise temporal grounding. We propose Listening with Entropy-guided Attention for Faithful explainability (LEAF-X), a model-intrinsic XAI framework for transformer-based ASR. LEAF-X combines entropy-guided attention weighting, multi-layer attention rollout, and optional causal ablations to identify low-entropy, high-impact heads and layers, producing sparse token-to-frame attributions. Unlike perturbation-based explainers or raw attention maps, LEAF-X exploits the internal structure of encoder-decoder and speech-augmented decoder-only models to generate explanations that better reflect model computation. Results show 32% improved faithfulness, 35-39% stronger locality/sparsity, and the most stable attributions, supporting more transparent and auditable ASR.