Mask-Induced Displacement in Audio XAI via Logit Trajectory Decomposition
Authors: Nico García-Peguinho, David Kelly, Fabrizio Smeraldi, Anna Xambó Sedó
Organizations: School of Electronic Engineering and Computer Science, Queen Mary University of London · Department of Informatics, King's College London · Department of Informatics, King’s College London
Perturbation-based XAI methods for audio classifiers often estimate feature importance by masking spectrogram regions and crediting output changes to the retained signal. Yet they typically assume the fill (the mask replacement) is negligible. We propose logit-space trajectory decomposition to examine this assumption. An on-axis component captures output along a line connecting the fully filled (occluded) spectrogram to the fully retained original; an off-axis component captures perpendicular displacement. We evaluate across three fills, three audio classifiers, and 1,100 AudioSet clips. We demonstrate that no fill is acoustically neutral under full occlusion: Zero fill activates silence and Gaussian Noise fill activates broadband noise. Under partial masking, 41-77% of output displacement is off-axis, with the direction of the residual stable across mask retention fraction and specific to each model-fill combination. Attribution methods are unevenly exposed to off-axis displacement through their sampling and weighting strategies, revealing apparatus-dependence. Where the off-axis residual is stable and low-dimensional, its structure affords mitigation.
Figures & tables
Figure 1: Method overview. Four spectrograms (original and three fills) are passed through classifier F , producing logit vectors v∈R527 . Each vα denotes the logit vector at retention fraction α . The logit-space trajectory is decomposed into a scalar on-axis component τ (along the attribution axis d=vorig−vfoc ) and a scalar off-axis magnitude d⊥ , whose direction is captured by the residual vector v⊥∈R527 .
Model
Fill
Top-1 class
qˉ
σ
n
Top-2 class
qˉ
σ
n
ast+sa
zero
Silence
.608
–
1100
Music
.269
–
1100
mean
Mains hum
.691
.220
857
Hum
.598
.236
864
gaussian noise
White noise
.466
.145
969
Noise
.176
.092
350
panns-sa
zero
Silence
.327
–
1100
Music
.186
–
1100
mean
Buzzer
.216
.123
273
Hum
.173
.129
178
gaussian noise
White noise
.342
.136
618
Static
.234
.121
545
Table 1: Top-2 predicted classes per fill at FOC ( N=1,100 clips, 22 classes). qˉ : mean sigmoid confidence. σ omitted for zero (signal-independent).
ast+sa
panns-sa
panns+sa
Frying food
Mains hum † – .83
Hum † – .07
Siren § – .13
Drum kit
Mains hum † – .78
Buzzer † – .06
Sound fx – .18
Owl
Sine wave § – .39
Hum † – .09
Sine wave § – .49
Heartbeat
Sine wave § – .37
Music – .12
Sine wave § – .47
Whispering
Mains hum † – .24
Hum † – .09
Sine wave § – .41
Bagpipes
Mains hum † – .14
Buzzer † – .17
Siren § – .14
Table 2: Modal top-1 class and mean qˉ per source class under mean at FOC. Colours indicate semantic group: † electrical interference (Mains hum, Hum, Buzzer), § tonal periodic (Sine wave, Siren).
Figure 2: Logit-space trajectory analysis across 95.04M partial masks (22 classes, 50 clips, 9,600 masks/clip). Fill: Z = zero, M = mean, G = Gaussian. Row 1: mask distribution and mean on-axis projection τ by α . Solid line = full logit vector; dotted = GT class ( τgt ); dashed = ideal ( τ=α ). Row 2: per-fill summary table with σˉτ , σˉd⊥ , d⊥/dtotal ; mean off-axis residual d⊥ by α .
Model
Fill
PR
cˉ
mincb,b′
ρˉ (%)
ast+sa
Z
19.14
.95
.74
18.5
M
13.37
.94
.76
23.5
G
15.13
.87
.25
20.0
panns-sa
Z
8.39
.95
.68
32.1
M
19.56
.86
.20
15.3
G
16.03
.90
.52
18.5
Table 3: Off-axis residual structure per model–fill condition. PR: participation ratio (effective dimensionality). cˉ : mean pairwise eigenvector cosine similarity. mincb,b′ : worst-case similarity. ρˉ : variance fraction explained by leading eigenvector. Fill: Z = zero , M = mean , G = gaussian noise .
Explainable AI (XAI) has achieved remarkable success in image classification, yet the audio domain lacks equally mature solutions. Current methods apply vision-based attribution techniques to spectrograms, overlooking fundamental differences between visual and acoustic signals. While prototype reasoning is promising, acoustic similarity remains multidimensional. We introduce APEX (Audio Prototype EXplanations), a post-hoc framework for interpreting pre-trained audio classifiers. Crucially, APEX requires no fine-tuning of the original backbone and strictly preserves output invariance. APEX disentangles explanations into four perspectives: Square-based prototypes to localize transient events, Time-based for temporal patterns, Frequency-based highlighting spectral bands, and Time-Frequency-based integrating both. This yields intuitive, example-based explanations that respect acoustic properties, providing greater semantic clarity than standard gradient-based methods.
Piotr Kawa, Kornel Howil, Piotr Borycki +3
Department of Artificial Intelligence, Wroclaw University of Science and Technology, Poland · IDEAS Research Institute, Poland · Faculty of Mathematics and Computer Science, Jagiellonian University, Poland +1
This paper investigates the fragility of post-hoc explanation methods in audio deepfake detection. While previous work on explanation manipulation focused on images using standard Lp metrics, we introduce a psychoacoustic framework that optimizes inaudible perturbations to decouple model attributions from final classifications. We evaluate this vulnerability across state-of-the-art architectures under strict prediction-preserving constraints. By evaluating the manipulation cost through domain-specific perceptual audio quality metrics alongside explanation alignment criteria, our framework demonstrates that an adversary can systematically distort automated explanation heatmaps while preserving the predicted deepfake label. Full code available at: https://github.com/cncPomper/Audio-XAI
Piotr Kitłowski, Dominik Wiącek, Mateusz Modrzejewski
Faculty of Electronics and Information Technology, Warsaw University of Technology, Warsaw, Poland.
Transformer-based automatic speech recognition (ASR) models such as Whisper are highly accurate, but their predictions remain difficult to interpret. Existing explainable AI (XAI) methods often lack faithfulness and precise temporal grounding. We propose Listening with Entropy-guided Attention for Faithful explainability (LEAF-X), a model-intrinsic XAI framework for transformer-based ASR. LEAF-X combines entropy-guided attention weighting, multi-layer attention rollout, and optional causal ablations to identify low-entropy, high-impact heads and layers, producing sparse token-to-frame attributions. Unlike perturbation-based explainers or raw attention maps, LEAF-X exploits the internal structure of encoder-decoder and speech-augmented decoder-only models to generate explanations that better reflect model computation. Results show 32% improved faithfulness, 35-39% stronger locality/sparsity, and the most stable attributions, supporting more transparent and auditable ASR.
Ravi Ranjan, Utkarsh Grover, Xiaomin Lin +1
Florida International University · Miami, USA · University of South Florida +1