cs.SDSep 27, 2026

Mask-Induced Displacement in Audio XAI via Logit Trajectory Decomposition

Authors: Nico García-Peguinho, David Kelly, Fabrizio Smeraldi, Anna Xambó Sedó

Organizations: School of Electronic Engineering and Computer Science, Queen Mary University of London · Department of Informatics, King's College London · Department of Informatics, King’s College London

Abstract

Perturbation-based XAI methods for audio classifiers often estimate feature importance by masking spectrogram regions and crediting output changes to the retained signal. Yet they typically assume the fill (the mask replacement) is negligible. We propose logit-space trajectory decomposition to examine this assumption. An on-axis component captures output along a line connecting the fully filled (occluded) spectrogram to the fully retained original; an off-axis component captures perpendicular displacement. We evaluate across three fills, three audio classifiers, and 1,100 AudioSet clips. We demonstrate that no fill is acoustically neutral under full occlusion: Zero fill activates silence and Gaussian Noise fill activates broadband noise. Under partial masking, 41-77% of output displacement is off-axis, with the direction of the residual stable across mask retention fraction and specific to each model-fill combination. Attribution methods are unevenly exposed to off-axis displacement through their sampling and weighting strategies, revealing apparatus-dependence. Where the off-axis residual is stable and low-dimensional, its structure affords mitigation.

Figures & tables

Explore similar work

CardsList
  1. APEX: Audio Prototype EXplanations for Classification Tasks

    May 11, 2026Piotr Kawa, Kornel Howil, Piotr Borycki +3Audio UnderstandingAudio Editing

  2. The Perceived Fragility of Explanations in Audio Models: Manipulation of Attribution with Unchanged Predictions

    Jun 12, 2026Piotr Kitłowski, Dominik Wiącek, Mateusz ModrzejewskiAudio Deepfake DetectionExplainable AI Methods

  3. Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models

    Jun 12, 2026Ravi Ranjan, Utkarsh Grover, Xiaomin Lin +1End-To-End Automatic Speech Recognition ModelsNeural Audio