Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?
Authors: Serli Kopar, Alkis Koudounas, Roshan P. Rane, Sam Gijsen, Paula A. Perez-Toro, Kerstin Ritter
Organizations: Hertie Institute for AI in Brain Health, Germany · Tübingen AI Center, Germany · Sony Group Corporation, Japan · Friedrich-Alexander-University Erlangen-Nürnberg, Germany
Speech-based Alzheimer's disease (AD) assessments increasingly rely on pretrained self-supervised learning (SSL) models that learn acoustic representations directly from raw audio, exposing the model to recording factors. We ask whether such factors are merely encoded in SSL representations or can systematically alter predictions. Using ADReSSo and three large SSL backbones, we apply controlled noise and reverberation interventions to participant-speech-only, non-speech, and full-recording audio. We combine layer-wise linear decoding, input- and representation-space interventions, and geometric alignment analysis to distinguish acoustic decodability from influence on AD prediction. Our results show that controlled acoustic interventions alter AD predictions across all three SSL backbones. Noise, despite showing no significant diagnostic-group difference in the original data, produces the strongest intervention effects. Importantly, these effects are systematically structured relative to the classifier's decision direction, replicate on the held-out test set and reverse when the representation-space intervention direction is reversed. Together, these findings show that high predictive performance and the absence of a significant diagnostic-group difference in a measured acoustic factor are not sufficient for robustness. We argue that intervention-based robustness tests should become standard for trustworthy clinical speech models.
Figures & tables
Train
Test
Feature
HC
AD
p
HC
AD
p
N
79
87
–
36
35
–
MMSE
29.0 ± 1.1
17.4 ± 5.3
< .001
28.9 ± 1.3
18.9 ± 5.8
< .001
SNR
76.2 ± 33.5
78.5 ± 31.9
.489
76.8 ± 33.6
86.9 ± 29.6
.073
SRMR
5.90 ± 2.65
8.10 ± 3.88
< .001
6.88 ± 4.53
5.77 ± 3.55
.216
Table 1: MMSE and recording characteristics for train and test splits
Figure 1: E1, linear decoding. Layer-wise SNR, SRMR, and MMSE R2 and AD BA across streams.
Figure 2: E2&E4, input-space intervention&alignment. (a) AD prediction flip rates across layers for SNR and SRMR interventions; colors denote streams, gradients degradation severity D1–D4, and the dashed line marks ℓAD∗=12 . (b) HC → AD and AD → HC flips at ℓAD∗=12 , averaged across levels. (c) Cosine alignment with the AD decision direction at ℓAD∗=12 ; gray regions show 95% random-direction intervals.
Figure 3: E3&E4, representation-space intervention&alignment. Same panel structure as Fig. 2 ; panels b–c show results at r∗=9 .
Speech-based Alzheimer's disease (AD) detection increasingly relies on speech-enhanced and curated versions of the Pitt Corpus, where speech enhancement, sample selection, and demographic balancing are often treated as beneficial preprocessing steps. However, whether these transformations improve real-world AD detection or instead affect model generalization and prediction behavior remains unclear. In this work, we revisit the role of speech preprocessing and dataset curation across widely used benchmarks for speech-based AD detection. We evaluate the speech quality of different datasets, the cross-dataset generalization of multiple deep learning models under matched and mismatched enhancement settings, and the behavior of several recent large audio-language models (LALMs). Experimental results show that across multiple supervised speech models, speech-enhanced datasets often improve in-domain performance while reducing robustness in cross-domain evaluation. Matched enhancement between training and test data alleviates, but does not eliminate, this degradation. LALMs show a similar sensitivity: enhanced datasets induce stronger class imbalance and prediction shifts than unprocessed data. These results suggest that speech preprocessing and dataset curation can substantially influence downstream AD detection behavior, indicating that ``cleaner'' speech datasets are not necessarily more reliable for real-world AD detection.
Luqi Sun, Shreeram Suresh Chandra, Lin Zhang +5
Center for Language and Speech Processing (CLSP), Johns Hopkins University · Research Center for Information Technology Innovation, Academia Sinica · University of Michigan, Ann Arbor +1
Acoustic biomarkers show promise for detecting Alzheimer's Disease (AD), yet whether the cues driving diagnostic AI align with those salient to human listeners is underexplored across languages and genders, where pathological markers and perceptual strategies differ. We train models to predict clinical AD status (pathology) and human perceptual scores across Mandarin and Greek, male and female speakers. Using SHAP for interpretability and statistical models for validation, we compare feature importance by subgroup. Results reveal a context-dependent divergence: pathological-perceptual alignment is significant for Mandarin and female speakers but disappears for Greek and male speakers, where pathology models did not exceed chance; this is a failure mode that population-specific auditing surfaces. Global Explainable AI (XAI) explanations can mask critical demographic divergences, highlighting the need for population-specific explainability auditing for equitable deployment of clinical speech AI.
Liu He, Yuanchao Li, Yin-Long Liu +5
University of Science and Technology of China · University of Edinburgh
Audio language models (ALMs) are increasingly used for speech-based understanding, yet their ability to perform semantic reasoning beyond transcription, Text-to-Audio Retrieval, Captioning, and Question-Answering accuracy remains insufficiently benchmarked. In particular, the effects of accent variation, domain shift, and semantic over-inference on audio reasoning are poorly understood. We evaluate audio language models across five semantic and paralinguistic reasoning tasks: entailment, consistency, plausibility, accent drift, and accent restraint. Collectively, these tasks assess a model's ability to reason over spoken audio as the primary evidence source, including whether a textual hypothesis can be inferred, contradicted, or left undetermined by the audio, whether statements align or conflict with spoken content, whether claims are plausible given the discourse, and whether model predictions remain stable or appropriately constrained across accent variation. These findings highlight critical limitations in current audio reasoning evaluations and hope to provide guidance for more robust and equitable ALM design and assessment