cs.SDMar 9, 2026

PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment

Authors: Bence Mark HalpernThomas TienkampDefne AburTomoki Toda

Organizations: 1Nagoya University, Japan · Department of Neurology, Faculty of Medicine and University Hospital Cologne, University of Cologne, Germany · University of Groningen, The Netherlands

Abstract

Automatic speech intelligibility assessment is crucial for monitoring speech disorders and therapy efficacy. However, existing methods are difficult to compare: research is fragmented across private datasets with inconsistent protocols. We introduce PathBench, a unified benchmark for pathological speech assessment using public datasets. We compare reference-free, reference-text, and reference-audio methods across three protocols (Matched Content, Extended, and Full) representing how a linguist (controlled stimuli) versus machine learning specialist (maximum data) would approach the same data. We establish benchmark baselines across six datasets, enabling systematic evaluation of future methodological advances, and introduce Dual-ASR Articulatory Precision (DArtP), achieving the highest average correlation among reference-free methods.

Explore similar work

Jun 23, 2026cs.SD

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge

Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches predominantly focus on holistic naturalness, leaving fine-grained paralinguistic distinctions underexplored. We introduce ParaPairAudioBench, a pairwise benchmark of 5,175 audio pairs across five paralinguistic dimensions: Style, Rate, Emphasis, Age, and Gender. Our experiments show that current LALM judges still lag behind human judgments by 32%p on average and exhibit severe calibration failures, particularly in Tie cases where the correct decision is to abstain. To further analyze lexical versus acoustic reliance, the benchmark includes both same-transcript and cross-transcript conditions. ParaPairAudioBench enables multi-dimensional, calibration-aware assessment of the reliability of LALM-as-a-Judge for paralinguistic speech evaluation.
Jisu Jeon, Seungyeon Jwa, Joosung Lee +6
Jun 15, 2026cs.AI

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI

Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor, respiratory, and vocal systems. Current clinical speech AI methods have largely progressed through isolated condition-specific studies, making results difficult to compare and generalization difficult to assess. We introduce SpeechDx, a large-scale benchmark for clinical speech AI spanning 12 datasets and 27 tasks across diverse health conditions. To enable evaluation across shared clinical mechanisms, SpeechDx structures tasks by the stage of speech production they disrupt: conceptualization, formulation, and articulation. The benchmark tests generalization by including tasks with limited labeled data and evaluating the same health condition across multiple datasets, distinguishing clinically meaningful patterns from dataset artefacts. We systematically evaluate 12 state-of-the-art audio encoders across all tasks and under zero-shot cross-condition transfer. Results show that large-scale speech models represent the strongest overall baselines, domain-specific models improve performance only on closely matched tasks, and no current representation generalizes reliably across the clinical speech landscape. SpeechDx establishes a shared evaluation framework for tracking progress toward general-purpose clinical speech representations
Sejal Bhalla, Larry Kieu, Aina Merchant +2
Sep 21, 2026cs.SD

ART-NAD: An Articulatory Inversion-based Neural Acoustic Distance for Pathological Speech Intelligibility Assessment

Speech assessment tools for speakers with speech pathology must be both accurate and interpretable if they are to be adopted in clinical practice. Existing reference-audio measures such as the Neural Acoustic Distance (NAD) reach high speaker-level correlations with listener intelligibility scores but operate on self-supervised features that are hard to interpret, providing only frame-level explanations. We propose ART-NAD, a reference-audio intelligibility metric that replaces the \texttt{wav2vec2} features of NAD with vocal-tract constriction variables (tract variables, TVs) predicted from audio by a speaker-independent acoustic-to-articulatory inversion model trained on the same \texttt{wav2vec2} features. ART-NAD is computed as the multivariate Dynamic Time Warping distance between the nine-channel quasi-TV trajectories of the test and one or more references. Across 20 reference-audio protocols spanning six pathological-speech datasets and five languages, ART-NAD with silence trimming (ART-NAD-FA) reaches the same average speaker-level Pearson correlation as NAD-FA (both r=0.71r=0.71) on the same self-supervised backbone, with no significant per-protocol difference (Wilcoxon p=0.18p=0.18), and is the strongest reference-audio metric on 6 of the 20 protocols. Beside the score itself, each TV channel visualizes which constriction deviates from the reference over time, providing interpretable information as to where articulation breaks down.
Bence Mark Halpern, Thomas Tienkamp, Defne Abur +1