Towards clinical adoption of voice and speech as measures of health: the need for harmonization
Organizations: Institute of Psychiatry, Psychology and Neuroscience, King’s College London, UK · NIHR Biomedical Research Centre: Maudsley, UK · Modality.AI Inc, San Francisco, California, USA · University of California San Francisco, San Francisco, California, USA · Child Mind Institute, New York, NY, USA · Department of Psychology, Harvard University, Cambridge, USA · McGovern Institute Massachusetts Institute of Technology (MIT) Cambridge, USA · Department of Language Science and Technology, The Hong Kong Polytechnic University, Hong Kong SAR of China · Arizona State University, Tempe, AZ, USA · INESC-ID, Lisbon, Portugal · Sword Health, Portugal · Department of Precision Health, Luxembourg Institute of Health, Strassen, Luxembourg · Pattern Recognition Lab, Friedrich-Alexander-Universitat Erlangen-Nurnberg (FAU), Erlangen, Germany · Chair for AI in Healthcare and Medicine, Technical University of Munich (TUM), Munich, Germany · Lincoln Laboratory, Massachusetts Institute of Technology, Lexington, MA, USA
Abstract
Speech and voice are multidimensional signals that capture both communicative intent and underlying physiological processes, providing a unique, non-invasive window into health. Analyzing these signals has the potential to yield digital biomarkers that (i) provide scalable, objective measurement tools for research and clinical care and (ii) reflect the presence or progression of diverse conditions, including neurological, psychiatric, respiratory, and cardiovascular disorders. Realizing this promise, however, requires the field to overcome pervasive reproducibility and generalizability issues due to heterogeneous data collection, processing, and analysis practices. A major source of this heterogeneity is how underlying acoustic measures themselves are defined and computed. In this paper, we outline key considerations across the speech biomarker discovery lifecycle, from data collection through machine learning modeling to clinical interpretation, needed to achieve reliable, reproducible, and clinically translatable results. Chief among these is the need for harmonization efforts to start from common, precisely specified measure definitions. As a first step, we therefore provide definitions, physiological correlates, and computational implementations for a minimal, clinically interpretable set of core speech measures spanning respiration, phonation, articulation, and fluency. We close by discussing ongoing standardization efforts and the open challenges that remain in advancing the adoption of speech- and voice-based digital biomarkers.