Alzheimer's disease (AD) and mild cognitive impairment (MCI), which may precede AD, manifest early through subtle linguistic and acoustic alterations. Traditional diagnostics, however, are often resource-intensive and lack scalability for mass screening. To address these challenges, we introduce a novel bilingual speech large language model framework for automated, explainable cognitive screening. Unlike conventional pipelines that rely on error-prone automatic speech recognition, our system directly processes raw speech to learn joint acoustic-semantic representations, preserving critical prosodic cues often lost in transcription. Utilising our newly collected PUTH-AD dataset alongside multiple open-source corpora, we implemented a multi-task learning objective that simultaneously performs cognitive status classification and generates clinician-understandable natural language explanations. Our system achieved the highest average accuracy and AUROC across six dataset/task conditions, comparing three representative baselines. The system demonstrated cross-task transfer to held-out PUTH-AD task subsets, maintaining classification accuracy on an entirely unseen cognitive task without task-specific fine-tuning. Furthermore, clinician evaluation confirms that the generated explanations are both clinically relevant and largely consistent with the underlying speech evidence, supporting their potential utility in clinical interpretation. This study provides a scalable, objective, and explainable framework for speech-based cognitive screening, combining cognitive status classification with natural language explanations that clinicians can assess and verify, bridging the gap between advanced AI and clinical utility.
Figures & tables
Figure 1 : Overview of the proposed bilingual speech-based cognitive screening framework. (a) Dataset composition includes a newly collected PUTH-AD dataset, completing four distinct cognitive tasks, alongside established open-source datasets covering English and Chinese languages with varying task types. (b) The system pipeline consists of two components: Evidence text generation uses Gemini to generate explanations from speech recordings and cognitive status labels, serving as training targets; the proposed SpeechLLM-based system processes raw speech with task prompts to simultaneously output cognitive status classification and natural language explanations.
Figure 2 : Demographic and cognitive status characteristics of datasets. (a) Distribution of cognitive status categories (HC/MCI/AD) across ADReSS, NCMMSC-AD, and PUTH-AD datasets. (b) Age distribution by cognitive status categories for ADReSS and PUTH-AD datasets (NCMMSC-AD lacks age information). (c) Gender distribution by cognitive status categories across all three datasets. (d) Mini-Mental State Examination (MMSE) score distribution by cognitive status categories for the ADReSS dataset. (e-f) Montreal Cognitive Assessment (MoCA) and Hong Kong Brief Cognitive Test (HKBC) score distributions by cognitive status categories for the PUTH-AD dataset.
Figure 3 : Classification accuracy comparison across datasets and speech tasks. Bar charts show the performance of the SSL-Model baseline, the Gemini-few-shot baseline, the Text-only-LLM baseline, and our system on six separate datasets and speech tasks. Bars show mean accuracy across five runs with different random seeds, with error bar showing standard deviation. Asterisks indicate statistical comparisons between each baseline and Ours-Kimi-Audio. Bars marked “NS” denote non-significant differences. ∗p<0.05;∗∗p<0.01;∗∗∗p<0.001 indicate statistical significance.
Dataset
Speech Task
Accuracy
AUROC
ADReSS
-
0.893 [0.825, 0.952]
0.921 [0.849, 0.980]
NCMMSC-AD
-
0.857 [0.807, 0.904]
0.943 [0.906, 0.973]
PUTH-AD
AFT
0.553 [0.427, 0.677]
0.675 [0.587, 0.763]
FPD
0.507 [0.404, 0.612]
0.673 [0.584, 0.763]
CTD
0.553 [0.442, 0.662]
0.681 [0.572, 0.792]
PR
0.638 [0.523, 0.750]
0.770 [0.674, 0.875]
Table 1 : Classification performance of the Kimi-Audio model on each dataset/task. Accuracy and AUROC across five random seeds are reported in formant of mean [95% CIs].
Figure 4 : Ablation study comparing classification-only, generation-only, and multi-task training for SALMONN and Kimi-Audio. Panels show accuracy, AUROC, BERT-Score, and hit rate (HR), with markers indicating mean performance across five random seeds. Squares, triangles, and circles denote classification-only, generation-only, and multi-task training, respectively; lines connect training variants within each backbone. Generation-only accuracy is derived from the generated explanations. AUROC is reported for classification-only and multi-task, which have continuous class scores, and generation metrics for generation-only and multi-task which produce explanations. ADR and NCM denote ADReSS and NCMMSC-AD; AFT, FPD, CTD, and PR denote the four PUTH-AD tasks.
Figure 5 : Comparison of separate and joint training with the Kimi-Audio backbone. Blue and red markers show mean performance across five random seeds under separate and joint training, respectively. Joint training includes ADReSS (ADR), NCMMSC-AD (NCM), and the PUTH-AD Personal Recall (PR) subset. The dashed line separates these data from the PUTH-AD AFT, FPD, and CTD subsets, which are excluded from joint training and evaluated on held-out test subjects. Lines connect the two training settings within each condition. Voting denotes participant-level hard voting across the four PUTH-AD tasks, evaluated with accuracy.
Figure 6 : Clinical evaluation for classification and explanation outputs by the joint-training Kimi-Audio model using the predefined random seed of 42. (a) Class-wise F1, sensitivity, and specificity computed in a one-vs-rest manner, after pooling test predictions from ADReSS, NCMMSC-AD, and PUTH-AD voting result. (b) Agreement between original Gemini-generated explanations and clinician-corrected Gemini explanations, measured by BERT-Score and hit rate. (c) Clinician scores for Gemini-generated explanation in two dimensions, clinical relevance and evidence faithfulness. (d) Clinician scores for Kimi-Audio explanations, shown overall and separately for correctly and incorrectly classified samples. Significance was assessed by Welch’s two-sample t-test between the correct and wrong groups. ∗∗∗p<0.001 ; n.s., not significant.
Figure 7 : Model Architecture. Dual output heads simultaneously do classification and generation. The model is optimised through both losses.
Supplementary Figure 1 : Two representative successful examples of generated explanations on the AFT task and the FPD task. Each example includes the ASR transcription of the participant’s speech, the system’s generated explanation output, and the ground truth category. The original transcription and output are in Chinese, and translated to English here.
Supplementary Figure 2 : Two representative failure cases of generated explanations on the CTD task and PR task. Each example includes the ASR transcription of the participant’s speech, the system’s generated explanation output, and the ground truth category. The original transcription and output are in Chinese, and have been translated to English here.
Supplementary Figure 3 : Prompt templated for generating explanation texts using Gemini-2.5-Flash. The generation process employs a two-stage pipeline to ensure thorough analysis and concise output: Stage 1 guides detailed step-by-step analysis of speech characteristics, while Stage 2 extracts the most salient points from the detailed explanation. The prompt is original in Chinese and translated into English here.
Supplementary Figure 4 : Prompt templated for Gemini-few-shot baseline, with NCMMSC-AD shown as an example. The prompt contains the same general clinical and task descriptions as the explanation generation prompt, but is adapted for direct cognitive status classification. Six in-context examples are provided, including two examples from each category (HC, MCI, and AD). The prompt is original in Chinese and translated into English here.
Supplementary Figure 5 : Web interface for clinician evaluation of generated explanations. The left panel presents the audio player, audio path, transcription, task instruction, and task image when applicable. The right panel presents the original Gemini-generated explanation, scoring fields for the Gemini explanation, an editable text box for clinician correction, the Kimi-Audio-generated explanation, and scoring fields for the Kimi-Audio explanation. The interface was implemented in Chinese and deployed on a local server. A default value of 0 indicates that a rating has not yet been assigned; valid clinician ratings range from 1 to 5.
Dataset
Speech Task
Separate Training
Unseen
Acc
AUROC
BERT
HR
ADReSS
-
-
0.893 ± 0.016
0.921 ± 0.013
0.772 ± 0.001
0.697 ± 0.037
NCMMSC-AD
-
-
0.857 ± 0.023
0.943 ± 0.004
0.727 ± 0.007
0.580 ± 0.032
PUTH-AD
AFT
-
0.553 ± 0.016
0.675 ± 0.031
0.731 ± 0.012
0.598 ± 0.053
FPD
-
0.507 ± 0.029
0.673 ± 0.011
0.743 ± 0.009
0.570 ± 0.043
CTD
-
0.553 ± 0.041
0.681 ± 0.032
0.752 ± 0.004
0.569 ± 0.022
Supplementary Table 3 : Comparison of separate training versus cross-domain joint training with Kimi-Audio backbone. Joint training data includes ADReSS, NCMMSC-AD, and PUTH-AD Personal Recall data. “Unseen” denotes PUTH-AD task subsets excluded from joint training, which are evaluated on the held-out PUTH-AD test subjects for generalisation assessment. Voting represents the result of hard voting across all PUTH-AD tasks. Results reported as mean ± standard deviation across the five different seeds.
Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for enabling timely intervention and improving patient outcomes. Speech-based CI detection has emerged as a promising non-invasive approach, as speech signals encode both linguistic and acoustic markers associated with cognitive decline. Recent advances in large language models (LLMs) further strengthen the potential of speech-based assessment by enabling more expressive representation learning and improved generalization across diverse speakers, recording devices, and clinical environments. Moreover, multimodal learning by jointly modeling linguistic and acoustic features allows for a more comprehensive characterization of cognitive and behavioral changes related to CI, leading to more reliable detection. In this work, we propose a multimodal CI detection framework based on open-source LLMs that integrates speech audio and corresponding transcripts while preserving patient privacy. Acoustic embeddings are extracted directly from speech signals, while textual embeddings are generated from automatically transcribed speech. These modality-specific embeddings are then concatenated to create a combined feature vector and used for downstream classification, without requiring access to raw or sensitive patient data. The proposed approach is evaluated on the ADReSS20 and ADReSSo21 benchmark datasets. Experimental results show that the proposed multimodal framework achieves an CI classification accuracy of 92.4% and consistently outperforms single-modality baselines. Our work establishes a new state-of-the-art for CI identification, with the proposed method demonstrating superior cross-dataset generalization. This advance highlights the power of an LLM-based multimodal framework that fuses linguistic and acoustic data to enable robust, scalable, and non-invasive screening.
Yingchao Huang, Xin Wang, Yuhan Su +1
Faculty of Digital Innovation, Arts & Sciences, Saskatchewan Polytechnic, Regina SK S4S 5X1, Canada · School of Basic Medical Sciences, Hebei University, Baoding 071000, China · Department of Civil & Environmental Engineering and School of Mining & Petroleum Engineering, University of Alberta, Edmonton AB T6G 2H5, Canada
Speech-based cognitive impairment detection offers a noninvasive, accessible alternative to costly biomarker assays, yet transformer-based models remain clinically uninterpretable. We propose a multi-stage explainability framework that translates black-box transformer predictions into clinically grounded narratives by integrating SHapley Additive exPlanations (SHAP)-based token attribution, theory-informed linguistic features, and a four-stage LLM reasoning pipeline using LLaMA-3.1-70B-Instruct. Built on the SpeechCARE-Adaptive Gating Network multimodal screening model (F1 = 72.11% on the NIA PREPARE benchmark), the framework maps model outputs to four cognitive-linguistic dimensions, including lexical richness, syntactic complexity, and semantic coherence. Physician evaluation on 70 stratified English samples demonstrated strong alignment with patient-level cognitive profiles, and a System Usability Scale score of 82/100 indicated high potential for clinical workflow integration.
Yasaman Haghbin, Sina Rashidi, Ali Zolnour +6
Independent Researcher · Columbia University, United States · Chalmers University of Technology, Sweden
Early diagnosis of Alzheimer's disease (AD) is critical for enabling timely interventions that may slow disease progression and improve patient outcomes. There is a growing need for AD detection methods that are non-invasive and cost-effective, especially in real-world clinical settings with diverse patient populations and recording conditions. Speech-based screening addresses these needs by using natural speech collected without specialized equipment. Recent advances in large language models (LLMs) have improved speech analysis by providing rich linguistic representations and strong generalization. In this study, we propose LSEAD, a speech-based AD detection framework using pretrained open-source LLMs. Speech recordings are automatically transcribed, and text embeddings are extracted using locally deployed LLMs. Principal component analysis (PCA) is applied to reduce dimensionality before classification. Because the framework relies only on speech transcripts and locally deployed models, it supports privacy-preserving AD risk assessment without external data exchange. We evaluate LSEAD on the ADReSS20 and ADReSSo2021 benchmark datasets. Experimental results show that LLM-based embeddings generalize well across datasets and improve AD classification accuracy by up to 5 percent over existing methods, especially for early-stage detection. These results demonstrate that LSEAD provides a practical, secure, and scalable approach for early AD screening.
Xin Wang, Yingchao Huang, Yuhan Su +2
Faculty of Digital Innovation, Arts & Sciences, Saskatchewan Polytechnic, Regina SK S4S 5X1, Canada · School of Basic Medical Sciences, Hebei University, Baoding 071000, China · Department of Civil & Environmental Engineering and School of Mining & Petroleum Engineering, University of Alberta, Edmonton AB T6G 2H5, Canada +1