MemoCare: An Interactive Multimodal Mobile System for Automated Cognitive Screening
Authors: Duy-Cat Can, Mau Minh Phuc Le, Tuan-Khoa Hoang, Hai-Dang Nguyen, Trung-Hieu Do, Dang Minh Ly, Minh-Duc Nguyen, Nghia TT Hoang, +5 more
Organizations: Lausanne University Hospital, Switzerland · Faculty of Biology and Medicine, University of Lausanne, Switzerland · VNU University of Engineering and Technology, Vietnam · VinUni-Illinois Smart Health Center, VinUniversity, Hanoi, Vietnam · University of Science, VNU-HCM, Vietnam · Hanoi Medical University, Vietnam · National Geriatric Hospital, Vietnam · Department of Neurology, Military Hospital 175, Vietnam · Department of Neurology, School of Medicine, University of Medicine and Pharmacy at Ho Chi Minh City, Vietnam · International University, VNU-HCM, Vietnam · Vietnam National University Ho Chi Minh City, Vietnam
MemoCare is an interactive mobile system for automated multimodal cognitive screening. A React Native application combines spoken responses, temporal and spatial orientation, touchscreen actions, and visuoconstruction in complete English and Vietnamese workflows. Speech is transcribed by Google Speech-to-Text and scored locally with deterministic task-specific natural language processing rules; GPS coordinates are resolved by the MemoCare spatial module before answer matching; touch tasks are scored from interaction events; and the drawing task uses a three-model convolutional neural network consensus with separate visual interpretation. Software tests pass 151/151 predefined cases across speech/language, spatial-answer, and touch-interaction scoring, while spatial regression passes 48/48 four-country coordinate-resolution cases. For the drawing module, validation-selected ShuffleNetV2 x1.5 achieved 91.33% mean balanced accuracy and 78.87% exact three-criterion accuracy on a locked 71-image test set. Four clinician co-authors additionally inspected the end-to-end workflow, yielding a pooled median rating of 4/5 across eight criteria, with item-level medians ranging from 3 to 4.5. At MMM, attendees can directly try a shortened multimodal screening workflow and inspect automatic item-level and total scoring.
Figures & tables
Figure 1: MemoCare architecture. (A) Speech, location, touch, and drawing are captured in one mobile app. (B) Specialized modules perform language scoring, geospatial reasoning, touch verification, and vision inference. (C) Item scores are aggregated into total score and local history.
Module
Evaluation
Result
Speech/language
Transcript-level scoring
138/138 passed
Spatial
Fixed-coordinate and answer scoring
48/48 and 8/8 passed
Touch
Interaction scoring
5/5 passed
Drawing, single model
ShuffleNetV2 x1.5, locked test ( n=71 )
Mean BA 91.33%; three-criterion exact 78.87%
Drawing, consensus
Locked test ( n=71 ), descriptive
Mean BA 95.17%
Clinician inspection
Eight 5-point criteria, n=4
Median 4/5 (item medians 3–4.5)
Table 1: Technical verification and expert inspection.
Alzheimer's disease (AD) and mild cognitive impairment (MCI), which may precede AD, manifest early through subtle linguistic and acoustic alterations. Traditional diagnostics, however, are often resource-intensive and lack scalability for mass screening. To address these challenges, we introduce a novel bilingual speech large language model framework for automated, explainable cognitive screening. Unlike conventional pipelines that rely on error-prone automatic speech recognition, our system directly processes raw speech to learn joint acoustic-semantic representations, preserving critical prosodic cues often lost in transcription. Utilising our newly collected PUTH-AD dataset alongside multiple open-source corpora, we implemented a multi-task learning objective that simultaneously performs cognitive status classification and generates clinician-understandable natural language explanations. Our system achieved the highest average accuracy and AUROC across six dataset/task conditions, comparing three representative baselines. The system demonstrated cross-task transfer to held-out PUTH-AD task subsets, maintaining classification accuracy on an entirely unseen cognitive task without task-specific fine-tuning. Furthermore, clinician evaluation confirms that the generated explanations are both clinically relevant and largely consistent with the underlying speech evidence, supporting their potential utility in clinical interpretation. This study provides a scalable, objective, and explainable framework for speech-based cognitive screening, combining cognitive status classification with natural language explanations that clinicians can assess and verify, bridging the gap between advanced AI and clinical utility.
Language agents are increasingly deployed over accumulating multimodal information, yet existing benchmarks assume a human-human form with sparse visuals and straightforward content, evaluating neither reasoning over authentic multimodal file interaction nor the interpretation of concealed user information. We therefore introduce M3Exam, a query-centric multimodal conversational memory benchmark built on realistic user-agent interaction, with multi-dimensional evaluation spanning cross-modal grounding and implicit information inference. Benchmarking MLLMs and memory systems reveals persistent gaps in cross-modal grounding, cross session reasoning, and the efficiency cost of accumulating multimodal context. We further propose M3Proctor, a multimodal memory method that detects query modality bias and consumes raw visual sources only on demand, improving accuracy by 13% while cutting index-construction time and retrieved tokens by over 70%.
Zhengjun Huang, Wenxuan Liu, Zhoujin Tian +6
The Hong Kong University of Science and Technology · Peng Cheng Laboratory · Beijing University of Chemical Technology +4
Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for enabling timely intervention and improving patient outcomes. Speech-based CI detection has emerged as a promising non-invasive approach, as speech signals encode both linguistic and acoustic markers associated with cognitive decline. Recent advances in large language models (LLMs) further strengthen the potential of speech-based assessment by enabling more expressive representation learning and improved generalization across diverse speakers, recording devices, and clinical environments. Moreover, multimodal learning by jointly modeling linguistic and acoustic features allows for a more comprehensive characterization of cognitive and behavioral changes related to CI, leading to more reliable detection. In this work, we propose a multimodal CI detection framework based on open-source LLMs that integrates speech audio and corresponding transcripts while preserving patient privacy. Acoustic embeddings are extracted directly from speech signals, while textual embeddings are generated from automatically transcribed speech. These modality-specific embeddings are then concatenated to create a combined feature vector and used for downstream classification, without requiring access to raw or sensitive patient data. The proposed approach is evaluated on the ADReSS20 and ADReSSo21 benchmark datasets. Experimental results show that the proposed multimodal framework achieves an CI classification accuracy of 92.4% and consistently outperforms single-modality baselines. Our work establishes a new state-of-the-art for CI identification, with the proposed method demonstrating superior cross-dataset generalization. This advance highlights the power of an LLM-based multimodal framework that fuses linguistic and acoustic data to enable robust, scalable, and non-invasive screening.
Yingchao Huang, Xin Wang, Yuhan Su +1
Faculty of Digital Innovation, Arts & Sciences, Saskatchewan Polytechnic, Regina SK S4S 5X1, Canada · School of Basic Medical Sciences, Hebei University, Baoding 071000, China · Department of Civil & Environmental Engineering and School of Mining & Petroleum Engineering, University of Alberta, Edmonton AB T6G 2H5, Canada