cs.LGJul 4, 2025

Detecting and explaining clinical-omics inconsistencies to improve patient cohort stratification: an application to Parkinson's disease

Authors: José A. Pardo-PérezTomás BernalJaime ÑiguezAna Luisa Gil-MartínezLaura IbañezJosé T. PalmaJuan A. BotíaAlicia Gómez-Pascual

Abstract

Discrepancies between clinical diagnoses and omics profiles within a characterized cohort may reflect misdiagnosis, hidden subgroups or prodromal disease states. We propose MLASDO, a tool to detect and characterize such discrepancies before downstream analyses. MLASDO (1) detects outliers using two unsupervised methods, thereby flagging potential poor-quality samples; and (2) identifies and characterizes anomalous samples (ASs), i.e., individuals whose molecular profile resembles the opposite clinical class. We applied MLASDO to the Parkinson's Progression Markers Initiative (PPMI) and Parkinson's Disease Biomarkers Program (PDBP) Parkinson's disease (PD) cohorts. In PPMI, it detected 26 outliers and 12 ASs: 5 anomalous healthy controls (AHCs) and 7 anomalous PD cases (APDs). AHCs exhibited higher cerebrospinal fluid (CSF) A\b{eta}1-42 levels than controls (P < 0.0396), suggesting resistance to cognitive decline. AHC-specific genes were enriched for the MAPK pathway (P<0.0145), implicated in PD pathogenesis. One AHC later received a different neurological disorder, three months after enrollment. In PDBP, it identified 10 outliers and 8 ASs: 2 AHCs and 6 APDs. One AHC exhibited 24 clinical features consistent with a PD-like phenotype, including severe motor impairment. Results are compiled in an interactive report for inspection and querying, highlighting clinically meaningful individuals otherwise overlooked in conventional analyses.

Explore similar work

Aug 10, 2026cs.LG

Label-Free Parkinson's Disease Screening from Face and Voice through Mechanistic Interpretability

Parkinson's disease (PD) is the second most common neurodegenerative disorder. Typical machine learning screening methods require PD labels, but the available data is limited by privacy concerns and the need for expert annotation. We propose a label-free face-plus-voice PD screen built entirely on frozen pretrained encoders--a face-expression Vision Transformer and HuBERT--in which no PD label touches any fit; the reference is training controls only. The voice modality uses a synthetic-dysarthria contrastive activation addition (CAA) direction built from time-stretch and breathy degradation of healthy speech; the face modality uses a k-nearest-neighbor anomaly score to the control embedding cluster. We introduce the alignment principle, a post-hoc analysis showing that a synthetic-degradation CAA detector works when the cosine similarity between the synthetic and real disease directions exceeds zero. Measured on the YouTubePD benchmark, this cosine is +0.37 for voice (CAA works, AUROC 0.765) and -0.48 for face (CAA fails; anomaly succeeds, AUROC 0.751). Equal-weight late fusion reaches AUROC 0.802 (95% CI [0.70,0.89]) with NPV 0.95, supporting a rule-out triage interpretation. An overfitting audit shows the voice detector transfers cleanly, while the face-side--and thus fused--AUROC is potentially optimistic pending external validation.
Jiaheng Su, Yu Sun
Apr 19, 2026cs.LG

STEP-PD: Stage-Aware and Explainable Parkinson's Disease Severity Classification Using Multimodal Clinical Assessments

Parkinson's disease (PD) is a progressive disorder in which symptom burden and functional impairment evolve over time, making severity staging essential for clinical monitoring and treatment planning. However, many computational studies emphasize binary PD detection and do not fully use repeated follow-up clinical assessments for stage-aware prediction. This study proposes STEP-PD, a severity-aware machine learning framework to classify PD severity using clinically interpretable boundaries. It leverages all available visits from the Parkinson's Progression Markers Initiative (PPMI) and integrates routinely collected subjective questionnaires and objective clinician-assessed measures. Disease severity is defined using Hoehn and Yahr staging and grouped into three clinically meaningful categories: Healthy, Mild PD (stages 1-2), and Moderate-to-Severe PD (stages 3-5). Three binary classification problems and a three-class severity task were evaluated using stratified cross-validation with imbalance-aware training. To enhance interpretability, SHAP was used to provide global explanations and local patient-level waterfall explanations. Across all tasks, XGBoost achieved the strongest and most stable performance, with accuracies of 95.48% (Healthy vs. Mild), 99.44% (Healthy vs. Moderate-to-Severe), and 96.78% (Mild vs. Moderate-to-Severe), and 94.14% accuracy with 0.8775 Macro-F1 for three-class severity classification. Explainability results highlight a shift from early motor features to progression-related axial and balance impairments. These findings show that multimodal clinical assessments within the PPMI cohort can support accurate and interpretable visit-level PD severity stratification.
Md Mezbahul Islam, John Michael Templeton, Christian Poellabauer +1
May 13, 2026eess.AS

A Benchmark for Early-stage Parkinson's Disease Detection from Speech

Early-stage Parkinson's disease (EarlyPD) detection from speech is clinically meaningful yet underexplored, and published results are hard to compare because studies differ in datasets, languages, tasks, evaluation protocols, and EarlyPD definitions. To address this issue, we propose the first benchmark for speech-based EarlyPD detection, with a speaker-independent split designed for fair and replicable cross-method evaluation on researcher-accessible datasets. The benchmark covers three common speech tasks and evaluates methods under different training-resource settings. We also present multi-dimensional evaluation breakdowns by dataset, aggregation level, gender, and disease stage to support fine-grained comparisons and clinical adoption. Our results provide a replicable reference and actionable insights, encouraging the adoption of this publicly available benchmark to advance robust and clinically meaningful EarlyPD detection from speech.
Terry Yi Zhong, Cristian Tejedor-Garcia, Khiet P. Truong +3