Organizations: Manning College of Information and Computer Sciences, UMass Amherst, Amherst, MA, USA · Center for Healthcare Organization and Implementation Research, VA Bedford Health Care, Bedford, MA, USA · Miner School of Computer and Information Sciences, UMass Lowell, Lowell, MA, USA
Opioid overdose remains a major clinical and public health burden, highlighting the need for scalable approaches to identify patients at high risk. Here, we investigate diagnosis-specific adaptation for 180-day opioid overdose risk prediction from patients' preceding one-year longitudinal ICD histories. We develop OODMAMBA and OODQWEN through continued pretraining on longitudinal diagnostic sequences followed by task-specific fine-tuning. Building on the stronger Qwen-based predictors, we further propose OVERDOSEMOE, a multi-expert framework that integrates models of different scales using complementary expert-weighting strategies. Diagnosis-specific adaptation consistently improved predictive performance over general-purpose language-model baselines, with OODQWEN achieving an AUPRC of 24.47 and an AUROC of 68.56. OVERDOSEMOE further improved discrimination and precision, achieving an AUPRC of 25.17 and an AUROC of 69.49 while outperforming the strongest single-model baselines. Among patients ranked in the top 5% of predicted risk, OVERDOSEMOE identified substantially enriched overdose risk, achieving a PPV of 25.38% while retaining meaningful recall. Evaluation on an independent MIMIC-IV cohort further demonstrated cross-cohort robustness, with complementary weighting strategies showing advantages across different performance measures. These findings demonstrate that diagnosis-specific language-model adaptation combined with multi-expert integration can improve opioid overdose risk stratification and support more robust prediction across heterogeneous electronic health record populations.
Figures & tables
Data source
Split
N
Overdose, n (%)
VHA
Training
62,928
3,724 (5.92%)
VHA
Validation
7,866
440 (5.59%)
VHA
Test
7,866
473 (6.01%)
MIMIC-IV
External
3,971
55 (1.38%)
Table 1: Cohort size and 180-day opioid-overdose prevalence in the VHA development cohort and the external MIMIC-IV cohort. The VHA cohort was partitioned at the patient level into training, validation, and held-out test sets.
Characteristic
Training
Validation
Test
MIMIC-IV
Age group, years
18–30
3,387 (5.38%)
385 (4.89%)
442 (5.62%)
569 (14.33%)
31–40
10,201 (16.21%)
1,352 (17.19%)
1,272 (16.17%)
601 (15.14%)
41–50
7,323 (11.64%)
845 (10.74%)
917 (11.66%)
775 (19.51%)
51–60
14,327 (22.77%)
1,812 (23.03%)
1,783 (22.67%)
916 (23.06%)
61–70
20,332 (32.31%)
2,558 (32.52%)
2,542 (32.32%)
599 (15.08%)
Table 2 : Demographic characteristics of the VHA training, validation, and held-out test cohorts and the external MIMIC-IV cohort. Values are reported as number of patients and percentage within each cohort.
Family
Model
AUPRC
AUROC
PPV
Classical ML
RandomForest
10.07
64.50
6.77
LogisticRegression
11.93
59.46
3.38
SVM
7.02
54.22
8.88
Deep sequence
GRU
11.26
62.48
7.61
TransformerEHR
13.51
60.13
7.44
Fine-tuned LM
Qwen3-1.7B
20.79
67.59
12.25
Table 3: Primary predictive performance for 180-day opioid overdose prediction on the held-out VHA test set. Models are grouped into classical machine-learning baselines, deep-sequence models, fine-tuned language models, diagnosis-adapted models, and the proposed OverdoseMoE framework. OverdoseMoE includes three expert-fusion variants: EW (Equal-Weight Fusion), PGA (Prior-Guided Adaptive Fusion), and GLA (Global-Local Adaptive Fusion). Risk-stratified PPV and recall at prespecified top- k thresholds are reported separately in Table 4 .
Model
PPV @ Top- k %
Recall @ Top- k %
Top-5%
Top-5%
1%
2%
5%
10%
1%
2%
5%
10%
Risk Enrichment
NNR
BioMistral-7B
8.86
6.96
6.34
5.20
1.47
2.32
5.28
8.66
1.05 ×
15.8
Qwen2.5-3B
7.59
6.32
6.85
6.48
1.26
2.11
5.70
10.78
1.14 ×
14.6
ClinicalMamba-2.8B
10.12
8.22
7.61
6.48
1.69
2.74
6.34
10.78
1.27 ×
13.1
Qwen3-1.7B
58.22
40.50
24.36
18.55
9.72
13.53
20.29
30.86
4.05 ×
4.1
Qwen3-4B
67.08
41.77
22.58
16.51
11.20
13.95
18.81
27.48
3.76 ×
4.4
Table 4 : Risk-stratified PPV and recall for the evaluated models among patients ranked in the top 1%, 2%, 5%, and 10% of predicted 180-day opioid-overdose risk. OverdoseMoE includes three expert-fusion variants: EW (Equal-Weight Fusion), PGA (Prior-Guided Adaptive Fusion), and GLA (Global-Local Adaptive Fusion). The top-5% threshold is underlined because it was prespecified for targeted risk review. Risk enrichment is calculated as PPV at the top-5% threshold divided by the overall opioid-overdose prevalence of 6.01%. Number needed to review (NNR) is calculated as 100 divided by PPV at the top-5% threshold. All PPV and recall values are reported as percentages.
OODQwen3-1.7B
OverdoseMoE-EW
OverdoseMoE-PGA
OverdoseMoE-GLA
Group
Subgroup
PPV
AUPRC
AUROC
PPV
AUPRC
AUROC
PPV
AUPRC
AUROC
PPV
AUPRC
AUROC
Age
18–30
13.33
25.45
56.52
13.12
23.03
62.14
12.68
23.50
59.78
12.82
22.87
61.92
31–40
11.89
21.73
64.32
11.42
20.59
64.64
12.53
22.07
65.65
11.16
20.43
64.57
41–50
7.73
14.86
65.08
7.46
12.76
63.66
7.92
16.56
62.17
7.53
15.03
63.46
51–60
12.53
22.42
67.70
13.11
23.84
68.27
13.23
23.44
68.44
12.81
23.54
68.09
61–70
14.25
24.44
69.19
13.93
26.56
72.25
15.96
27.36
72.33
13.90
26.44
71.84
Table 5 : Subgroup performance of OODQwen3-1.7B and the three OverdoseMoE variants for 180-day opioid overdose prediction across age, sex, and race/ethnicity strata in the held-out VHA test set. EW denotes Equal-Weight Fusion, PGA denotes Prior-Guided Adaptive Fusion, and GLA denotes Global-Local Adaptive Fusion. Performance is reported as PPV, AUPRC, and AUROC, with all values expressed as percentages. Dashes indicate subgroup results that were not available or not reported. Results for strata with very few positive events should be interpreted descriptively because of the instability of subgroup-level estimates.
Family
Model
AUPRC
AUROC
PPV
Fine-tuned LM
Mamba-2.8B
10.96
55.08
7.50
BioMistral-7B
10.17
57.64
7.69
Qwen2.5-3B
7.27
50.12
7.24
ClinicalMamba-2.8B
8.50
48.78
7.00
Qwen3-1.7B
10.70
54.94
9.14
Qwen3-4B
14.11
58.94
10.01
Table 6 : External validation on an independent OUD cohort constructed from MIMIC-IV for 180-day opioid overdose prediction. OverdoseMoE includes three expert-fusion variants: EW (Equal-Weight Fusion), PGA (Prior-Guided Adaptive Fusion), and GLA (Global-Local Adaptive Fusion). The best result for each metric is shown in bold.
Figure 1 : Ablation of Input Representations for OODQwen3-1.7B.
Figure 2 : Ablation of Input Representations for Qwen3-4B.
Figure 3 : Ablation of Input Representations for OODQwen3-8B.
Figure 4 : Overview of the OverdoseMoE framework for 180-day opioid overdose risk prediction. The framework includes four components. (1) Longitudinal diagnosis data : nationwide VHA electronic health record data are converted into longitudinal sequences of clinical diagnosis descriptions for continued pretraining. (2) Cohort construction : an OUD cohort is used for model fine-tuning and evaluation, with one-year longitudinal diagnosis histories as input and opioid overdose within 180 days as the prediction outcome; an independent MIMIC-IV OUD cohort is used for external evaluation. (3) Continued pretraining and task-specific adaptation : ClinicalMamba-2.8B and Qwen3-1.7B are continued pretrained on large-scale VHA diagnosis sequences to obtain OODMamba and OODQwen, followed by fine-tuning for binary opioid overdose prediction. (4) Multi-expert fusion : OODQwen and independently fine-tuned Qwen3-4B and Qwen3-8B models provide complementary prediction logits, which are combined using equal-weight (EW), prior-guided adaptive (PGA), or global–local adaptive (GLA) fusion to produce the final overdose risk prediction.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S2 : Cohort construction for the VHA opioid use disorder study population.
Outcome Ascertainment
ICD-10 Diagnosis Codes
EHR opioid overdose event (composite event captured within the health system)
Feature selection is a critical step in electronic health record (EHR)-based predictive modeling, where input variables are often high-dimensional, sparse, noisy, and redundant. Large feature sets not only increase computational burden and overfitting risk, but also make model interpretation difficult, leading to limited usefulness in clinical settings. In this study, we focus on diagnosis-related features and compare five feature selection paradigms for opioid use disorder (OUD) prediction: recurrence enrichment, NTK-motivated early gradient sensitivity, LightGBM-SHAP, Elastic Net, and large language model (LLM)-guided semantic selection. We use a unified preprocessing and evaluation framework and assess each method by downstream predictive performance, resampling stability, and representation of infrequent diagnosis codes. Our results demonstrate that performance improves with larger feature budgets with diminishing returns beyond a moderate size. NTK sensitivity provides the best overall balance of accuracy and stability, and LLM-guided selection contributes complementary clinically meaningful signals despite lower standalone performance.
Zihan Ding, Yinan Liu, Tengfei Ma +6
Stony Brook University · Rutgers University · University of Vermont
Electronic health records (EHRs) may incompletely capture patient-reported factors associated with opioid use disorder (OUD). We evaluated whether survey data improve prediction of a first recorded OUD diagnosis among 267,747 All of Us participants with documented opioid exposure, including 15,287 OUD cases. We compared EHR-only and EHR+survey models across 6-, 12-, and 24-month look-back windows using logistic regression, random forest, XGBoost, LightGBM, multilayer perceptron, LSTM, GRU, and Transformer. Survey augmentation improved PR-AUC across all 24 model-window combinations by 0.0087-0.0505; the best 24-month LightGBM model improved from 0.6219 to 0.6603. Survey coverage increased with longer windows and differed by OUD status (24 months: 21.7% OUD-positive vs. 60.7% OUD-negative). Permutation analysis ranked survey features as the second most important information domain at 24 months in both evaluated models. Patient-reported data provide complementary predictive signals beyond structured EHRs while highlighting the importance of survey availability.
Xiyue Jiang, Zihan Ding, Grace Han +3
Department of Applied Mathematics & Statistics, Stony Brook University, Stony Brook, NY, USA · Department of Computer Science, Stony Brook University, Stony Brook, NY, USA · Department of Biomedical Informatics, Stony Brook University, Stony Brook, NY, USA +1
Longitudinal clinical notes contain rich evidence of how patients evolve over time, but converting this signal into training supervision for clinical prediction remains challenging. We extend Foresight Learning to clinical prediction by converting time-ordered MIMIC-III notes into examples consisting of past patient context, a natural-language question about a possible future event, and a label resolved from later documentation. This process yields 6,900 prediction examples from 702 admissions across medications, procedures, organ support, microbiology, and mortality. A small LoRA adapter trained on these examples improves over the prompted base model, reducing expected calibration error from 0.1269 to 0.0398 and Brier score from 0.199 to 0.145, while slightly outperforming GPT-5 point estimates on held-out questions. The approach enables reusable clinical prediction supervision from longitudinal notes without hand-engineered structured features or endpoint-specific classifiers.