Accurate medication recommendation is central to clinical decision-making, directly determining therapeutic efficacy and patient safety. However, existing methods suffer from two key limitations: drugs are often abstracted as discrete tokens, ignoring their molecular structures and pharmacological mechanisms, and the commonly used "flat" recommendation paradigm fails to leverage the hierarchical logic of the internationally standardized Anatomical Therapeutic Chemical (ATC) classification system. To address these issues, we propose HADRec, a Hierarchy-Aware Drug Recommendation framework that integrates molecular knowledge with electronic health records (EHRs). HADRec employs LLaMA-7B to encode clinical notes for rich patient representations and ChemBERTa to encode drug Simplified Molecular Input Line Entry System strings, building a global molecular knowledge base. A cross-attention mechanism then performs deep multimodal fusion between patient states and drug features. The framework further incorporates a hierarchical predictor and a novel consistency constraint loss to enforce strict adherence to ATC logical dependencies. Extensive experiments on MIMIC-III demonstrate that HADRec achieves state-of-the-art performance across Jaccard, F1, and PR-AUC. External validation on MIMIC-IV confirms strong generalization under distribution shifts, and calibration analysis shows well-calibrated predictive confidence on MIMIC-IV with ECE = 0.04, and Brier = 0.06. Counterfactual evaluation reveals clinically aligned reasoning, disentangling disease-specific treatments from general care. Together, these results establish HADRec as a high-performance, interpretable, and clinically grounded pathway toward safe and reliable AI-driven medication recommendation.
Figures & tables
Figure 1: The workflow of the HADRec framework: (a) Multimodal representation learning employs LLaMA-7B and ChemBERTa to extract patient clinical profiles and drug molecular features, respectively. (b) Multimodal fusion performs deep interactions between patient states and drug molecular knowledge via a cross-attention mechanism. (c) The hierarchical prediction module constructs a four-level predictor to generate drug recommendations following the top-down ATC taxonomy. (d) The consistency constraint module incorporates a consistency loss function to enforce clinical logicality in the recommendations.
Figure 2: Illustration of the prompting template used for model input. The placeholders <…> are populated with patient EHR data.
Category
Item
Count
Patients
Total patients
6350
Average visits patient
2.37
Clinical Events
Unique diagnosis codes
1958
Unique procedure codes
1430
Medications
Unique medications (ATC-L4)
110
Average medications per visit
19.57
Table 1: Summary statistics of the processed dataset derived from MIMIC-III.
Category
Model
Key-Technique
Drug-Knowledge Used
Use ATC?
Knowledge-Graph & EHR-based
GAMENet [ 27 ]
Memory-Network & GNN
Knowledge-Graph (KG)
×
SafeDrug [ 37 ]
DDI-controllable-Loss
Molecular-Graph
×
RASNet [ 45 ]
Reverse-Time Attention
–
×
EGNeT [ 25 ]
GNN & Explainable Module
External-Medical-KG
×
Molecular-Structure-Aware
DEPOT [ 42 ]
Motif-Tree & Graph-Transformer
Molecular-Motifs
×
MoleRec [ 39 ]
Substructure Interaction & Annealed Loss
Molecular Substructures
×
Table 2: Categorization and key characteristics of the baseline models.
Category
Notation
Description
Value
Model Architecture
ddim
Hidden layer dimension
4096
dmol
Molecular embedding dimension
768
Attention Mechanism
λ
Residual connection scaling factor
0.7
τcross
Soft masking threshold for cross-attention
0.1
Loss Weighting
(λi)i=14
Weights for hierarchical BCE loss (L1-L4)
0.4, 0.4, 1.0, 1.5
λcons
Global weight for consistency constraint loss
0.01
Table 3: Hyperparameter settings of the HADRec framework
Category
Model
Jaccard ↑
F1 ↑
PR-AUC ↑
DDI ↓
Avg.Med ↓
Knowledge-Graph & EHR-based
GAMENet [ 27 ]
0.5067±0.0025
0.6626±0.0025
0.7631±0.0030
0.0864±0.0006
27.21±0.11
SafeDrug [ 37 ]
0.5213±0.0030
0.6768±0.0027
0.7647±0.0025
0.0589±0.0005
19.92±0.16
RASNet [ 45 ]
0.5401±0.0021
0.6931±0.0019
0.7882±0.0025
0.0599±0.0009
21.04±0.15
EGNeT [ 25 ]
0.5431±0.001
0.6955±0.0007
0.7842±0.0023
0.0693±0.0012
23.33±0.01
Molecular-Structure-Aware
DEPOT [ 42 ]
0.5352±0.0022
0.6890±0.003
0.7816±0.0035
0.0684±0.0007
20.62±0.19
MoleRec [ 39 ]
0.5301±0.0025
0.6841±0.0022
0.7748±0.0022
0.0756±0.0006
22.22±0.17
Table 4: Performance comparison between HADRec and baseline models. Results are reported as mean ± standard deviation across multiple runs. The best performances are highlighted in bold.
MKR
HP
CL
Jaccard ↑
F1 ↑
PR-AUC ↑
×
×
×
0.5397±0.0063
0.6904±0.0063
0.7946±0.0060
✓
×
×
0.5405±0.0090
0.6915±0.0087
0.7944±0.0066
✓
✓
×
0.5472±0.0051
0.6967±0.0050
0.7986±0.0052
×
✓
✓
0.5450±0.0054
0.6950±0.0053
0.7971±0.0049
✓
✓
✓
0.5517±0.0059
0.7010±0.0054
0.7994±0.0057
Table 5: Ablation study of key components of HADRec
Figure 3: Performance comparison demonstrates the superiority of full hierarchical supervision. The figure shows Jaccard (a-c), F1-score (d-f), and PR-AUC (g-i) across four ATC levels (L1-L4) under three supervision settings: L1+L4 (a, d, g), L1+L3+L4 (b, e, h), and full supervision L1+L2+L3+L4 (c, f, i).
Figure 4: Hierarchical attention distribution for a representative patient case. The heatmap visualizes the attention weights across Levels 1 to 4 for the ten drugs with the highest predicted probabilities at Level 4. Darker shading indicates stronger attention, clearly revealing the structured “parent-to-child” information flow constrained by the ATC hierarchy. Cells sharing the same background pattern correspond to drugs belonging to the same therapeutic class, highlighting consistent information transfer within hierarchical groups. Specifically, dotted (·), horizontal (—), and vertical (|) patterns denote A0-, A1-, and B0-level classes, respectively.
method
Threshold
Jaccard ↑
F1 ↑
Precision ↑
Recall ↑
Global Threshold
0.1
0.4578±0.0032
0.6186±0.0027
0.4789±0.0031
0.9184±0.0039
0.2
0.5253±0.0057
0.6791±0.0050
0.5848±0.0052
0.8443±0.0065
0.3
0.5491±0.0065
0.6992±0.0060
0.6589±0.0061
0.7741±0.0083
0.35
0.5517±0.0059
0.7010±0.0054
0.6891±0.0064
0.7403±0.0072
0.4
0.5470±0.0054
0.6969±0.0050
0.7126±0.0063
0.7069±0.0073
0.5
0.5275±0.0051
0.6789±0.0050
0.7595±0.0075
0.6349±0.0062
Table 6: Performance comparison of global thresholding (0.1–0.9) and class-wise dynamic thresholding on HADRec. Results are reported as mean ± standard deviation across multiple runs. The best performances are highlighted in bold.
Table 7: Counterfactual medication recommendations under selective disease removal.
Model
Jaccard ↑
F1 ↑
PR-AUC ↑
DDI ↓
Avg.Med ↓
KNN [ 27 ]
0.3978±0.0023
0.5473±0.0022
0.3917±0.0023
0.0815±0.0013
12.78±0.2203
LR [ 24 ]
0.4499±0.0033
0.5985±0.0027
0.7259±0.0024
0.0769±0.0008
10.36±0.1111
ECC [ 26 ]
0.4350±0.0017
0.5809±0.0020
0.7193±0.0021
0.0681±0.0003
8.76±0.2115
RETAIN [ 7 ]
0.4182±0.0020
0.5729±0.0016
0.6775±0.0012
0.0835±0.0004
11.29±0.2007
LEAP [ 41 ]
0.4123±0.0022
0.5661±0.0019
0.6026±0.0025
0.0667±0.0005
12.13±0.1530
GAMENet [ 27 ]
0.4594±0.0015
0.6126±0.0011
0.7229±0.0020
0.0803±0.0006
18.86±0.0002
Table 8: External validation and cross-dataset transferability evaluation of HADRec on MIMIC-IV.
Dataset
ECE
Brier Score
MIMIC-III
0.09
0.12
MIMIC-IV
0.04
0.06
Transfer (III→IV)
0.04
0.07
Table 9: Calibration and uncertainty evaluation results.
Model
Backbone
GPU Device
Memory Usage
Avg. Inference Time
HADRec
LLaMA-7B
NVIDIA RTX3090
12.75GB
3.14 s/sample
Table 10: Inference efficiency and deployment cost analysis.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Longitudinal EHR context analysis across different visit window configurations on the MIMIC-III cohort (N=6,350). (A) Distribution of patient counts by number of visits (log scale). (B) Cumulative patient coverage as a function of window size; window size 3 covers 89.0% of patients. (C) Marginal coverage gain diminishes rapidly beyond window size 3. (D) Information retention rate for truncated patients (n=701) with mean 63.1%. (E) Temporal distance of visits to last admission; truncated visits are significantly older (median 2.4 years). (F) Key summary statistics.
Medication recommendation predicts medications for patient visits, but existing methods still face two key challenges. At the model level, traditional drug recommendation methods only predict structured drug codes with limited evidence grounding, while LLM agents can use richer clinical context but may lack safety verification and traceability. At the task level, existing benchmarks often use broad medication categories, which ignore subgroup-level safety differences and can lead to risk overestimation. We introduce the first fine-grained medication recommendation setting based on fourth-level ATC code generation. We propose Safe Prescription Agent (SafeRx-Agent), a knowledge-grounded multi-agent framework that uses patient context, external clinical knowledge, and safety verification to recommend traceable medication sets. Experimental results on MIMIC-III and MIMIC-IV datasets show that SafeRx-Agent improves fine-grained medication prediction accuracy while controlling drug interactions, contraindications, and medication set size.
Xinyu Wang, Hanwei Wu, Zhenghan Tai +7
McGill University · McMaster University · University of Toronto +1
Medication recommendation from electronic health records must balance predictive accuracy against the risk of adverse drug-drug interactions (DDIs) under polypharmacy. Existing safety-aware recommenders operate at one of two granularities: the drug code, which treats each medication as an indivisible token, or the molecular substructure, which is finer than pharmacological interaction knowledge is actually organized. We argue that the active ingredient is the missing granularity, and introduce GRAIN, a medication recommendation framework built around it. GRAIN encodes longitudinal patient trajectories (diagnoses, procedures, past medications) with a selective state space backbone that handles long, irregular visit sequences in linear time. On top of it we introduce a joint objective unifying three knowledge sources aligned to a common medication vocabulary: a drug-level DDI graph, an ingredient-level DDI graph obtained by normalizing medication codes to active ingredients via RxNorm, and an EHR-derived co-prescription graph. A proportional controller adapts the accuracy-safety trade-off to the observed validation DDI rate rather than fixing it a priori. Under strictly matched settings -- identical preprocessing, cohort, vocabulary, split, and evaluation code -- GRAIN improves over a re-implemented MambaHealth baseline on MIMIC-IV across all standard multi-label metrics (Jaccard 0.4488 to 0.4983, PRAUC 0.6911 to 0.7485, F1 0.5989 to 0.6453) while reducing the drug-level DDI rate from 0.1875 to 0.0948. We further define an ingredient-level DDI rate, a safety measure invisible to drug-code-level evaluation. The results indicate that ingredient-level normalization recovers predictive signal erased by code-level aggregation, and that it is complementary to, rather than in competition with, accurate sequence modeling.
Juao Fan, Jinhan Li, Shengxin Zhu
Guangdong Provincial Key Laboratory of Interdisciplinary Research and Application for Data Science, Beijing Normal–Hong Kong Baptist University, Zhuhai, China
Large language models (LLMs) exhibit strong natural-language reasoning abilities for clinical decision support, but struggle to effectively model structured longitudinal electronic health records (EHRs). In contrast, EHR foundation models can learn predictive patient representations, yet lack interpretable language-based reasoning. To bridge this gap, we propose ChatHealthAI, a multimodal reasoning framework that aligns structured EHR representations from a pretrained EHR foundation model with the semantic space of a frozen LLM through a task-aware resampler. By integrating longitudinal patient representations with refined clinical event descriptions, ChatHealthAI enables clinically grounded natural-language reasoning while maintaining accurate patient prediction. We evaluated ChatHealthAI on three clinical predictive tasks from the EHRSHOT benchmark. Results show that ChatHealthAI improves reasoning quality and interpretability while preserving competitive predictive performance. These findings highlight the potential of integrating EHR foundation models with pretrained LLMs for interpretable clinical prediction.
Bo-Hong Wang, Baicheng Peng, Ruilin Wang +3
School of Computer Science, McGill University, Montreal, QC, Canada · Mila - Quebec AI Institute, Montreal, QC, Canada