Accurate medication recommendation is central to clinical decision-making, directly determining therapeutic efficacy and patient safety. However, existing methods suffer from two key limitations: drugs are often abstracted as discrete tokens, ignoring their molecular structures and pharmacological mechanisms, and the commonly used "flat" recommendation paradigm fails to leverage the hierarchical logic of the internationally standardized Anatomical Therapeutic Chemical (ATC) classification system. To address these issues, we propose HADRec, a Hierarchy-Aware Drug Recommendation framework that integrates molecular knowledge with electronic health records (EHRs). HADRec employs LLaMA-7B to encode clinical notes for rich patient representations and ChemBERTa to encode drug Simplified Molecular Input Line Entry System strings, building a global molecular knowledge base. A cross-attention mechanism then performs deep multimodal fusion between patient states and drug features. The framework further incorporates a hierarchical predictor and a novel consistency constraint loss to enforce strict adherence to ATC logical dependencies. Extensive experiments on MIMIC-III demonstrate that HADRec achieves state-of-the-art performance across Jaccard, F1, and PR-AUC. External validation on MIMIC-IV confirms strong generalization under distribution shifts, and calibration analysis shows well-calibrated predictive confidence on MIMIC-IV with ECE = 0.04, and Brier = 0.06. Counterfactual evaluation reveals clinically aligned reasoning, disentangling disease-specific treatments from general care. Together, these results establish HADRec as a high-performance, interpretable, and clinically grounded pathway toward safe and reliable AI-driven medication recommendation.
Figures & tables
Figure 1: The workflow of the HADRec framework: (a) Multimodal representation learning employs LLaMA-7B and ChemBERTa to extract patient clinical profiles and drug molecular features, respectively. (b) Multimodal fusion performs deep interactions between patient states and drug molecular knowledge via a cross-attention mechanism. (c) The hierarchical prediction module constructs a four-level predictor to generate drug recommendations following the top-down ATC taxonomy. (d) The consistency constraint module incorporates a consistency loss function to enforce clinical logicality in the recommendations.
Figure 2: Illustration of the prompting template used for model input. The placeholders <…> are populated with patient EHR data.
Category
Item
Count
Patients
Total patients
6350
Average visits patient
2.37
Clinical Events
Unique diagnosis codes
1958
Unique procedure codes
1430
Medications
Unique medications (ATC-L4)
110
Average medications per visit
19.57
Table 1: Summary statistics of the processed dataset derived from MIMIC-III.
Category
Model
Key-Technique
Drug-Knowledge Used
Use ATC?
Knowledge-Graph & EHR-based
GAMENet [ 27 ]
Memory-Network & GNN
Knowledge-Graph (KG)
×
SafeDrug [ 37 ]
DDI-controllable-Loss
Molecular-Graph
×
RASNet [ 45 ]
Reverse-Time Attention
–
×
EGNeT [ 25 ]
GNN & Explainable Module
External-Medical-KG
×
Molecular-Structure-Aware
DEPOT [ 42 ]
Motif-Tree & Graph-Transformer
Molecular-Motifs
×
MoleRec [ 39 ]
Substructure Interaction & Annealed Loss
Molecular Substructures
×
Table 2: Categorization and key characteristics of the baseline models.
Category
Notation
Description
Value
Model Architecture
ddim
Hidden layer dimension
4096
dmol
Molecular embedding dimension
768
Attention Mechanism
λ
Residual connection scaling factor
0.7
τcross
Soft masking threshold for cross-attention
0.1
Loss Weighting
(λi)i=14
Weights for hierarchical BCE loss (L1-L4)
0.4, 0.4, 1.0, 1.5
λcons
Global weight for consistency constraint loss
0.01
Table 3: Hyperparameter settings of the HADRec framework
Category
Model
Jaccard ↑
F1 ↑
PR-AUC ↑
DDI ↓
Avg.Med ↓
Knowledge-Graph & EHR-based
GAMENet [ 27 ]
0.5067±0.0025
0.6626±0.0025
0.7631±0.0030
0.0864±0.0006
27.21±0.11
SafeDrug [ 37 ]
0.5213±0.0030
0.6768±0.0027
0.7647±0.0025
0.0589±0.0005
19.92±0.16
RASNet [ 45 ]
0.5401±0.0021
0.6931±0.0019
0.7882±0.0025
0.0599±0.0009
21.04±0.15
EGNeT [ 25 ]
0.5431±0.001
0.6955±0.0007
0.7842±0.0023
0.0693±0.0012
23.33±0.01
Molecular-Structure-Aware
DEPOT [ 42 ]
0.5352±0.0022
0.6890±0.003
0.7816±0.0035
0.0684±0.0007
20.62±0.19
MoleRec [ 39 ]
0.5301±0.0025
0.6841±0.0022
0.7748±0.0022
0.0756±0.0006
22.22±0.17
Table 4: Performance comparison between HADRec and baseline models. Results are reported as mean ± standard deviation across multiple runs. The best performances are highlighted in bold.
MKR
HP
CL
Jaccard ↑
F1 ↑
PR-AUC ↑
×
×
×
0.5397±0.0063
0.6904±0.0063
0.7946±0.0060
✓
×
×
0.5405±0.0090
0.6915±0.0087
0.7944±0.0066
✓
✓
×
0.5472±0.0051
0.6967±0.0050
0.7986±0.0052
×
✓
✓
0.5450±0.0054
0.6950±0.0053
0.7971±0.0049
✓
✓
✓
0.5517±0.0059
0.7010±0.0054
0.7994±0.0057
Table 5: Ablation study of key components of HADRec
Figure 3: Performance comparison demonstrates the superiority of full hierarchical supervision. The figure shows Jaccard (a-c), F1-score (d-f), and PR-AUC (g-i) across four ATC levels (L1-L4) under three supervision settings: L1+L4 (a, d, g), L1+L3+L4 (b, e, h), and full supervision L1+L2+L3+L4 (c, f, i).
Figure 4: Hierarchical attention distribution for a representative patient case. The heatmap visualizes the attention weights across Levels 1 to 4 for the ten drugs with the highest predicted probabilities at Level 4. Darker shading indicates stronger attention, clearly revealing the structured “parent-to-child” information flow constrained by the ATC hierarchy. Cells sharing the same background pattern correspond to drugs belonging to the same therapeutic class, highlighting consistent information transfer within hierarchical groups. Specifically, dotted (·), horizontal (—), and vertical (|) patterns denote A0-, A1-, and B0-level classes, respectively.
method
Threshold
Jaccard ↑
F1 ↑
Precision ↑
Recall ↑
Global Threshold
0.1
0.4578±0.0032
0.6186±0.0027
0.4789±0.0031
0.9184±0.0039
0.2
0.5253±0.0057
0.6791±0.0050
0.5848±0.0052
0.8443±0.0065
0.3
0.5491±0.0065
0.6992±0.0060
0.6589±0.0061
0.7741±0.0083
0.35
0.5517±0.0059
0.7010±0.0054
0.6891±0.0064
0.7403±0.0072
0.4
0.5470±0.0054
0.6969±0.0050
0.7126±0.0063
0.7069±0.0073
0.5
0.5275±0.0051
0.6789±0.0050
0.7595±0.0075
0.6349±0.0062
Table 6: Performance comparison of global thresholding (0.1–0.9) and class-wise dynamic thresholding on HADRec. Results are reported as mean ± standard deviation across multiple runs. The best performances are highlighted in bold.
Table 7: Counterfactual medication recommendations under selective disease removal.
Model
Jaccard ↑
F1 ↑
PR-AUC ↑
DDI ↓
Avg.Med ↓
KNN [ 27 ]
0.3978±0.0023
0.5473±0.0022
0.3917±0.0023
0.0815±0.0013
12.78±0.2203
LR [ 24 ]
0.4499±0.0033
0.5985±0.0027
0.7259±0.0024
0.0769±0.0008
10.36±0.1111
ECC [ 26 ]
0.4350±0.0017
0.5809±0.0020
0.7193±0.0021
0.0681±0.0003
8.76±0.2115
RETAIN [ 7 ]
0.4182±0.0020
0.5729±0.0016
0.6775±0.0012
0.0835±0.0004
11.29±0.2007
LEAP [ 41 ]
0.4123±0.0022
0.5661±0.0019
0.6026±0.0025
0.0667±0.0005
12.13±0.1530
GAMENet [ 27 ]
0.4594±0.0015
0.6126±0.0011
0.7229±0.0020
0.0803±0.0006
18.86±0.0002
Table 8: External validation and cross-dataset transferability evaluation of HADRec on MIMIC-IV.
Dataset
ECE
Brier Score
MIMIC-III
0.09
0.12
MIMIC-IV
0.04
0.06
Transfer (III→IV)
0.04
0.07
Table 9: Calibration and uncertainty evaluation results.
Model
Backbone
GPU Device
Memory Usage
Avg. Inference Time
HADRec
LLaMA-7B
NVIDIA RTX3090
12.75GB
3.14 s/sample
Table 10: Inference efficiency and deployment cost analysis.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Longitudinal EHR context analysis across different visit window configurations on the MIMIC-III cohort (N=6,350). (A) Distribution of patient counts by number of visits (log scale). (B) Cumulative patient coverage as a function of window size; window size 3 covers 89.0% of patients. (C) Marginal coverage gain diminishes rapidly beyond window size 3. (D) Information retention rate for truncated patients (n=701) with mean 63.1%. (E) Temporal distance of visits to last admission; truncated visits are significantly older (median 2.4 years). (F) Key summary statistics.
Guangdong Provincial Key Laboratory of Interdisciplinary Research and Application for Data Science, Beijing Normal–Hong Kong Baptist University, Zhuhai, China