Multimodal AI predicts clinical outcomes of drug combinations from preclinical data
Authors: Yepeng Huang, Xiaorui Su, Varun Ullanat, Intae Moon, Ivy Liang, Lindsay Clegg, Damilola Olabode, Ruthie Johnson, +5 more
Organizations: Department of Biomedical Informatics, Harvard Medical School, Boston, MA · Program in Biological and Biomedical Sciences, Harvard Medical School, Boston, MA · Harvard College, Cambridge, MA · Clinical Pharmacology and Quantitative Pharmacology, Clinical Pharmacology & Safety Sciences, R&D, AstraZeneca, Gaithersburg, MD · Clinical Pharmacology and Quantitative Pharmacology, Clinical Pharmacology & Safety Sciences, R&D, AstraZeneca, Waltham, MA · Program in Computational Biology, Carnegie Mellon University, Pittsburgh, PA · Department of Medical Oncology, Dana-Farber Cancer Institute and Harvard Medical School, Boston, MA · Broad Institute of MIT and Harvard, Cambridge, MA · Imaging and Data Analytics, Clinical Pharmacology & Safety Sciences, R&D, AstraZeneca, Waltham, MA · Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University, Allston, MA · Harvard Data Science Initiative, Cambridge, MA
Predicting clinical outcomes from preclinical data is essential for selecting safe and effective drug combinations and for reducing late-stage failures. AI models use molecular structure and target annotations, and do not leverage the perturbation readouts that report how a compound acts in a cellular context. Here we introduce Madrigal, a multimodal AI model that learns from structural, pathway, cell-viability, and transcriptomic data. Madrigal aligns these modalities across 21,842 compounds into a shared latent space and predicts combination outcomes even for drugs observed in only a subset of the data modalities. Trained on 158 expert-curated and 795 patient-reported combination outcomes, Madrigal outperforms single-modality and state-of-the-art multimodal methods. Ablations show that modality alignment and multimodal input each improve predictive performance. Madrigal predicts elevated risk for combinations that share membrane transporters. In head-to-head trials that compare two combination arms,the arm with the higher observed incidence of neutropenia, anemia, alopecia, or hypoglycemia receives the higher predicted risk in 25 of 28 comparisons. In MASH, Madrigal ranks resmetirom among the candidates with favorable predicted safety when paired with type 2 diabetes drugs. Madrigal also improves adverse-event prediction in a longitudinal patient cohort and an independent oncology cohort and predicts efficacy in primary acute myeloid leukemia samples and patient-derived xenografts.
Figures & tables
Figure 1: Madrigal integrates multimodal preclinical data to predict clinical outcomes of drug combinations. a, Overview of Madrigal . Comprising modality-specific encoders, a fusion module, and a prediction module, Madrigal is trained on expert-curated and patient-reported drug combination datasets to predict clinical outcomes. Modalities are fused with attention bottleneck modules (Methods and Supplementary Note ). The model enables three key applications: prediction of polypharmacy safety outcomes, efficacy prediction in patient-derived models, and personalized drug combination outcome predictions in patients. b, Overview of data modalities and the modality alignment framework. The modality-specific encoders are aligned via contrastive learning. Madrigal extracts information from multimodal data using specialized encoders (Methods Sec. 2 ). c, The missing data modality problem is evident in the scarcity of drugs with more than two available modalities. d, UMAP of modality-specific latent embeddings of ten randomly sampled drugs, before and after modality alignment in Madrigal . Prior to alignment, embeddings cluster by data type, while post-alignment, they cluster based on drug identity, enabling cross-modal integration.
Figure 2: Benchmarking Madrigal and performance analyses. a, Dataset splitting strategy for predicting safety outcomes of drug combinations. In the split-by-drugs setup, during training, all available modalities are used ( d2,d3,d4 ); at test time, test drugs are forced to only use the structure and transcriptomics modalities (patterned boxes) ( d1 ). b, Rank counts of Madrigal and existing methods across the 24 (dataset, split, metric) comparisons (2 datasets, 4 splits, and 3 metrics). c, Ablation of Madrigal components in the split-by-drugs (targets) setting. “W/o CL”, ablation model without modality alignment; “Struc. only”, ablation model with only structure modality during finetuning (all modalities available during modality alignment); “Struc. only w/o CL”, ablation model without modality alignment and with only structure modality available during finetuning. AUROC, area under the receiver operating characteristic curve; AUPRC, area under the precision-recall curve; Fmax, maximum of F-measure. d, Test performance of Madrigal and existing methods in the split-by-drugs (targets) and split-by-drugs (ATC) settings in the DrugBank (expert-curated drug combination outcomes) and TWOSIDES (patient-reported drug combination outcomes) datasets. The brackets highlight Madrigal’s improvements upon the best-performing existing method per (dataset, split, metric) comparison. ** p -value <0.005 , * p -value <0.05 , exact value for p -value <0.2 , n.s. otherwise; two-sided Mann-Whitney U test. e,f, Test performance increase for test drugs with increasing similarity to train drugs in terms of structure (e) or target profile (f). The progression from “Struc. only w/o CL”, “W/o CL”, to Madrigal represents progressive additions of multimodal input and modality alignment upon a simple model only taking molecular structure as input. For each test drug, its structural similarity (with the train set) is calculated as the average of the highest 5 Tanimoto similarities between its Morgan fingerprint with any train drug’s fingerprint. Target profile similarity is similarly defined as the average of the highest 5 Jaccard similarities between the drug’s target profile with any train drug’s profile. Target profiles of drugs are set of targets annotated to the drugs in DrugBank [ 47 ] . Analyses in (e-g) are performed with models trained on the DrugBank dataset, split-by-drugs (random) setting. Error bars show 95% confidence interval. Two-sided Mann-Whitney U test; ** p -value <0.005 . g, Test performance of Madrigal ablation with only structure modality and without modality alignment, ablation without modality alignment but with multi-modality, and full model, stratified by safety outcomes. Two-sided Wilcoxon signed rank test; ** p -value <0.005 .
Figure 3: Evaluating Madrigal predictions on external patient safety datasets. a-c, Model predictions of individual drug’s organ-specific adverse effects correlate with concern levels in three organ-specific adverse effect datasets (drug-induced liver injury (a), drug-induced cardiotoxicity (b), drug-induced QT prolongation (c)). Error bars show 95% confidence interval. Two-sided Mann-Whitney U test; ** p -value <0.005 . Higher prediction score indicates greater predicted concern. d, Model predictions of transporter-mediated DDIs (Methods Sec. 4.2 ) in combinations involving doxycycline (Dox). Piracetam is included as a reference as it is structurally highly similar to levetiracetam. For each drug pair, the prediction score column denotes the maximal prediction score among transporter-mediated safety outcomes that are relevant to the increase in serum concentration from DrugBank. Percentiles compare the max prediction score of Dox + X among either X + any curated DrugBank drug (“drug + X group”) or Dox + any curated DrugBank drug (“Dox + X group”). e-g, Drugs sharing the same transporters (e), carriers (f), or enzymes (g) are predicted to have a higher tendency to have relevant safety outcomes (Methods Sec. 4.2 ). Transporter, carrier, and enzyme information of drugs are obtained from DrugBank [ 47 ] . The highest prediction score among all potential transporter-, carrier-, or enzyme-mediated safety outcomes is considered for each drug pair (Methods Sec. 4.2 ). Two-sided Mann-Whitney U test; ** p -value <0.005 . h, Drugs sharing specific transporters are predicted to have a higher tendency of both common and specific transporter-related safety outcomes. Safety outcomes shown are ranked in the highest 10 for at least one transporter (across all drug pairs sharing it). The color gradient reflects the median prediction score (across all drug pairs sharing corresponding transporter), and the number in each cell is the ranking (max=141) of the median prediction score of the corresponding safety outcome among all safety outcomes for drug pairs sharing the corresponding transporter. Outcome signs are neutralized before ranking, collapsing the 158 DrugBank safety outcomes into 141 (Methods Sec. 4.2 ). The safety profiles of drugs sharing three enzymes, carriers, and targets, respectively, are also shown for comparison.
Figure 4: Madrigal predicts safety for clinically tested drug combinations. a, A candidate drug pair is evaluated in two ways. Top: head-to-head clinical trial arms yielding observed AE incidence. Bottom: Madrigal predicts safety scores with regard to the same adverse outcome. The model predictions are not calibrated to match the percentage. Agreement is assessed by whether the safer arm in the trial also receives a lower Madrigal score. b, Comparing Madrigal predictions with clinical trials (CT) AE data in advanced-stage clinical trials with multiple combination arms for neutropenia, hypoglycemia, anemia, and alopecia. c, Comparative safety assessment across different classes of drug combinations. Left to right bars within each group represent: (1) drug combinations containing at least one cancer drug that have been investigated in advanced stage (above Phase I), (2) drug combinations indicated for cancer that have been investigated in advanced stage, (3) FDA-approved drug combinations, (4) FDA-approved drug combinations containing at least one cancer drug, (5) PARPi combinations that have been investigated in advanced stage, and (6) pairwise combinations of a PARPi with any other cancer drug. Each drug combination’s safety profile is represented by the average of the five highest prediction scores for outcomes of each organ system. Higher prediction score indicates greater predicted concern. d, Predicted safety profiles of combination involving pioglitazone or rosiglitazone with T2D drugs. Each point represents the mean prediction score of pioglitazone or rosiglitazone, when combined with any T2D drug of a different mechanism of action, with regard to each relevant safety outcome. One-sided Wilcoxon signed rank test; ** p -value <0.005 . e, Predicted hyperkalemia-related safety profiles of drug combination involving a heart failure drug and any T2D drug of a different mechanism of action. SGLT2i, sodium/glucose cotransporter 2 inhibitor; ARB, angiotensin II receptor blocker; ARNi, angiotensin receptor/neprilysin inhibitor; ACEi, angiotensin-converting enzyme inhibitor; SZC, sodium zirconium cyclosilicate; HF, heart failure. Error bars show 95% confidence interval. One-sided Mann-Whitney U test; ** p -value <0.005 . f, Predicted safety of MASH clinical candidates from [ 83 ] in combination with T2D drugs, calculated from the average of highest 5 prediction scores across outcomes (Methods Sec. 4.4 ). Shown all MASH candidates, ranked by prediction scores.
Figure 5: Madrigal predicts individualized drug combination efficacy in patient-derived cancer models and adverse events in real-world patient cohorts. a, Using Madrigal to predict individualized drug combination efficacy in patient-derived cancer models. b, Performance of Madrigal in predicting synergistic drug combinations in BeatAML [ 39 ] , where prediction target is combination synergy (Methods Sec. 4.6 ). Model evaluation is conducted on randomly held-out patients. Patient-centric and drug-centric denote two ways of calculating AUROC to evaluate the model (Methods Sec. 4.6 ). Error bars show standard deviation across 25 runs (five Madrigal models × five training seeds). c, Performance of Madrigal in predicting drug combination efficacy in PDX Encyclopedia [ 41 ] when leaving each drug combination out. The prediction target is treatment response (BestAvgResponse). d, Predicted progression-free survival (PFS, TimeToDouble) for individual patient models treated with the (BKM120 + encorafenib) and (LEE011 + encorafenib) combinations. The predictor is trained on other drug combinations, with PFS as the prediction target. Predictions are color-coded by the observed best response category (calculated from response according to mRECIST [ 41 ] ) of each patient model. PD, progressive disease; SD, stable disease; PR, partial response; CR, complete response. e, Kaplan-Meier survival estimates stratified by predicted treatment response for the (BKM120 + encorafenib) and (encorafenib + binimetinib) combinations. The predictor is trained on other drug combinations with treatment response as the prediction target (same predictor as in (c)). f, Using Madrigal to predict personalized drug combination adverse events in real-world patient cohorts. g, Model training and inference for predicting drug combination outcomes in patients. h, Performance of TransformerEHR and TransformerEHR with Madrigal drug embeddings across re-admission prediction, mortality prediction, and adverse event prediction (anemia, hypoglycemia, hyperkalemia, hyponatremia, thrombocytopenia) tasks in the longitudinal event-time cohort. Error bars show standard deviation. i, Performance of combining Madrigal with patient information to predict adverse events for individual patients in the independent oncology cohort, compared with using Morgan fingerprint or one-hot regimen encoding. Error bars show standard deviation.
Multimodal drug discovery enables drug representation learning beyond chemical structure by incorporating cellular responses such as gene expression and cell morphology. However, direct fusion and instance-level contrastive alignment may mix mechanism-related signals with modality-specific noise and incorrectly separate structurally dissimilar but biologically related compounds. This limitation can obscure transferable mechanism patterns required for predicting the properties of unseen compounds. We introduce PMRD, a pharmacological response domain-guided framework for multimodal zero-shot drug property prediction. PMRD separates mechanism-consistent factors from modality-specific information and constructs a consensus response domain across three modalities. Mechanism candidate augmentation identifies locally stable factors, while retrieval-geometry attribution dynamically reweights the alignment and augmentation objectives according to whether their updates preserve inter-drug discriminability.This feedback suppresses training signals that conflict with mechanism-discriminative retrieval. PMRD further combines complementary representations through reliability-aware multiview retrieval. Experiments on public datasets show improved zero-shot property prediction and more biologically coherent drug neighborhoods. Hard-negative analysis further indicates fewer conflicts between structurally dissimilar but response-related compounds. These results support PMRD as an effective framework for mechanism-aware multimodal drug representation learning.\footnote{The code will be released upon publication.}
Jintao Huang, Lu Leng, Ziyuan Yang
Nanchang Hangkong University, Nanchang, Jiangxi, China · Sichuan University, Chengdu, Sichuan, China
Drug-drug interaction (DDI) prediction is a critical task in computational biomedicine, as adverse interactions between co-administered drugs can cause severe side effects and clinical risks. A key challenge is unseen-drug generalization, where interactions must be predicted for drugs not observed during training. Although multimodal DDI models exploit diverse drug-related information, their fusion mechanisms are often tied to specific prediction architectures, limiting their reuse across models. To address this, we propose AIM-DDI, an architecture-independent multimodal integration module that represents heterogeneous modality information as tokens in a shared latent space. By modeling dependencies across modality tokens through a unified fusion module, AIM-DDI enables model-agnostic integration of structural, chemical, and semantic drug signals across different DDI prediction architectures. Extensive evaluations across diverse DDI models and DrugBank-based settings show that AIM-DDI consistently improves prediction performance, with the strongest gains under the most challenging both-unseen setting where neither drug in a test pair is observed during training. These results suggest that treating multimodal integration as a reusable module, rather than a model-specific fusion component, is an effective strategy for robust unseen-drug DDI prediction.
Yerin Park, Sangseon Lee
Department of Artificial Intelligence, Inha University
Treating patients with combinations of drugs reduces the risk of resistance to any individual drug. Finding effective combinations is difficult because the large search space makes combinatorial screens prohibitively expensive, time consuming, and often technically infeasible. Predictive models can fill this gap, yet existing methods typically require molecular profiling of each sample and per-cohort training, limiting their applicability when time and tissue are scarce. To address this challenge, we introduce ScreenShot, a hierarchical transformer pretrained on 40 drug screening datasets covering 3,700 drugs and 6,000 biological samples, whose architecture mirrors the nested structure of screening data. Given a few-shot context of observations from a new patient, ScreenShot predicts the response of the sample to combination therapies through in-context learning, operating directly on functional measurements with no fine-tuning and no molecular profiling. On four held-out datasets, ScreenShot outperforms all baselines in both prediction accuracy and identification of selectively effective treatments. ScreenShot's internal representations are directly useful for experimental design: we use them to drive a weighted k-means++ active learning strategy that selects which experiments to run, achieving the same hit detection as uniform screening with a third of the budget. Source code and interactive dashboard: https://github.com/tansey-lab/screenshot.
Antoine de Mathelin, Christopher Tosh, Wesley Tansey
Computational Oncology Memorial Sloan Kettering Cancer Center New York, NY 10065