Organizations: School of Medicine, The Chinese University of Hong Kong, Shenzhen, Longgang District, Shenzhen, Guangdong 518172, China · School of Data Science, The Chinese University of Hong Kong, Shenzhen, Longgang District, Shenzhen, Guangdong 518172, China · Warshel Institute for Computational Biology, School of Medicine, The Chinese University of Hong Kong, Shenzhen, Longgang District, Shenzhen, Guangdong 518172, China · Guangdong Provincial Key Laboratory of Digital Biology and Drug Development, The Chinese University of Hong Kong, Shenzhen, Longgang District, Shenzhen, Guangdong 518172, China · Department of Endocrinology, Key Laboratory of Endocrinology of National Health Commission, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College, Beijing, 100730, P.R. China
Drug mechanism-of-action (MoA) modeling commonly relies on perturbational transcriptomes, but matched microRNA (miRNA) measurements are often unavailable. Inferred regulatory features offer a scalable way to reuse these data. Here, we present MIRCID, a framework comparing gene expression with inferred transcription factor (TF) activity and miRNA expression across pathway classification and similarity-based MoA retrieval. HubmiRNet infers 414 pan-cancer hub miRNAs (HubmiRs) from 977 L1000 landmark genes, achieving a Pearson correlation coefficient of 87.72%; its 1,298-output variant also outperformed SiCmiR on the full-miRNA task (71.21% versus 67.30%). In the evaluated comparisons, miRNA augmentation provided more consistent gains than TF activity. Generic embedding controls showed model-dependent utility, while complementarity analyses identified a distinct, partially linearly recoverable representation that retained gene-derived structure. Illustrative rescue cases linked improved classification to biologically plausible miRNA patterns in samples with weak transcriptional signatures. These findings support inferred HubmiRs as a biologically informed recoding of transcriptomic data for perturbational drug modeling, while leaving recovery of measured perturbational miRNA responses to further validation.
Figures & tables
Figure 1: Workflow of the proposed framework for integrating regulatory layers into drug mechanistic modeling. The workflow comprises four stages: data integration; miRNA/TF regulatory feature inference (including HubmiRNet); systematic comparison of regulatory feature layers across downstream perturbational drug modeling tasks; and model training with performance and feature-contribution evaluation.
Figure 2: HubmiRNet enables accurate inference of miRNA expression from mRNA profiles. a, mRNA–miRNA co-expression basis and model development. Building on the 414 hub miRNAs identified in SiCmiR 5 , we benchmarked multiple architectures for inferring miRNA expression from L1000 landmark-gene inputs. b, Pearson correlation coefficient (PCC) and root mean square error (RMSE) for the 414-HubmiR architecture benchmark and for a matched comparison between SiCmiR and a full-output HubmiRNet variant on the original 1,298-miRNA SiCmiR task. c, Principal component analysis of predicted and observed 414-HubmiR profiles, showing overlap in their global variance structure.
Figure 3: Systematic benchmarking identifies a robust TF activity inference configuration. a, Benchmarking workflow for TF activity (TFA) inference: gene expression profiles are combined with alternative prior knowledge networks (PKNs) to evaluate three representative TFA algorithms under a unified pipeline. b, Summary of benchmarking results across three perturbation datasets, showing that the DoRothEA+Priori configuration achieves the best or second-best performance for most evaluation metrics.
Figure 4: miRNA features consistently improve drug pathway inference performance across models. a, Pathway-classification workflow using Gene alone or Gene combined with inferred TF activity (TFA), inferred HubmiR expression, or both. b,c, Held-out accuracy ( b ) and macro-F1 ( c ) for random forest (RF), multilayer perceptron (MLP), Kolmogorov–Arnold network (KAN), and residual network (ResNet) classifiers. Points denote individual held-out runs; boxes show the median and interquartile range, whiskers the minimum and maximum, and diamonds with bars the mean ± s.d. Dashed lines show the mean gene-only PROGENy linear-model performance across 7 splits.
Figure 5: Generic embedding controls and complementary geometry of the inferred HubmiR representation. a, Held-out macro-F1 for RF and ResNet using Gene, Gene plus 414-dimensional principal-component analysis (PCA), autoencoder (AE), or random projection (RP), HubmiR alone, or Gene+HubmiR. Controls used the same 977 landmark-gene inputs; RP is averaged over five projections. b, Regularized canonical correlation analysis (CCA) between Gene and each representation across the first ten held-out modes (mean and bootstrap 95% CI; 20 splits). c, Debiased linear and radial-basis-function centred-kernel alignment (CKA) with Gene (seven evenly spaced display splits from 20). d, Held-out variance-weighted R2 from cross-validated multi-output ridge regression predicting HubmiR from full Gene or the 977-gene input for the same seven splits. e, Paired change in held-out macro-F1 after adding HubmiR to Gene or Gene to HubmiR across RF, MLP, KAN, and ResNet. In a , boxes show the median and interquartile range and whiskers the observed range; in e , diamonds denote means and lines the observed range.
Figure 6: Inferred HubmiR features rescue pathway classification across EGFR, MAPK, and hypoxia perturbations. The schematic summarizes the represented signaling contexts. Heatmaps show nine strict-rescue cases (three per pathway) in which Gene-only ResNet was incorrect and Gene+HubmiR was correct on the same held-out split. Each case displays eight pathway-associated PROGENy genes and three pathway-related inferred HubmiRs as training-fold empirical-percentile deviations ( −1 to +1 ; 0, training-fold median). Arrows denote activating or inhibiting perturbations; red and green frames identify the displayed gene and HubmiR features, respectively. When pathway-related gene signals were weak, the corresponding pathway-related HubmiRs showed larger deviations, supporting rescue by complementary miRNA features.
Algorithm
Feature Space
A375
A549
HA1E
HCC515
MCF7
VCAP
weightedKS
LM+TF
0.0459
0.0182
0.0514
0.0090
0.0546
0.0117
LM only
0.0668
0.0272
0.0636
0.0343
0.0990
0.0188
LM+Mi
0.0803
0.0377
0.0622
0.0465
0.1037
0.0264
LM+Mi+TF
0.0523
0.0169
0.0509
0.0094
0.0586
0.0226
KS
LM+TF
0.0394
0.0160
0.0275
0.0029
0.0533
0.0165
LM only
0.0662
0.0201
0.0643
0.0197
0.0829
0.0203
Table 1: Drug MoA similarity prediction performance ( AUC0.01 ) across four algorithms, four feature-selection strategies, and six cell lines. Bold and underlined values indicate the best and second-best results, respectively, for each algorithm. LM: landmark gene signature; Mi: inferred miRNA signature; TF: inferred transcription factor signature.
Figure S1: Embedding-control performance across learners and linear recoverability of inferred HubmiR features. a, Held-out pathway-classification accuracy for RF, MLP, KAN, and ResNet. b, Held-out macro-F1 for KAN and MLP. Six configurations are compared: Gene, Gene augmented with PCA, AE, or RP, HubmiR alone, and Gene+HubmiR ( n=7 runs per configuration). Gene and Gene+HubmiR reuse Figure 4 results. RP denotes the mean across five projection seeds within each run. Boxes show medians and interquartile ranges, whiskers the observed range, circles individual runs, and diamonds the mean ± s.d. c, AE training and validation reconstruction MSE, with the minimum validation error marked. d, Median held-out R2 across 20 splits for 414 inferred miRNAs predicted from full Gene and the 977-gene input; the dashed line denotes equal recovery. e, Aggregate held-out R2 across 20 frozen splits for cross-validated multi-output ridge models; bars show the mean ± s.d.
Figure S2: Pathway-wide rescue outcomes and selected strict-rescue catalogue across seven ResNet runs. a, Joint correctness of the paired Gene-only and Gene+HubmiR predictions across 595 held-out sample–run occurrences. Rescued denotes Gene-only incorrect and Gene+HubmiR correct; harmed denotes Gene-only correct and Gene+HubmiR incorrect. Counts and percentages use all 595 paired occurrences as the denominator. b, Rescued and harmed occurrence counts stratified by the true pathway label. c, Catalogue of the 27 selected strict-rescue profiles displayed in Supplementary Figure S3, annotated by pathway, dataset/profile identifier, perturbation, cellular context, and the incorrect Gene-only prediction. At most three profiles were selected per pathway; pathways with fewer available strict rescues were not padded. The nine profiles examined in Figure 6 were retained in the catalogue.
Figure S3: Feature profiles of selected strict-rescue cases across 11 pathways. Fold-relative empirical-percentile deviations for eight selected genes and three inferred HubmiRs across the same 27 strict-rescue profiles catalogued in Supplementary Figure S2c. Red dashed frames denote the Gene-only features, whereas green dashed frames encompass the genes together with the inferred HubmiRs available to the combined classifier. Values range from −1 to +1 , with zero denoting the corresponding outer-training-fold median; values for profiles rescued in multiple runs are averaged across their strict-rescue occurrences.
Predicting transcriptomic responses to small-molecule perturbations across cell lines is central to drug discovery, but exhaustive profiling of drug-cell combinations is infeasible. We frame molecular perturbation prediction as retrieve-and-aggregate: approximate an unmeasured drug's response in a cell line by aggregating measured responses of a small set of biologically related compounds. We propose LLM-Guided Retrieval (LGR), where a large language model (LLM) ranks candidate neighbor drugs (restricted to those profiled in the target cell line); after which a fixed mean aggregator combines their observed expression deltas to form the prediction. We evaluate on the Tahoe-100M single-cell perturbation atlas under unseen-drug, unseen-cell-line, and open-world regimes. LGR consistently improves over drug mean, ChemCPA, and chemistry-based kNN baselines, with the strongest gains for unseen cell-line generalization, where it achieves higher correlation and lower error than mean baselines. Across settings, LGR improves directional (sign) accuracy of gene regulation, indicating better recovery of biologically meaningful perturbation effects even when magnitude-based metrics are similar. These results suggest that retrieval quality, rather than predictor complexity, is a key driver of zero-shot molecular perturbation prediction, and that LLMs can provide a useful biological prior when used as constrained retrieval modules.
Department of Biomedical Data Science Stanford University Stanford, CA 94305, USA · Biology Research & AI Development, Genentech DNA Way, South San Francisco, CA 94080, USA
Predicting how a cell's transcriptome responds to a drug it has never seen is a core, hard problem in computational cell biology: recent benchmarks show complex models often fail to beat trivial baselines once test compounds are held out by chemistry. We study one cell line and assay, THP-1 cells profiled by DRUG-seq, scored by the active-compound weighted MSE(wMSE) of the VCPI prediction contest. We propose a staged approach: dumb baselines (untreated control and mean training-compound response) that the field keeps failing to beat; non-parametric retrieval (a Tanimoto-weighted average of a held-out compound's nearest training compounds); and a fusion stage combining a frozen chemistry embedding with retrieval-support features to predict the residual over the mean, with an uncertainty head and gene programs. On the released VCPI THP-1 drug-seq data (14,026 training compounds), under a Bemis-Murcko scaffold split, the model ranking inverts depending on the metric. Under an inverse-variance per-gene proxy, a regularized linear regression on Morgan fingerprints appears to win over the deep models, retrieval, and ChemBERTa -- the textbook "simple baselines win" result. But under the contest's true active-set metric (per-(gene, compound) Mejia weights, validated against the official scorer; mean baseline 0.535 vs the organizers' 0.507 reference), that reverses: the deep models win, our fusion decoder significantly beats the linear fingerprint baseline (-0.012 wMSE, paired bootstrap p < 10^-4), and the proxy's winner becomes the worst chemistry-aware predictor. Picking the metric picks the winner -- to our knowledge the first demonstration on real held-out drug chemistry of the metric-calibration effect established largely on genetic perturbation. We release a reproducible pipeline wired to the official scorer that emits a valid submission over the real 1064 x 12,995 grid.
Accurate prediction of patient-specific therapeutic response from pre-treatment transcriptomes is hindered by the scarcity of matched clinical response labels and post-treatment molecular profiles. Preclinical transfer-learning models can simulate drug-induced expression changes but are often hard to interpret and unstable, whereas knowledge-graph methods provide mechanistic context yet remain static and fail to capture drug-induced transcriptomic perturbation dynamics. We propose PREDIKTOR, a patient-centered multi-view framework that aligns a personalized network view with a transferable transcriptomic perturbation view to predict clinical drug response. For each patient, we construct an individualized gene regulatory network from tumor expression using DysRegNet and augment it with drug-target links from DrugBank; a graph neural encoder yields a drug-centric, mechanistically grounded embedding. In parallel, a frozen condition-specific gene-gene attention model pretrained on LINCS L1000 generates a simulated post-perturbation transcriptomic profile for the same patient-drug pair. We align the two views in a shared latent space via a CLIP-style contrastive objective with drug-context hard negatives, then concatenate the representations for end-to-end response classification. On TCGA, PREDIKTOR consistently outperforms state-of-the-art baselines under patient-, drug-, and tissue-split evaluations, and transfers zero-shot to the I-SPY2 trial, improving AUROC by 5.6% over competing methods. The aligned embeddings yield stable gene and pathway attributions that recover known mechanisms, supporting actionable and interpretable precision oncology.
Dongmin Bang, Sugyun An, Inyoung Sung +3
Interdisciplinary Program in Bioinformatics, Seoul National University, Seoul, Republic of Korea · AIGENDRUG Co., Ltd., Seoul, Republic of Korea · BK21 FOUR Intelligence Computing, Seoul National University, Seoul, Republic of Korea +2