ColdDDI: Evaluating Knowledge Utilization in Cold-Start Drug-Drug Interaction Prediction
Organizations: Data Science, William & Mary · Computer Science, Baylor University · Southwest University · Hefei University of Technology · Florida Atlantic University
Abstract
Cold-start drug-drug interaction (DDI) prediction tests whether models can identify clinically significant interactions for drugs without training-time interaction history. Existing benchmarks mostly report aggregate edge-prediction scores, leaving a key evaluation question unanswered: when models receive molecular, textual, or knowledge-graph (KG) evidence, do they actually use the evidence that pharmacologically supports the interaction? We introduce ColdDDI, a reconstructible diagnostic benchmark built from DrugBank 5.1.13, with 1,900 approved small-molecule drugs and 565,731 positive DDI pairs. ColdDDI evaluates pairs with zero, one, or two unseen drugs. It also annotates each interaction by whether it changes drug exposure or drug effect, and by whether the biomedical knowledge graph contains shared enzymes, transporters, or targets that can plausibly mediate the interaction. These annotations separate evidence availability from predictive dependence. We evaluate eight conventional DDI methods and 13 LLMs; for open-weight LLMs, we test five prompt patterns and use masking, drug replacement, and channel-sensitivity metrics to probe knowledge utilization. ColdDDI exposes that, in the hardest split where both drugs are unseen, the main performance divide is mediator availability. A fine-tuned 1B LLM recovers 89-93% of interactions with a shared enzyme, transporter, or target, but only 40-62% without such a mediator. More importantly, KG-provided evidence is not always used; several KG-augmented baselines change little when the shared mediator is masked or disrupted, whereas fine-tuned LLMs respond strongly to this intervention. Thus, ColdDDI evaluates knowledge utilization rather than knowledge access alone, showing where cold-start DDI models rely on mechanistic evidence and where they fail despite receiving it. Code is available at https://github.com/0217ljh/ColdDDI-NeurIPS2026.
Figures & tables
| Benchmark | Year | Source | Scale | D1 | D2 | D3 | D4 |
| (drugs / pairs) | Setting | Mechanism | Modality | Model | |||
| TWOSIDES [ 22 , 23 ] | 2012 | FAERS / SIDER | 645 / 63K | ✗ | ✗ | ✗ | ✗ |
| DeepDDI [ 1 ] | 2018 | DrugBank 5.0 | 1,710 / 192K | ✗ | ✗ | ✗ | ✗ |
| OGBL-DDI [ 24 ] | 2020 | DrugBank 5.0 | 4,267 / 1.3M † | ✗ | ✗ | ✗ | ✗ |
| TDC-DDI [ 25 ] | 2021 | DrugBank 5.0 | 1,706 / 192K | ✗ | ✗ | ✗ | ✗ |
| DDInter [ 26 ] | 2022 | FDA / Lit. | 1,972 / 237K | ✗ | ✗ | ✗ |
| S0 (Transductive) | S1 (Semi-Inductive) | S2 (Fully Inductive) | |
| Train (+) | |||
| Train ( ) | , resampled each epoch | ||
| Val/Test (+) | |||
| Val/Test ( ) | fixed from | fixed from | fixed from |
| AUC-ROC | Recall | ||||||
| Group | Method | S0 | S1 | S2 | S0 | S1 | S2 |
| Matrix | DeepDDI [ 1 ] | ||||||
| Mol-Graph | SSI-DDI [ 7 ] | ||||||
| DSN-DDI [ 8 ] † | |||||||
| HDN-DDI [ 9 ] | |||||||
| KG-only | EmerGNN [ 14 ] | ||||||
| Method | PK-A | PK-B | PD-A | PD-B | A–B gap | |||||
| Recall | AUC-ROC | Recall | AUC-ROC | Recall | AUC-ROC | Recall | AUC-ROC | Recall | AUC-ROC | |
| DeepDDI | ||||||||||
| SSI-DDI | ||||||||||
| DSN-DDI | ||||||||||
| HDN-DDI | ||||||||||
| EmerGNN | ||||||||||
| Group | Entity mask | Name mask | Both mask | KSAI | Dominant |
| (R0 R2) | (R0 R1) | (R0 R3) | channel | ||
| PK-A | Entity | ||||
| PK-B | Name | ||||
| PD-A | Entity | ||||
| PD-B | Name | ||||
| ALL |
| Group | Method | Absolute KPS-F | A–B gap |
| Matrix | DeepDDI | ||
| Mol-Graph | SSI-DDI | ||
| DSN-DDI | |||
| HDN-DDI | |||
| KG-only | EmerGNN | ||
| KG-integrated | TIGER |
| Intervention | PK-A | PK-B | PD-A | PD-B |
| Zero KG attention | ||||
| Zero name attention |
| Llama | Qwen 2.5 | Gemma 3 | Closed-source | |||||||
| Size | Inf | FT | Size | Inf | FT | Size | Inf | FT | Model | Inf |
| 3.2-1B | 0.525 | 0.776 | 0.5B | 0.489 | 0.766 | 1B | 0.520 | 0.756 | GPT-4o | 0.710 |
| 3.2-3B | 0.459 | 0.785 | 3B | 0.517 | 0.769 | 4B | 0.496 | 0.781 | Claude Sonnet 4.6 | 0.751 |
| 2-7B | 0.474 | 0.731 | 7B | 0.486 | 0.740 | 12B | 0.548 | 0.798 | ||
| 2-13B | 0.513 | 0.781 | 14B | 0.510 | 0.750 | |||||
Appendix figures & tables67 assets
Supplementary material from the paper’s appendix.
Appendix
| Step | Operation | Drugs | DDI Edges | DDI Types |
| 0 | DrugBank 5.1.13 (raw) | 17,430 | 1,427,655 | n/a |
| 1 | Retain small molecules | 13,166 | 1,205,013 | n/a |
| 2 | Require valid RDKit SMILES | 12,303 | 1,161,857 | n/a |
| 3 | Restrict to approved compounds | 2,643 | 613,514 | n/a |
| 4 | Extract DDI edges & normalize descriptions | 2,643 | 613,514 | 451 |
| 5 | Remove DDI types with 10 occurrences | 2,151 | 613,055 | 221 |
| Label | # DDI Types | # DDI Edges | Edge % |
| PK | 30 | 295,406 | 52.2% |
| PD | 184 | 270,241 | 47.8% |
| Mixed (resolved as PK) | 1 | 84 | 0.1% |
| Total | 215 | 565,731 | 100% |
| Group | # Edges | Edge % | # Unique Key Entities | Dominant Entity |
| PK-A | 167,718 | 29.6% | 79 (enzyme: 38, transporter: 41) | CYP3A4 (48.0% of PK-A) |
| PK-B | 127,772 | 22.6% | n/a | n/a |
| PD-A | 21,686 | 3.8% | 282 (target: 282) | Histamine H1 receptor (9.7% of PD-A) |
| PD-B | 248,555 | 43.9% | n/a | n/a |
| Task | Subset | Agreement | F1 (pos. 1) | F1 (pos. 2) | |
| PK/PD per-pair | all | 500 | 89.0% | 0.884 (PK) | 0.895 (PD) |
| PK/PD type-level | majority-vote types | 50 | 94.0% | 0.936 (PK) | 0.943 (PD) |
| A/B per-pair | overall | 500 | 93.8% | 0.893 (A) | 0.956 (B) |
| Task | Subset | Agreement / | |
| A/B per-quadrant (auto vs. consensus) | |||
| A/B per-pair | PK-A | 131 | 77.9% |
| A/B per-pair | PK-B | 119 | 100.0% |
| A/B per-pair | PD-A | 29 | 93.1% |
| A/B per-pair | PD-B | 221 | 100.0% |
| Inter-annotator (Cohen’s ) | |||
| Path | Contents | Paper anchor |
| Entry scripts (top level) | ||
| reconstruct.py | DrugBank XML filtered edges + KG + annotations + splits | Appendices A.1 and B.1 |
| evaluate.py | Conventional baseline training and evaluation | Section 5.1 |
| sanity_check.py | Checks data consistency and reconstruction checksums | Appendix A.6.3 |
| Library package | ||
| coldddi/data/filter.py | 7-step filtering pipeline | Appendix A.1 |
| Full 1,900-drug | 800-drug subset | |||||
| Quantity | Seed 42 | Seed 43 | Seed 44 | Seed 42 | Seed 43 | Seed 44 |
| Edge pool sizes | ||||||
| 361,485 | 366,641 | 360,330 | 59,715 | 62,861 | 61,554 | |
| 181,571 | 177,769 | 182,380 | 30,515 | 29,867 | 32,448 | |
| 22,675 | 21,321 | 23,021 | 3,837 | 3,450 | 4,188 | |
| Per-split positive counts | ||||||
| Tanimoto cutoff | Train–Test | Train–Train |
| Method | 1:1 | 1:3 | 1:5 |
| DeepDDI | 0.678 | 0.680 | 0.674 |
| SSI-DDI | 0.656 | 0.646 | 0.642 |
| DSN-DDI | 0.814 | 0.755 | 0.793 |
| HDN-DDI | 0.675 | 0.668 | 0.667 |
| EmerGNN | 0.724 | 0.714 | 0.706 |
| TIGER | 0.581 | 0.589 | 0.557 |
| Method | Random | Hard | Structure-matched |
| DeepDDI | 0.678 | 0.558 | 0.672 |
| SSI-DDI | 0.656 | 0.581 | 0.645 |
| DSN-DDI | 0.814 | 0.718 | 0.770 |
| HDN-DDI | 0.675 | 0.623 | 0.621 |
| EmerGNN | 0.724 | 0.619 | 0.723 |
| TIGER | 0.581 | 0.572 | 0.569 |
| Statistic | Full (1,900) | Seed 42 | Seed 43 | Seed 44 |
| Total drugs | 1,900 | 800 | 800 | 800 |
| (seen) | 1,520 | 640 | 640 | 640 |
| (unseen) | 380 | 160 | 160 | 160 |
| Total positive DDI pairs | 565,731 | 94,067 | 96,178 | 98,190 |
| DDI types covered | 215 | 165 | 152 | 165 |
| Train positive | 325,336 | 53,743 | 56,574 | 55,398 |
| Method | AUC | Recall | F1 |
| DeepDDI | |||
| SSI-DDI | |||
| DSN-DDI | |||
| HDN-DDI | |||
| EmerGNN | |||
| TIGER |
| Method | S0 | S1 | S2 | Mean |
| DeepDDI | 0.011 | 0.005 | 0.005 | 0.007 |
| SSI-DDI | 0.033 | 0.021 | 0.031 | 0.028 |
| DSN-DDI | 0.081 | 0.141 | 0.190 | 0.137 |
| HDN-DDI | 0.041 | 0.007 | 0.032 | 0.026 |
| EmerGNN | 0.001 | 0.005 | 0.009 | 0.005 |
| TIGER | 0.007 | 0.047 | 0.053 | 0.036 |
| Drugs | Split | AUROC | AUPRC | Recall | Accuracy | Precision | F1 | |
| 100 | S0 | 0.823 | 0.817 | 0.630 | 0.717 | 0.763 | 0.690 | 46 |
| 100 | S1 | 0.654 | 0.620 | 0.568 | 0.616 | 0.628 | 0.597 | 241 |
| 100 | S2 | 0.541 | 0.542 | 0.138 | 0.500 | 0.500 | 0.216 | 29 |
| 200 | S0 | 0.885 | 0.868 | 0.801 | 0.804 | 0.805 | 0.803 | 186 |
| 200 | S1 | 0.762 | 0.735 | 0.752 | 0.691 | 0.670 | 0.709 | 975 |
| 200 | S2 | 0.654 | 0.662 | 0.702 | 0.621 | 0.604 | 0.649 | 124 |
| Split | Top-1 | Top-3 | Top-5 | Macro-F1 | Macro-AUROC | Macro-AUPRC | |
| S0 | 3,069 | ||||||
| S1 | 15,472 | ||||||
| S2 | 1,913 |
| Method | Optim | LR | Weight decay | Batch | Epochs | Patience | Key arch dims | Params | Wall-clock (5090) | |
| 800 | 1,900 | |||||||||
| DeepDDI | Adam | 256 | 100 | 15 | SSP=50, hidden=2048, layers=9, dropout=0.3 | 29.6M | 10 min | 1 h | ||
| SSI-DDI | Adam | 1024 | 150 | 50 | atom feats=55, KGE dim=64 | 0.7M | 50 min | 10 h | ||
| DSN-DDI | Adam | 512 | 50 | 10 | hidden=128, KGE dim=128, dropout=0.2 | 1.4M | 1 h | 14 h | ||
| HDN-DDI | Adam | 512 | 50 | 10 | hidden=128, KGE dim=128 | 1.6M | 1 h | 14 h | ||
| EmerGNN | Adam | 32 | 40 | 10 | =64, length=3, feat=Morgan FP | 0.5M | 40 min | 10 h | ||
| Method | S0 | S1 | S2 | |||
| AUC-ROC | AUC-PRC | AUC-ROC | AUC-PRC | AUC-ROC | AUC-PRC | |
| DeepDDI | 0.988 0.003 | 0.989 0.002 | 0.814 0.008 | 0.797 0.014 | 0.659 0.018 | 0.646 0.011 |
| SSI-DDI | 0.876 0.026 | 0.871 0.027 | 0.716 0.015 | 0.699 0.011 | 0.614 0.037 | 0.610 0.025 |
| DSN-DDI | 0.981 0.004 | 0.980 0.005 | 0.876 0.013 | 0.876 0.008 | 0.758 0.052 | 0.747 0.061 |
| HDN-DDI | 0.909 0.011 | 0.908 0.015 | 0.741 0.005 | 0.721 0.008 | 0.634 0.035 | 0.628 0.041 |
| EmerGNN | 0.982 0.001 | 0.984 0.002 | 0.805 0.012 | 0.793 0.024 | 0.716 0.020 | 0.701 0.032 |
| Method | S0 | S1 | S2 | |||
| F1 | Acc | F1 | Acc | F1 | Acc | |
| DeepDDI | 0.941 0.007 | 0.943 0.006 | 0.721 0.005 | 0.735 0.004 | 0.555 0.013 | 0.611 0.005 |
| SSI-DDI | 0.801 0.025 | 0.798 0.026 | 0.656 0.004 | 0.659 0.012 | 0.548 0.079 | 0.585 0.037 |
| DSN-DDI | 0.929 0.004 | 0.928 0.005 | 0.764 0.034 | 0.784 0.018 | 0.617 0.019 | 0.674 0.029 |
| HDN-DDI | 0.828 0.011 | 0.826 0.008 | 0.670 0.016 | 0.680 0.005 | 0.606 0.035 | 0.590 0.037 |
| EmerGNN | 0.932 0.002 | 0.933 0.002 | 0.714 0.022 | 0.735 0.015 | 0.614 0.039 | 0.659 0.022 |
| Method | S0 | S1 | S2 | |||
| AUC-ROC | AUC-PRC | AUC-ROC | AUC-PRC | AUC-ROC | AUC-PRC | |
| DeepDDI | 0.997 0.000 | 0.997 0.000 | 0.819 0.003 | 0.804 0.008 | 0.672 0.009 | 0.663 0.016 |
| SSI-DDI | 0.829 0.010 | 0.819 0.013 | 0.706 0.006 | 0.688 0.006 | 0.627 0.003 | 0.624 0.014 |
| DSN-DDI | 0.950 0.038 | 0.947 0.050 | 0.827 0.056 | 0.818 0.065 | 0.711 0.063 | 0.696 0.062 |
| HDN-DDI | 0.853 0.006 | 0.846 0.010 | 0.736 0.003 | 0.718 0.003 | 0.645 0.002 | 0.638 0.009 |
| EmerGNN | 0.980 0.001 | 0.982 0.001 | 0.799 0.000 | 0.776 0.007 | 0.707 0.006 | 0.680 0.012 |
| Method | S0 | S1 | S2 | |||
| F1 | Acc | F1 | Acc | F1 | Acc | |
| DeepDDI | 0.973 0.000 | 0.973 0.000 | 0.726 0.007 | 0.743 0.004 | 0.575 0.019 | 0.624 0.006 |
| SSI-DDI | 0.756 0.012 | 0.751 0.013 | 0.643 0.011 | 0.652 0.007 | 0.587 0.021 | 0.593 0.002 |
| DSN-DDI | 0.887 0.055 | 0.879 0.064 | 0.742 0.041 | 0.744 0.059 | 0.667 0.076 | 0.656 0.061 |
| HDN-DDI | 0.775 0.011 | 0.772 0.007 | 0.669 0.001 | 0.675 0.002 | 0.597 0.018 | 0.603 0.002 |
| EmerGNN | 0.932 0.004 | 0.932 0.004 | 0.728 0.002 | 0.734 0.001 | 0.644 0.011 | 0.657 0.007 |
| Type | Tier | DeepDDI | SSI-DDI | DSN-DDI | HDN-DDI | EmerGNN | TIGER | MKG-FENN | TextDDI | |
| PK-A | Low | 183 | ||||||||
| PK-A | Mid | 250 | ||||||||
| PK-A | High | 255 | ||||||||
| PK-B | Low | 157 | ||||||||
| PK-B | Mid | 129 | ||||||||
| PK-B | High | 95 |
| Metric | LLM-FT (P4) | EmerGNN | ||||
| C1 | C2 | C3 | C1 | C2 | C3 | |
| Recall | ||||||
| AUC-ROC | ||||||
| AUPRC | ||||||
| MCC | ||||||
| F1 | ||||||
| Method | Recall gap | 95% CI |
| DeepDDI | ||
| SSI-DDI | ||
| DSN-DDI | ||
| HDN-DDI | ||
| EmerGNN | ||
| TIGER |
| Added source | Edge increase (%) | Drugs | Recall gap | |||
| Enz. | Tgt. | Trn. | LLM-FT | EmerGNN | ||
| None | ||||||
| PharmGKB | ||||||
| ChEMBL | ||||||
| KEGG | ||||||
| All three | ||||||
| Training-frequency bin | Llama-3.2-1B | Gemma-3-12B | EmerGNN | |
| Shared mediator: | ||||
| – | ||||
| – | ||||
| – | ||||
| – | ||||
| Parameter | Search space | Best |
| learning rate | (step ) | |
| dropout | (step ) | |
| target modules | ||
| batch size |
| Parameter | Value |
| Adapter (PEFT/LoRA) | |
| rank | 16 |
| 16 | |
| dropout | 0.10 |
| target modules | q_proj, k_proj, v_proj, o_proj |
| layers transformed | 11 (last 11 transformer blocks) |
| Model | P1 | P2 | P3 | P4 | P5 |
| Llama-3.2-1B | 0.488 | 0.523 | 0.465 | 0.525 | 0.485 |
| Llama-3.2-3B | 0.504 | 0.488 | 0.457 | 0.459 | 0.468 |
| Llama-2-7B | 0.535 | 0.489 | 0.549 | 0.474 | 0.496 |
| Llama-2-13B | 0.545 | 0.533 | 0.450 | 0.513 | 0.479 |
| Qwen2.5-0.5B | 0.525 | 0.473 | 0.511 | 0.489 | 0.513 |
| Qwen2.5-3B | 0.502 | 0.511 | 0.555 | 0.517 | 0.579 |
| Model | P1 | P2 | P3 | P4 | P5 |
| Llama-3.2-1B | 0.000 | 0.078 | 0.850 | 0.857 | 0.102 |
| Llama-3.2-3B | 0.000 | 0.003 | 0.000 | 0.000 | 0.000 |
| Llama-2-7B | 0.000 | 0.008 | 0.000 | 0.000 | 0.000 |
| Llama-2-13B | 1.000 | 0.411 | 0.964 | 0.995 | 1.000 |
| Qwen2.5-0.5B | 1.000 | 0.980 | 1.000 | 1.000 | 1.000 |
| Qwen2.5-3B | 0.999 | 0.631 | 1.000 | 1.000 | 0.000 |
| Model | AUC | Acc. | Prec. | Rec. | F1 |
| GPT-4o | 0.710 | 0.666 | 0.664 | 0.672 | 0.668 |
| Claude Sonnet 4.6 | 0.751 | 0.617 | 0.886 | 0.267 | 0.411 |
| Model | P1 | P2 | P3 | P4 | P5 |
| AUC-ROC | |||||
| Llama-3.2-1B | 0.982 0.003 | 0.982 0.005 | 0.987 0.002 | 0.987 0.001 | 0.980 0.003 |
| Llama-3.2-3B | 0.987 0.002 | 0.965 0.036 | 0.946 0.071 | 0.991 0.001 | 0.984 0.003 |
| Llama-2-7B | 0.883 0.055 | 0.975 0.014 | 0.962 0.023 | 0.893 0.081 | 0.967 0.031 |
| Llama-2-13B | 0.891 0.059 | 0.930 0.007 | 0.915 0.071 | 0.903 0.044 | 0.907 0.027 |
| Qwen2.5-0.5B | 0.972 0.003 | 0.974 0.003 | 0.983 0.004 | 0.979 0.007 | 0.972 0.004 |
| Model | P1 | P2 | P3 | P4 | P5 |
| AUC-ROC | |||||
| Llama-3.2-1B | 0.806 0.014 | 0.804 0.024 | 0.836 0.001 | 0.843 0.009 | 0.798 0.004 |
| Llama-3.2-3B | 0.819 0.021 | 0.807 0.012 | 0.827 0.010 | 0.838 0.014 | 0.828 0.005 |
| Llama-2-7B | 0.789 0.025 | 0.797 0.006 | 0.827 0.008 | 0.771 0.067 | 0.812 0.011 |
| Llama-2-13B | 0.804 0.017 | 0.827 0.010 | 0.839 0.024 | 0.836 0.019 | 0.811 0.007 |
| Qwen2.5-0.5B | 0.779 0.012 | 0.788 0.019 | 0.841 0.007 | 0.839 0.016 | 0.788 0.016 |
| Model | P1 | P2 | P3 | P4 | P5 |
| AUC-ROC | |||||
| Llama-3.2-1B | 0.685 0.027 | 0.696 0.008 | 0.772 0.013 | 0.776 0.020 | 0.709 0.025 |
| Llama-3.2-3B | 0.725 0.030 | 0.729 0.025 | 0.773 0.004 | 0.785 0.018 | 0.737 0.023 |
| Llama-2-7B | 0.700 0.015 | 0.704 0.009 | 0.762 0.010 | 0.731 0.039 | 0.716 0.023 |
| Llama-2-13B | 0.721 0.022 | 0.731 0.005 | 0.781 0.017 | 0.781 0.015 | 0.726 0.017 |
| Qwen2.5-0.5B | 0.673 0.012 | 0.675 0.022 | 0.759 0.018 | 0.766 0.013 | 0.693 0.021 |
| Setting | AUC-ROC | Accuracy | F1 | Precision | Recall |
| S0 | |||||
| S1 | |||||
| S2 |
| Method / Prompt | AUC-ROC | Recall |
| TextDDI (RoBERTa) | ||
| Llama-3.2-1B, P1 (names) | ||
| Llama-3.2-1B, P7 (names + descriptions) | ||
| Llama-3.2-1B, P4 (KG context) | ||
| Llama-3.2-1B, P6 (KG context + descriptions) |
| Type | Tier | Llama-3.2-1B | Llama-3.2-3B | Llama-2-7B | Llama-2-13B | |
| PK-A | Low | 183 | 0.838 0.113 | 0.701 0.191 | 0.776 0.157 | 0.731 0.119 |
| PK-A | Mid | 250 | 0.908 0.082 | 0.716 0.173 | 0.818 0.125 | 0.731 0.148 |
| PK-A | High | 255 | 0.911 0.076 | 0.782 0.141 | 0.829 0.134 | 0.740 0.128 |
| PK-B | Low | 157 | 0.305 0.108 | 0.240 0.049 | 0.335 0.079 | 0.265 0.055 |
| PK-B | Mid | 129 | 0.433 0.168 | 0.352 0.132 | 0.436 0.088 | 0.351 0.097 |
| PK-B | High | 95 | 0.511 0.151 | 0.379 0.121 | 0.506 0.134 | 0.429 0.149 |
| Type | Tier | Qwen2.5-0.5B | Qwen2.5-3B | Qwen2.5-7B | Qwen2.5-14B | |
| PK-A | Low | 183 | 0.820 0.070 | 0.708 0.076 | 0.726 0.186 | 0.794 0.111 |
| PK-A | Mid | 250 | 0.861 0.072 | 0.766 0.104 | 0.821 0.121 | 0.863 0.055 |
| PK-A | High | 255 | 0.863 0.054 | 0.780 0.073 | 0.838 0.099 | 0.879 0.054 |
| PK-B | Low | 157 | 0.310 0.075 | 0.296 0.072 | 0.281 0.143 | 0.318 0.126 |
| PK-B | Mid | 129 | 0.488 0.054 | 0.396 0.069 | 0.418 0.150 | 0.438 0.138 |
| PK-B | High | 95 | 0.438 0.016 | 0.420 0.100 | 0.491 0.189 | 0.511 0.138 |
| Type | Tier | Gemma-3-1B | Gemma-3-4B | Gemma-3-12B | |
| PK-A | Low | 183 | 0.643 0.255 | 0.753 0.133 | 0.669 0.108 |
| PK-A | Mid | 250 | 0.689 0.235 | 0.760 0.121 | 0.733 0.089 |
| PK-A | High | 255 | 0.709 0.212 | 0.799 0.050 | 0.776 0.072 |
| PK-B | Low | 157 | 0.262 0.126 | 0.275 0.095 | 0.236 0.060 |
| PK-B | Mid | 129 | 0.380 0.119 | 0.358 0.110 | 0.367 0.067 |
| PK-B | High | 95 | 0.444 0.142 | 0.378 0.128 | 0.407 0.070 |
| Setup | Agreement | Pearson | AUC | AUC | ||
| P1 (zero-shot) | 556 | |||||
| P4 (one-hop KG) | 556 |
| Approval bucket | Drugs | pos | Recall | AUC |
| Pre-2000 | 388 | 1240 | ||
| 2000–2009 | 103 | 138 | ||
| 2010–2019 | 127 | 134 | ||
| 2020+ | 62 | 20 | ||
| Unknown | 120 | 381 | ||
| ALL | 800 | 1913 |
| Cutoff | AUROC | F1 | ||
| 1985 | 237 / 443 | 31,358 | ||
| 1995 | 335 / 345 | 19,106 | ||
| 2005 | 436 / 244 | 9,396 | ||
| 2010 | 491 / 189 | 5,956 |
| ID | Method directory | KG | Drug name | Key entity | Token cap |
| R0 | One_Hop_KG_Sequence | Top-3 | real | real | 1250 |
| R1 | OHKS_Mask_Name | Top-3 | [DRUG_*] | real | 1250 |
| R2 | OHKS_Mask_Entity | Top-3 | real | [ENTITY] | 1250 |
| R3 | OHKS_Mask_Name_Entity | Top-3 | [DRUG_*] | [ENTITY] | 1250 |
| R4 | OHKS_Full | Full | real | real | 4096 |
| R5 | OHKS_Full_Mask_Name | Full | [DRUG_*] | real | 4096 |
| Bucket | Top-3 KG | Full KG | ||||||
| R0 | R1 | R2 | R3 | R4 | R5 | R6 | R7 | |
| PK-A | ||||||||
| PK-B | ||||||||
| PD-A | ||||||||
| PD-B | ||||||||
| ALL | ||||||||
| Method | PK-A | PK-B | PD-A | PD-B | ALL | A–B gap |
| DeepDDI | ||||||
| SSI-DDI | ||||||
| DSN-DDI | ||||||
| HDN-DDI | ||||||
| EmerGNN | ||||||
| TIGER |
| Indicator | PK-A | PK-B | PD-A | PD-B | ALL |
| KPS-Name (KG present) | |||||
| KPS-KG (Name present) | |||||
| KPS-KG (Name masked) | |||||
| KSAI-masking |
| Method | Channel | PK-A | PK-B | PD-A | PD-B | ALL |
| MKG-FENN | KPS-mol | |||||
| KPS-KG | ||||||
| TIGER | KPS-mol | |||||
| KPS-KG |
| Function | Head | KG attn | Name attn | Layer |
| KG-focused | L7 H3 | 60.1% | 12.5% | Mid (7) |
| KG-focused | L7 H8 | 50.8% | 14.5% | Mid (7) |
| KG-focused | L7 H11 | 53.5% | 17.2% | Mid (7) |
| Name-focused | L14 H21 | 18.7% | 70.4% | Deep (14) |
| Name-focused | L13 H5 | 24.5% | 33.0% | Deep (13) |
| Name-focused | L10 H23 | 21.9% | 17.1% | Mid-deep (10) |
| Condition | ALL | PK-A | PK-B | PD-A | PD-B |
| Baseline | 0.798 | 0.939 | 0.690 | 0.930 | 0.733 |
| Suppress KG heads (3) | 0.776 | 0.920 | 0.689 | 0.875 | 0.702 |
| Suppress Name heads (3) | 0.799 | 0.948 | 0.685 | 0.937 | 0.730 |
| Suppress ALL KG | 0.673 | 0.720 | 0.630 | 0.739 | 0.652 |
| Suppress ALL Name | 0.774 | 0.939 | 0.650 | 0.928 | 0.697 |
| Suppress L7 H3 only | 0.786 | 0.945 | 0.690 | 0.898 | 0.704 |
| Cat. | Method | Best lr | ALL | PK-A | PK-B | PD-A | PD-B | ALL |
| — | Baseline (LoRA-FT) | — | 0.798 | 0.939 | 0.690 | 0.930 | 0.733 | — |
| A | SAB | 5e-4 | 0.799 | 0.934 | 0.688 | 0.935 | 0.739 | 0.1 |
| A | PHT | 1e-3 | 0.797 | 0.927 | 0.686 | 0.931 | 0.739 | 0.2 |
| A | THB | 1e-3 | 0.799 | 0.939 | 0.690 | 0.930 | 0.734 | 0.0 |
| A | ISB | 1e-2 | 0.799 | 0.937 | 0.687 | 0.935 | 0.737 | 0.1 |
| B | IAD ( ) | — | 0.796 | 0.939 | 0.688 | 0.925 | 0.730 | 0.2 |
| Condition | ALL | PK-A | PK-B | PD-A | PD-B |
| Llama-3.2-1B: AUROC | |||||
| Baseline | |||||
| Suppress KG heads (3) | |||||
| Suppress Name heads (3) | |||||
| Suppress ALL KG | |||||
| Suppress ALL Name | |||||