Protein-molecule virtual screening is increasingly cast as a problem of representation learning in a shared embedding space. Existing methods rely on dense holistic alignment, entangling invariant binding determinants with nuisance correlations and limiting transfer to new targets. It has been noted that binding in protein-molecule systems involves sparse cross-modality interactions: binding is governed by a small contact interface and a few decisive local interactions (e.g., hydrogen bonds, hydrophobic contacts, and salt bridges) rather than the global structures of the protein and molecule. We hypothesize that uncovering and leveraging sparse interaction patterns is critical for generalization beyond the training data, as these patterns are reusable and expected to improve performance across different scenarios. In this paper, we aim to identify and leverage sparse interaction patterns, and verify our hypothesis. Since the training data contain only observed binding pairs, we formalize this prior via a V-structure causal model under Heckman-style selection, and establish three theoretical results: (i) the latent concepts of interacting proteins and molecules are not identifiable without appropriate sparsity constraints; (ii) these concepts and their sparse interactions are component-wise identifiable under structural sparsity conditions; and (iii) a low-rank relaxation of these conditions yields subspace identifiability of the concepts and interactions. Inspired by these principles, we propose CausalBind with three implementation variants. Extensive experiments on DUD-E and LIT-PCBA benchmarks show that all variants consistently outperform strong retrieval baselines, with the largest gains on LIT-PCBA early enrichment, and further generalize to target- and scaffold-level out-of-distribution splits. Code is available at https://github.com/lokali/CausalBind.
Figures & tables
Figure 1: Dense alignment versus sparse binding. (a) Holistic dense alignment. Existing retrieval-based virtual screening methods compare every protein surface feature with every molecule feature in the shared embedding space; this “full-surface” alignment entangles binding-relevant signals with nuisance correlations. (b) Biology-inspired sparse binding. Protein-molecule binding is governed by a small contact interface and a handful of key local interactions (e.g., hydrogen bonds, hydrophobic contacts, salt bridges, and π -stacking), which together account for the majority of the binding affinity.
Figure 2: Binder-selected sparse-interaction view. Observing only binders ( B=1 ) induces cross-modal dependence through the collider B . Green edges denote active local interactions. Red dashed edges blocked interactions. Purple dashed curves allowed within-modality dependence.
Figure 3: Implementation-level view of CausalBind . Each modality-specific encoder maps its input to atomic concepts {cvk} , each with dimension dc≥1 . The indicators Bij at the bottom visualize the branch-wise cross-modal mask Ma,b for (a,b)∈{(mol,poc),(mol,seq)} : Bij=1 when the (i,j) entry of Ma,b survives the forward threshold ( τ=0.5 ) (green solid edges, active pair), and Bij=0 otherwise (red dashed edges, blocked pair). Active pairs form the contrastive score, while Lspa encourages the interaction sparsity.
Method
DUD-E ( n=102 )
LIT-PCBA ( n=15 )
AUROC
BEDROC 80.5
EF 1%
AUROC
BEDROC 80.5
EF 1%
Glide-SP ( Friesner et al., 2004 )
0.767
0.407
16.18
0.532
0.040
3.41
DeepDTA ( Öztürk et al., 2018 )
0.584
0.051
2.28
0.563
0.025
1.47
Gnina ( McNutt et al., 2021 )
0.782
0.299
17.73
0.609
0.054
4.63
RTMScore ( Shen et al., 2022 )
0.753
0.434
27.10
0.525
0.039
2.94
TankBind ( Lu et al., 2022 )
0.751
0.330
13.00
0.597
0.039
2.90
Table 1: Virtual-screening comparison on two benchmarks between baselines and CausalBind SP ( K=128,dc=1024 ), CausalBind LR ( K=64,r=1 ), and CausalBind EMB ( K=64,dc=256 ), and CausalBind ATOM ( K=128,dc=1024 ). † S 2 Drug results are paper-reported ( He et al., 2026 ) rather than reproduced in our pipeline.
Method
DUD-E ( n=102 )
LIT-PCBA ( n=15 )
AUROC
BEDROC 80.5
EF 1%
AUROC
BEDROC 80.5
EF 1%
LigUnity ( Feng et al., 2025 )
0.896 ± 0.005
0.660 ± 0.018
43.14 ± 1.42
0.586 ± 0.010
0.077 ± 0.001
6.46 ± 0.18
HypSeek ( Wang et al., 2026 )
0.914 ± 0.013
0.690 ± 0.081
44.12 ± 6.14
0.603 ± 0.009
0.067 ± 0.007
5.62 ± 0.71
CausalBind SP
0.939 ± 0.004
0.746 ± 0.004
47.99 ± 0.39
0.635 ± 0.006
0.088 ± 0.007
7.55 ± 0.73
CausalBind SP −Lspa
0.935 ± 0.004
0.732 ± 0.014
46.93 ± 1.07
0.623 ± 0.005
0.084 ± 0.004
7.17 ± 0.36
CausalBind LR
0.937 ± 0.004
0.732 ± 0.011
46.87 ± 0.91
0.630 ± 0.006
0.085 ± 0.006
7.22 ± 0.31
Table 2: Ablation and diagnostic study on DUD-E and LIT-PCBA. CausalBind SP is the sparse mask, CausalBind LR is the low-rank mask, and CausalBind EMB is the pooled embedding-space mask. −Lspa removes the sparsity penalty from CausalBind SP , while +Lspa tests an optional sparsity penalty on CausalBind LR without changing its rank-one parameterization. Membr=1 replaces the pooled embedding mask with a rank-one mask. Rows report mean ±std over five runs.
Figure 4: Hyperparameter sweeps for CausalBind SP on DUD-E and LIT-PCBA. Solid lines show BEDROC 80.5 and dashed lines show EF@1%; shaded bands mark the selected hyperparameters.
λspa
DUD-E
LIT-PCBA
Cross-modal mask statistics
AUC
BEDROC 80.5
EF@1%
AUC
BEDROC 80.5
EF@1%
L1 avg
%zeros
r@95%
r@99%
Anti. viol.
0
0.938
0.746
48.14
0.617
0.080
6.83
0.93
0.0
1
25.0
32,512
10−6
0.940
0.749
47.94
0.635
0.088
7.26
0.92
0.0
1
25.0
32,512
10−5
0.938
0.735
47.09
0.607
0.078
6.29
0.88
0.0
1
26.0
32,512
10−4
0.935
0.744
47.69
0.639
0.097
8.56
0.61
0.0
1
34.0
32,512
10−3
0.943
0.768
49.73
0.614
0.073
5.49
0.21
20.5
45.5
79.5
0
Table 3: λspa -scan for CausalBind SP with paired cross-modal mask statistics, averaged over the two branches. Small λspa preserves a low-rank regime, while large λspa drives hard-zero collapse. ‘Anti. viol.’ counts antichain violations, i.e., pairs of row (or column) supports in which one contains the other, over 32,512 pairwise tests.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Drug discovery pipeline funnel. Each stage progressively filters compounds, with numbers indicating typical compound counts. Our CausalBind method (green box) operates at the virtual screening stage, improving the quality of top-ranked candidates before expensive downstream validation. By identifying invariant binding determinants and suppressing spurious correlations, our method aims to increase the hit rate in subsequent experimental testing.
Physics-based docking
Structure-aware ML scoring
Retrieval-based learning
Inference input
(protein, ligand) pair + pose search
(protein, ligand) pair, often pose-conditioned
protein and ligand encoded independently
Binding score
physics + empirical scoring on the pair
one forward pass through a joint network on the pair
dot product (or cosine) of two precomputed vectors
Pre-computable embeddings?
No (pose-coupled)
No (pair-coupled)
Yes (dual-tower)
Cost per (P, L) at inference
seconds–minutes (pose search dominates)
∼ 10–100 ms (one full forward)
sub-millisecond (one dot product)
Practical library scale
104 – 105
105 – 106
108–109
Where the inductive bias lives
force field + sampling protocol
complex graph / 3D-voxel architecture
shape of the shared embedding space and the similarity used on it
Appendix
Table A1: Core differences between the three virtual-screening regimes. The retrieval regime decouples encoders into two towers, enabling pre-computed library embeddings and millisecond-scale similarity scoring; the other two regimes condition on the (protein, ligand) pair at inference. CausalBind SP stays in the retrieval regime but replaces the dense holistic similarity with a sparse bipartite concept-pair interaction.
Method
Core Design
Relevance Here
Physics-based docking systems
Glide-SP [ Friesner et al., 2004 ]
Hierarchical sampling with empirical scoring function
Industry standard; accurate on small candidate sets but expensive at library scale
AutoDock Vina [ Trott and Olson, 2010 ]
Empirical scoring + iterated local search
Open-source, widely used; not designed for ultra-large libraries
Structure-aware machine-learning scoring models
DeepDTA [ Öztürk et al., 2018 ]
Sequence/graph-based affinity prediction without pairwise retrieval design
Useful DTA baseline; not optimized for target-wise early enrichment
Gnina [ McNutt et al., 2021 ]
3D CNN scoring on docking-grid voxels
Strong pose-rescoring; depends on docked poses
Appendix
Table A2: Representative methods in the three virtual-screening regimes discussed in the main text. The table highlights the main design choice of each method and the role it plays in the comparison, with CausalBind SP differing from prior retrieval baselines by imposing a sparse bipartite interaction prior rather than a dense holistic similarity.
Figure A2: Three-view concept-level schematic of CausalBind . The inputs are pocket structure, molecule structure, and protein sequence. Each view is decomposed into modality-specific concepts cpoc , cmol , and cseq . Local binary indicators summarize whether a concept pair is active ( 1 ) or blocked ( 0 ). We show two interaction branches, pocket-molecule and molecule-sequence, which are then aggregated into the final pair-level binding event B . Green solid arrows denote active local interactions, red dashed arrows denote blocked interactions, and gray arrows denote aggregation to the overall binding variable.
Variant
Mask object
Forward bottleneck
Closest theory link
Role in the paper
CausalBind SP
Dense K×K concept-pair mask with L1 sparsity and hard thresholding.
Concept axis before pooling; each active entry selects a protein-molecule concept pair.
Direct sparse-support form of Theorem 2 ; learned masks can also enter the low-rank regime of Theorem 3 .
Primary method in the main text.
CausalBind LR
Rank- r concept-basis mask UV⊤ with U,V∈RK×r .
Concept axis before pooling; rank controls the number of interaction subspaces.
Direct implementation of the low-rank subspace-identifiability structure in Theorem 3 .
Appendix variant for explicit low-rank implementation.
CausalBind EMB
Dense dc×dc mask in pooled embedding space.
After mean-pooling; the mask acts on pooled features rather than concept pairs.
Compatible with a compressed interaction view, but not a direct instantiation of concept-pair sparse support.
Appendix alternative and robustness check.
Appendix
Table A3: Comparison of the three CausalBind implementations. The main paper presents CausalBind SP ; CausalBind LR and CausalBind EMB are appendix variants.
Method
Protocol
DUD-E
LIT-PCBA
AUROC
BEDROC 80.5
EF@1%
AUROC
BEDROC 80.5
EF@1%
LigUnity
last
0.897
0.674
44.20
0.599
0.075
6.50
best_auc
0.915
0.716
46.42
0.581
0.089
8.15
best_bedroc
0.896
0.671
44.00
0.599
0.076
6.38
HypSeek
last
0.909
0.605
37.83
0.603
0.061
4.62
best_auc
0.913
0.629
39.63
0.608
0.063
5.13
Appendix
Table A4: Three-protocol evaluation for main comparisons. The main text reports the fixed-endpoint last protocol; best_auc and best_bedroc are complementary validation-selected references. All rows are seed-1 results.
K
Protocol
DUD-E
LIT-PCBA
AUROC
BEDROC 80.5
EF@1%
AUROC
BEDROC 80.5
EF@1%
32
last
0.938
0.749
48.41
0.620
0.081
6.90
best_auc
0.939
0.750
48.38
0.625
0.083
6.85
best_bedroc
0.938
0.750
48.57
0.621
0.081
6.90
64
last
0.930
0.724
46.55
0.637
0.085
6.41
best_auc
0.933
0.715
45.90
0.638
0.088
7.88
Appendix
Table A5: CausalBind SP concept-level K -scan (fix dc=512 ), seed=1, all three protocols. K=256 underperforms consistently.
dc
Protocol
DUD-E
LIT-PCBA
AUROC
BEDROC 80.5
EF@1%
AUROC
BEDROC 80.5
EF@1%
128
last
0.919
0.589
35.75
0.615
0.062
5.25
best_auc
0.922
0.629
39.43
0.630
0.060
4.75
best_bedroc
0.920
0.592
36.03
0.620
0.062
5.35
256
last
0.912
0.600
36.93
0.614
0.062
5.68
best_auc
0.922
0.683
43.07
0.620
0.067
4.73
Appendix
Table A6: CausalBind SP concept-level dc -scan (fix K=128 ), seed=1, all three protocols. dc=1024 gives the strongest LIT-PCBA results.
λ
Protocol
DUD-E
LIT-PCBA
AUROC
BEDROC 80.5
EF@1%
AUROC
BEDROC 80.5
EF@1%
0
last
0.938
0.746
48.14
0.617
0.080
6.83
best_auc
0.943
0.744
47.68
0.629
0.073
6.26
best_bedroc
0.934
0.732
47.03
0.613
0.075
6.10
10−6
last
0.940
0.749
47.94
0.635
0.088
7.26
best_auc
0.941
0.750
47.75
0.628
0.086
7.41
Appendix
Table A7: CausalBind SP concept-level λ -scan, seed=1, all three protocols. K=128 , dc=1024 , L=4 . Rows below the dotted rule show mask collapse under last and best_bedroc .
L
Protocol
DUD-E
LIT-PCBA
AUROC
BEDROC 80.5
EF@1%
AUROC
BEDROC 80.5
EF@1%
2
last
0.940
0.763
49.20
0.625
0.087
6.98
best_auc
0.946
0.776
49.99
0.616
0.082
6.53
best_bedroc
0.942
0.757
48.48
0.622
0.085
7.95
4 (default)
last
0.935
0.744
47.69
0.639
0.097
8.56
best_auc
0.934
0.722
46.41
0.628
0.074
5.84
Appendix
Table A8: CausalBind SP concept-level Perceiver-depth ( L ) scan, seed=1, all three protocols. K=128 , dc=1024 , λ=10−4 .
λ
branch
σ(M)
max σ(M)
L1 (sum)
zeros
%zeros
r@95%
r@99%
0
mol
0.920
1.311
15068
000 0
0 0.0%
1
5
poc
0.993
1.398
16267
000 0
0 0.0%
1
1
seq
1.023
1.374
16758
000 0
0 0.0%
1
1
10−6
mol
0.910
1.301
14914
000 0
0 0.0%
1
5
poc
0.991
1.395
16228
000 0
0 0.0%
1
1
seq
1.023
1.374
16755
000 0
0 0.0%
1
1
Appendix
Table A9: Mask structural analysis across λ for the concept-level model. Per-stream numbers are reported at K=128 under the last protocol, seed=1. L1 and rank statistics use the continuous pre-gate masks; zeros use the hard-gated forward masks. λ=10−4 remains element-wise dense but has a richer spectral tail, while λ≥10−3 produces hard zeros and rank-1 masks.
Configuration
Primary metrics
ROC-enrichment (RE)
AUROC
BEDROC 80.5
EF@1%
@0.5%
@1%
@2%
@5%
DrugCLIP [ Gao et al., 2023 ]
0.809
0.505
31.89
73.97
41.79
23.68
11.16
LigUnity [ Feng et al., 2025 ]
0.897
0.674
44.20
104.69
57.47
33.76
13.88
HypSeek [ Wang et al., 2026 ]
0.909
0.605
37.83
95.20
54.49
30.74
14.01
CausalBind SP ( K=128,dc=512,L=4 )
0.943
0.752
47.93
127.56
69.94
37.83
16.39
CausalBind SP ( K=128,dc=1024,L=4 )
0.935
0.744
47.69
128.06
69.26
37.07
15.99
Appendix
Table A10: ROC-enrichment (RE) on DUD-E under the last checkpoint protocol. Values are 102 -target means. AUROC / BEDROC 80.5 / EF@1% are reproduced from the main DUD-E comparison for context.
Method
Pearson r
Spearman ρ
LigUnity
0.532
0.499
HypSeek
0.423
0.390
CausalBind SP
0.598
0.548
Appendix
Table A11: Merck-FEP-8 affinity-ranking correlations. Higher Pearson r and Spearman ρ are better. CausalBind SP uses the main K=128,dc=1024 configuration under seed 1 and the last checkpoint.
Method
EF@0.5%
EF@1%
S 2 Drug (paper-reported)
11.44
7.38
CausalBind SP
10.88
8.56
CausalBind LR
11.59
8.84
CausalBind EMB
12.20
9.60
Appendix
Table A12: LIT-PCBA EF@0.5% comparison with the paper-reported S 2 Drug result.
Method
Target-level OOD (9 targets)
Scaffold-level OOD (6 targets)
AUROC
BEDROC 80.5
EF@1%
AUROC
BEDROC 80.5
EF@1%
HypSeek [ Wang et al., 2026 ]
0.819
0.593
20.06
0.880
0.630
44.47
LigUnity [ Feng et al., 2025 ]
0.859
0.570
19.82
0.915
0.612
43.11
CausalBind SP
0.860
0.658
23.00
0.945
0.712
43.66
CausalBind LR
0.867
0.653
22.71
0.949
0.716
44.70
CausalBind ATOM
0.837
0.666
22.94
0.941
0.744
50.77
Appendix
Table A13: Zero-shot out-of-distribution evaluation on DEKOIS 2.0. CausalBind ATOM is the atom-attentive version of CausalBind SP (Section 4 ).
Stage
Evidence
Observation
Purpose
Local structure
Heavy-atom contacts ( ≤4 Å)
98 contacts involving 27 ligand and 47 pocket atoms
Identify the local binding interface
Learned concept
Contact concentration in the most contact-aligned concept pair
0.0505 vs. 0.00881 uniform baseline (5.73 × enrichment)
Test whether contacts concentrate in concepts
Predicted score
Mask the associated 4 ligand and 8 pocket atoms (encoder fixed)
Score decreases 1.878→0.851 ( −54.7% )
Test contribution to the prediction
Random control
Mask 64 matched random sets of 4 ligand and 8 pocket atoms
Mean decrease 0.026±0.073 ; guided decrease 1.027 (99.2nd percentile)
Exclude an arbitrary masking effect
Appendix
Table A14: Local-to-global attribution on the DUD-E THRB complex.
Figure A3: Local-to-global attribution on the DUD-E THRB complex. (A) Ligand–pocket heavy-atom contacts within 4 Å in a 2D projection; highlighted atoms are those occluded at the concept bottleneck. (B) Attention-weighted contact probability over ligand–pocket concept pairs, with the contact-selected concepts and the top pair marked. (C) Decrease in the final binding score after masking the local contact atoms, matched random atoms, or the contact-selected concepts.
r
Params/branch
Protocol
DUD-E
LIT-PCBA
AUROC
BEDROC 80.5
EF@1%
AUROC
BEDROC 80.5
EF@1%
1
256
last
0.935
0.720
45.86
0.630
0.087
8.41
best_auc
0.946
0.756
48.78
0.636
0.079
6.53
best_bedroc
0.935
0.724
46.28
0.632
0.089
8.56
4
1,024
last
0.934
0.731
47.14
0.632
0.085
7.45
best_auc
0.933
0.710
45.14
0.633
0.080
7.67
Appendix
Table A15: CausalBind LR rank scan, seed=1, all three checkpoint protocols. All rows use K=128 , dc=1024 , L=4 , and λ=10−4 .
Method
n
DUD-E
LIT-PCBA
AUROC
BEDROC 80.5
EF@1%
AUROC
BEDROC 80.5
EF@1%
CausalBind LR ( r=1 )
3
0.937±0.004
0.732±0.011
46.87±0.91
0.630±0.006
0.085±0.006
7.22±0.31
CausalBind LR ( r=1 ) +Lspa
3
0.935±0.003
0.727±0.007
46.39±0.58
0.623±0.007
0.083±0.005
7.51±0.94
CausalBind LR ( r=8 ) +Lspa
3
0.940±0.001
0.750±0.011
48.30±0.97
0.621±0.005
0.087±0.003
7.73±0.40
CausalBind SP main
4
0.938±0.004
0.746±0.004
47.99±0.39
0.631±0.006
0.088±0.007
7.55±0.73
Appendix
Table A16: Multi-seed last -checkpoint comparison for CausalBind LR and the main CausalBind SP reference.
K
Params/branch
Protocol
DUD-E
LIT-PCBA
AUROC
BEDROC 80.5
EF@1%
AUROC
BEDROC 80.5
EF@1%
64
128
last
0.935
0.704
44.87
0.632
0.103
8.84
best_auc
0.934
0.702
44.63
0.638
0.099
9.03
best_bedroc
0.936
0.716
45.92
0.631
0.095
6.54
128
256
last
0.935
0.720
45.86
0.630
0.087
8.41
best_auc
0.946
0.756
48.78
0.636
0.079
6.53
Appendix
Table A17: CausalBind LR K -scan with r=1 , seed=1, all three protocols. All rows use dc=1024 , L=4 , and λ=10−4 .
r capacity
σ(M)
%zeros
r@95%
r@99%
Interpretation
1
0.86
0.0%
1
1
rank-one by construction
4
0.63
0.0%
1
1
trained solution collapses to rank one
8
0.54
8.9%
1
1
some thresholded zeros, still rank one
16
0.51
23.3%
1
1
more zeros, same effective rank
32
0.51
34.3%
1
1
unused rank capacity
64
0.50
39.7%
1
1
unused rank capacity
Appendix
Table A18: Mask diagnostics for the CausalBind LR rank scan. The reported branch uses the same protocol as Table A9 .
K
Protocol
DUD-E
LIT-PCBA
AUROC
BEDROC 80.5
EF@1%
AUROC
BEDROC 80.5
EF@1%
32
last
0.942
0.743
47.35
0.638
0.090
7.48
best_auc
0.942
0.747
47.84
0.640
0.085
6.86
best_bedroc
0.937
0.716
45.79
0.643
0.066
4.38
64
last
0.939
0.754
48.56
0.636
0.102
9.60
best_auc
0.941
0.765
49.47
0.635
0.097
8.18
Appendix
Table A19: CausalBind EMB K -scan (fix dc=256 ), seed=1, all three protocols.
dc
Protocol
DUD-E
LIT-PCBA
AUROC
BEDROC 80.5
EF@1%
AUROC
BEDROC 80.5
EF@1%
128
last
0.943
0.732
47.09
0.606
0.074
5.88
best_auc
0.944
0.726
46.59
0.617
0.075
6.97
best_bedroc
0.943
0.738
47.49
0.602
0.071
6.03
256
last
0.939
0.754
48.56
0.636
0.102
9.60
best_auc
0.941
0.765
49.47
0.635
0.097
8.18
Appendix
Table A20: CausalBind EMB dc -scan (fix K=64 ), seed=1, all three protocols. dc=256 is strongest on LIT-PCBA early enrichment.
Model
Protocol
DUD-E
LIT-PCBA
AUROC
BEDROC 80.5
EF@1%
AUROC
BEDROC 80.5
EF@1%
HypSeek [ Wang et al., 2026 ]
last
0.909
0.605
37.83
0.603
0.061
4.62
best_auc
0.913
0.629
39.63
0.608
0.063
5.13
best_bedroc
0.923
0.676
42.71
0.607
0.063
4.56
CausalBind SP
last
0.935
0.744
47.69
0.639
0.097
8.56
best_auc
0.934
0.722
46.41
0.628
0.074
5.84
Appendix
Table A21: Three-protocol comparison among HypSeek, CausalBind SP , and CausalBind EMB . CausalBind SP uses K=128 , dc=1024 , λ=10−4 , L=4 , which is the main operating point analyzed in the paper. CausalBind EMB uses K=64 , dc=256 , λ=10−2 , selected as the strongest setting in the embedding-level sweep. All rows are seed-1 results.
Branch
%zeros
r@95%
r@99%
mol–pocket
0.0%
1
46
mol–sequence
0.0%
1
46
Appendix
Table A22: Spectral diagnostics for the embedding-level mask ( K=64 , dc=256 , λ=10−2 , seed 1). We report the two branchwise product masks used in the paper’s interaction view, evaluated under the last checkpoint.
Protein-ligand modeling underpins computational drug discovery and molecular design. Existing protein-ligand benchmarks typically evaluate whether a protein and ligand interact and how strongly they bind, through tasks such as binary binding prediction and affinity regression. However, these evaluations provide limited evidence of whether models can localize binding sites or identify the non-covalent interactions underlying molecular recognition. To address this gap, we introduce InteractBind, a large-scale protein-ligand dataset comprising approximately 100k protein-ligand pairs, together with a benchmark for fine-grained evaluation. The core fine-grained task is that of binding-site localization, which uses protein-residue and ligand-atom interaction maps spanning six major types of non-covalent interactions to assess whether model-derived interaction maps localize binding sites. InteractBind further includes binding affinity and protein similarity-controlled splits to support realistic generalization assessment. Using InteractBind, we evaluate eight existing sequence-based and interaction-aware models, assessing binary binding prediction and binding-site localization. Results reveal limited binding-site localization despite strong binary binding prediction, with marked variation across non-covalent interaction types. Overall, InteractBind establishes a benchmark paradigm that encourages the development of more interpretable and physically grounded protein-ligand models.
Zhaohan Meng, Zhen Bai, Ke Yuan +4
School of Computing Science, University of Glasgow · School of Life Science and Technology, Institute of Science Tokyo · School of Cancer Sciences, University of Glasgow +4
Despite the high accuracy of 'black box' deep learning models, drug discovery still relies on protein-ligand interaction principles and heuristics. To improve interpretability of protein-small molecule binding predictions, we developed the PWRules framework, which applies binding affinity data to identify privileged small molecule fragments and subsequently defines complementary pairing rules between these fragments and protein words (semantic sequence units) through an interpretability module. The resulting word-fragment rules are then ranked by the PWScore function to prioritize active compounds. Evaluations on benchmark datasets show that PWScore achieves competitive performance comparable to the physics-based model (Glide) and the deep learning model (PSICHIC) and shows broad applicability for protein targets outside the training dataset, e.g., SARS-CoV-2 main protease. Notably, PWScore captures complementary interaction information, yielding superior enrichment performance when integrated with these established methods. Structural analysis of protein-ligand complexes indicates that learned word-fragment rules are significantly enriched near ligand-binding pockets, despite training without explicit structural guidance. By extracting and applying complementary pairing rules, PWRules provides an interpretable framework for drug discovery.
Jingke Chen, Jingrui Zhong, Tazneen Hossain Tani +3
MOE Key Laboratory of Bioinformatics, State Key Laboratory of Molecular Oncology, Beijing Frontier Research Center for Biological Structure, School of Pharmaceutical Sciences, Tsinghua University, Beijing, 100084, China
Protein-protein interactions (PPIs) govern nearly all cellular processes, yet computational methods for identifying binding partners typically produce ranked predictions without mechanistic justification. This creates a fundamental barrier to adoption because biologists cannot assess whether predictions reflect genuine biochemical insight or spurious correlations. We present \textbf{Protein Thoughts}, a framework that reformulates PPI discovery as an interpretable search problem with explicit reasoning. The system decomposes binding evidence into four biologically meaningful signals: sequence similarity reflecting evolutionary relationships, structural complementarity capturing geometric fit, interface balance, and chemical compatibility encoding residue-level interactions. Rather than collapsing these signals into an opaque score, we preserve their individual contributions through a transparent value function that enables both ranking and auditing. To navigate large candidate spaces efficiently, we introduce hypothesis-guided entropy-regularized Tree-of-Thoughts search. A fine-tuned language model generates search directives from embedding-derived features, classifying candidates as high-priority, exploratory, or skippable. These directives condition a Boltzmann policy that balances exploitation with entropy-driven exploration, while hypothesis-aware pruning prevents premature abandonment of promising candidates. For candidates exhibiting score disagreement, hypothesis-conditioned embedding-space flow matching transports protein embeddings toward the binder manifold. On the SHS148k benchmark, Protein Thoughts achieves mean best-binder rank of 11.2 versus 47.7 for an entropic tree search baseline, a 76% improvement, and for binding prediction the trained value function achieves 91.08±0.19 Micro-F1, outperforming existing PPI methods on the same dataset.
Kingsley Yeon, Xuefeng Liu, Promit Ghosal
Department of Statistics and CCAM University of Chicago Chicago, IL 60637 · School of Medicine Stanford University Stanford, CA 94305 · Department of Statistics University of Chicago Chicago, IL 60637