Proteins populate conformational ensembles, yet structure-based biomolecular design typically optimizes candidates against a single target conformation. Consequently, a candidate that fits one state can lose favorable interactions or develop steric clashes when the target adopts another. We introduce FlexEvo, a model-agnostic evolutionary framework that adapts candidates once at inference time from a single target conformation to improve compatibility with alternative natural conformations unseen during adaptation, without retraining the source model or requiring a conformational ensemble. FlexEvo casts cross-state adaptation as geometry-constrained bi-objective optimization, balancing preservation of input-state interactions against robustness to plausible conformational perturbations. To limit the search space and reduce invalid structural edits, geometry-derived FlexBoxes define protected anchor regions, adaptable regions for local exploration, and forbidden regions for clash avoidance. A unified all-atom representation supports topology-preserving adaptation across diverse binder categories, while Pareto selection preserves nondominated candidates across the two objectives. We evaluate FlexEvo across multiple generation baselines and nine representative binder categories spanning diverse molecular sizes and structural topologies. FlexEvo reduces the category-balanced mean relative performance degradation from 47.8% to 4.4%, while adding only 1.4--3.1 minutes of adaptation per sample. These results establish single-state inference-time adaptation as a practical route toward robust biomolecular complex design across protein conformational landscapes.
Figures & tables
Figure 1: Motivation for cross-conformation ligand design. (a) Comparison of FlexEvo and single-state design methods, illustrating ligand compatibility across alternative target conformations. (b) Alignment of known protein–ligand complexes highlights a shared stable binding region, motivating ligand design that preserves conserved interactions while accommodating conformational variation.
Figure 2: Overview of the FlexEvo framework. Starting from a single target conformation, FlexEvo filters an initial candidate pool and constructs geometry-derived FlexBoxes defining protected, adaptable, and forbidden regions. Iterative structural editing, repair, and Pareto-based selection balance input-state interaction preservation with robustness to conformational perturbations.
Figure 3: Cross-conformation evaluation and analysis of FlexEvo. (a–c) Distributions of scores derived from Vina, contact F1, and DockQ before and after adaptation for small molecules, peptides, and protein binders, respectively, under the paired target conditions shown. (e–g) Initial, intermediate, and final populations in the corresponding two-objective spaces. (d,h) Mean residue RMSF in the lowest and highest within-protein quartiles of packing sparsity ( 1/WCN ) and distance from the whole-chain Cα geometric center, respectively, for 200 ATLAS proteins. Quartile means are normalized by each protein’s whole-chain mean RMSF. In (d,h), points denote proteins, faint lines pair quartiles, and diamonds with thick lines indicate equal-weight means across proteins.
Figure 4: Metric trajectories and structural progression. (a) High-affinity fraction (blue, left axis) and logP (green, right axis) over adaptation time. (b,c) On-target pAE (blue, left axis) and off-target pAE (orange, right axis). Time is shown in minutes. (d–f) Early, intermediate, and late binding poses for one ligand example, with the corresponding Vina scores shown below.
Base model
Vina ↓
High affinity (%) ↑
Vinardo ↑
AutoDock4 ↑
RTMScore ↑
QED ↑
Reference
Matched
Divergent
Reference
Matched
Divergent
Reference
Matched
Divergent
Reference
Matched
Divergent
Reference
Matched
Divergent
Reference
Matched
Divergent
Small molecule
DynamicFlow
-9.18 -2.57
-8.36 -1.89
-7.80 -2.95
58.2 +6.9
56.7 +16.8
61.7 +35.4
.698 +.282
.537 +.131
.587 +.232
.597 +.019
.570 +.141
.672 +.298
.619 +.248
.594 +.135
.571 +.129
.594 +.060
–
–
FlexSBDD
-9.22 -2.32
-8.27 -1.58
-7.74 -2.89
55.4 +8.3
57.1 +17.8
56.5 +17.6
.552 +.142
.607 +.218
.603 +.242
.585 +.244
.643 +.201
.588 +.182
.532 +.204
.572 +.110
.576 +.211
.589 +.215
–
–
YuelDesign
-9.18 -3.46
-8.44 -2.86
-7.67 -3.72
62.5 +13.0
61.3 +24.5
59.5 +28.3
.584 +.192
.643 +.269
.578 +.185
.652 +.220
.616 +.177
.619 +.268
.656 +.188
.584 +.171
.587 +.198
.612 +.191
–
–
Nonpeptidic macrocycle
Table 1: Cross-conformation performance across nine binder categories. Reference is the adaptation conformation; Matched and Divergent are alternative natural conformations held out from adaptation. Black entries report post- FlexEvo scores. Superscripts give changes from the unadapted baseline ( Δ=adapted−unadapted ); green and red indicate improvement and deterioration according to each metric’s preferred direction. Changes in percentage-valued metrics are expressed in percentage points. Vina, DockQ, and QED use means; LRMSD uses medians. Values and changes are rounded independently from the source values. QED is reported once per method in the Reference columns and applies to all three conformations; dashes denote repeated values omitted from display.
Appendix figures & tables41 assets
Supplementary material from the paper’s appendix.
Appendix
Base model
Vina Mean ↓
Vina Median ↓
High affinity (%) ↑
Vinardo ↑
AutoDock4 ↑
RTMScore ↑
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Reference conformation
DynamicFlow
-6.6062
-9.1775
-7.0317
-8.7901
51.2524
58.1689
0.4161
0.698
0.5782
0.5971
0.3716
0.6193
FlexSBDD
-6.906
-9.2245
-6.7984
-8.8448
47.0322
55.3719
0.4102
0.5524
0.3411
0.5847
0.3287
0.5322
YuelDesign
-5.7174
-9.1754
-5.4763
-8.8673
49.5198
62.5176
0.392
0.584
0.4317
0.6519
0.4674
0.6556
Matched conformation
Appendix
Table 2: Small-molecule drugs. Black: base model. For directional metrics with FlexEvo, green: improved or tied; red: worse. Adapted logP is descriptive and shown in black; arrows indicate preferred directions where defined. QED, SA, Lipinski, logP, and PAINS summaries are presented under Reference. Binding scores are reported separately for Reference, Matched, and Divergent. Diversity summarizes the frozen candidate pool and is reported once under Reference; dashes indicate that it is not repeated for Matched or Divergent.
Base model
Vina Mean ↓
Vina Median ↓
High affinity (%) ↑
Vinardo ↑
AutoDock4 ↑
RTMScore ↑
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Reference conformation
DynamicFlow
-5.9197
-8.3205
-6.3195
-7.8654
46.5292
51.6061
0.3698
0.6333
0.5181
0.5385
0.3295
0.5594
FlexSBDD
-6.2431
-8.2445
-6.0054
-8.0434
41.9284
50.0367
0.3668
0.4865
0.3012
0.5305
0.2965
0.4753
YuelDesign
-5.1565
-8.2256
-4.8197
-7.9627
44.0439
56.2662
0.3476
0.5204
0.3911
0.5904
0.4201
0.5806
Matched conformation
Appendix
Table 3: Nonpeptidic macrocyclic drugs. Black: base model. For directional metrics with FlexEvo, green: improved or tied; red: worse. Adapted logP is descriptive and shown in black; arrows indicate preferred directions where defined. QED, SA, Lipinski, logP, and PAINS summaries are presented under Reference. Binding scores are reported separately for Reference, Matched, and Divergent. Diversity summarizes the frozen candidate pool and is reported once under Reference; dashes indicate that it is not repeated for Matched or Divergent.
Base model
DockQ Mean ↑
DockQ Median ↑
LRMSD <2 Å (%) ↑
LRMSD Median (Å) ↓
MM/GBSA (kcal/mol) ↓
HADDOCK Score ↓
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Reference conformation
BoltzGen
0.6566
0.8739
0.4699
0.7672
42.1851
40.1549
5.4802
5.3169
-40.3664
-44.5812
-127.0134
-126.5549
Chai-1
0.6434
0.7716
0.4811
0.6535
39.6334
39.6517
5.8455
4.9851
-40.563
-43.8717
-122.8868
-124.9151
Protenix
0.7154
0.8198
0.6704
0.7799
40.0431
39.7062
4.2893
4.5944
-45.0344
-46.2925
-123.1837
-123.3106
Matched conformation
Appendix
Table 4: Linear peptide drugs. Black: base model. With FlexEvo, green: improved or tied; red: worse. Arrows indicate the preferred direction. Diversity summarizes the frozen candidate pool and is reported once under Reference; dashes indicate that it is not repeated for Matched or Divergent.
Base model
DockQ Mean ↑
DockQ Median ↑
LRMSD <2 Å (%) ↑
LRMSD Median (Å) ↓
MM/GBSA (kcal/mol) ↓
HADDOCK Score ↓
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Reference conformation
BoltzGen
0.6831
0.9018
0.3784
0.8291
43.2014
42.1781
5.6779
5.6422
-42.8373
-46.3974
-132.5755
-132.5663
Chai-1
0.856
0.8613
0.7376
0.7577
41.2663
40.9744
5.6419
4.7438
-46.2865
-44.6798
-129.835
-126.3605
Protenix
0.7567
0.8618
0.7788
0.7398
41.0919
42.0591
4.4517
4.7119
-47.4363
-48.4915
-129.0268
-131.1541
Matched conformation
Appendix
Table 5: Cyclic peptide drugs. Black: base model. With FlexEvo, green: improved or tied; red: worse. Arrows indicate the preferred direction. Diversity summarizes the frozen candidate pool and is reported once under Reference; dashes indicate that it is not repeated for Matched or Divergent.
Base model
DockQ Mean ↑
DockQ Median ↑
LRMSD <2 Å (%) ↑
LRMSD Median (Å) ↓
MM/GBSA (kcal/mol) ↓
HADDOCK Score ↓
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Reference conformation
BoltzGen
0.6432
0.8617
0.4529
0.7936
38.5198
42.7348
5.4221
4.2243
-39.9824
-40.6781
-117.7648
-119.6118
Chai-1
0.731
0.8331
0.5799
0.7596
41.3627
44.8159
4.6758
3.8565
-38.3406
-39.3865
-119.3965
-111.0687
Protenix
0.6655
0.8178
0.5644
0.7984
44.2448
45.3202
3.814
3.7879
-40.5226
-41.3208
-119.1796
-113.8899
Matched conformation
Appendix
Table 6: Non-antibody protein drugs. Black: base model. With FlexEvo, green: improved or tied; red: worse. Arrows indicate the preferred direction. Diversity summarizes the frozen candidate pool and is reported once under Reference; dashes indicate that it is not repeated for Matched or Divergent.
Base model
DockQ Mean ↑
DockQ Median ↑
LRMSD <2 Å (%) ↑
LRMSD Median (Å) ↓
MM/GBSA (kcal/mol) ↓
HADDOCK Score ↓
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Reference conformation
BoltzGen
0.6095
0.7995
0.3297
0.6427
36.2452
39.8647
5.1441
3.9774
-37.4543
-37.7412
-109.6532
-113.6292
Chai-1
0.6808
0.7909
0.4454
0.6112
39.1404
42.0743
4.3318
3.6037
-35.4734
-37.0303
-111.8543
-103.6798
Protenix
0.6178
0.7616
0.3397
0.5605
40.9499
42.9017
3.5419
3.5284
-37.6326
-39.1405
-110.3069
-106.1588
Matched conformation
Appendix
Table 7: Full-length antibody drugs. Black: base model. With FlexEvo, green: improved or tied; red: worse. Arrows indicate the preferred direction. Diversity summarizes the frozen candidate pool and is reported once under Reference; dashes indicate that it is not repeated for Matched or Divergent.
Base model
DockQ Mean ↑
DockQ Median ↑
LRMSD <2 Å (%) ↑
LRMSD Median (Å) ↓
MM/GBSA (kcal/mol) ↓
HADDOCK Score ↓
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Reference conformation
BoltzGen
0.6169
0.8239
0.3372
0.6666
37.1832
41.5067
5.2336
4.0466
-38.1485
-39.1582
-112.0832
-114.8399
Chai-1
0.6969
0.8022
0.4589
0.6322
39.8487
43.1363
4.4744
3.7165
-36.8421
-37.5865
-113.4421
-105.6989
Protenix
0.6342
0.7832
0.3542
0.5768
42.3102
43.6115
3.6699
3.5971
-38.8703
-40.1934
-115.6324
-108.4364
Matched conformation
Appendix
Table 8: Antibody fragments and nanobodies. Black: base model. With FlexEvo, green: improved or tied; red: worse. Arrows indicate the preferred direction. Diversity summarizes the frozen candidate pool and is reported once under Reference; dashes indicate that it is not repeated for Matched or Divergent.
Base model
DockQ Mean ↑
DockQ Median ↑
LRMSD <2 Å (%) ↑
LRMSD Median (Å) ↓
MM/GBSA (kcal/mol) ↓
HADDOCK Score ↓
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Reference conformation
BoltzGen
0.5812
0.7918
0.3225
0.6375
35.4166
38.8602
4.8907
3.8668
-36.9099
-37.2423
-108.2698
-108.9199
Chai-1
0.6744
0.7572
0.4361
0.5988
37.3752
40.4985
4.3123
3.5139
-35.0184
-36.1978
-107.8819
-100.5058
Protenix
0.6139
0.7419
0.3305
0.5449
39.9387
40.8061
3.4706
3.4549
-37.2316
-38.0442
-108.3421
-105.2553
Matched conformation
Appendix
Table 9: Oligonucleotide drugs. Black: base model. With FlexEvo, green: improved or tied; red: worse. Arrows indicate the preferred direction. Diversity summarizes the frozen candidate pool and is reported once under Reference; dashes indicate that it is not repeated for Matched or Divergent.
Base model
DockQ Mean ↑
DockQ Median ↑
LRMSD <2 Å (%) ↑
LRMSD Median (Å) ↓
MM/GBSA (kcal/mol) ↓
HADDOCK Score ↓
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Base
+FlexEvo
Reference conformation
BoltzGen
0.6435
0.8616
0.3526
0.7935
41.2659
45.2835
5.7891
4.4419
-42.5524
-43.7174
-126.0592
-126.9724
Chai-1
0.7316
0.8332
0.4798
0.7594
44.0388
47.8695
4.9724
4.0948
-40.3513
-41.9244
-125.9534
-117.1707
Protenix
0.6651
0.8176
0.3642
0.6987
47.3985
47.6928
4.0711
4.0487
-43.3554
-43.9896
-127.3749
-119.7496
Matched conformation
Appendix
Table 10: mRNA and long nucleic acid drugs. Black: base model. With FlexEvo, green: improved or tied; red: worse. Arrows indicate the preferred direction. Diversity summarizes the frozen candidate pool and is reported once under Reference; dashes indicate that it is not repeated for Matched or Divergent.
Symbol
Meaning
P0 , R
Observed target coordinates and number of represented target residues.
G , RG
Spatial-prior genome and its residue support.
cr , Qr , zr
Fixed residue centroid, fixed local frame, and evolvable box scale.
x , n(x) , Hx
Candidate record, represented coordinate count, and edit history.
X0 , XG
Shared screened source pool and pool adapted under G .
PGm , K
Target probe m and total probe count; K=7 includes Reference.
Appendix
Table 12: Principal objects used in the implementation and analysis.
Component
Parameter
Default
Randomization
Seed
31
Screening
Pool cap / neighbor threshold
64/0.35
Screening
Molecular / geometric similarity
0.92/0.90
FlexBox
Scale bounds
[0.55,1.75]
FlexBox
Stable / flexible / forbidden extent
3.0/4.5/2.9 Å
Inner adaptation
Rounds / maximum population
2/12
Appendix
Table 13: Default configuration of the coordinate-level FlexEvo core.
Figure 5: Structural and functional characteristics of the FlexEvo structure pool. RCSB Protein Data Bank (PDB) annotations characterize the small-molecule, peptide and protein-binder subsets. a, Numbers of unique current PDB entries in each subset: 19,715 for small molecules, 9,858 for peptides and 8,726 for protein binders. b, Proportions of structures determined by X-ray diffraction, solution nuclear magnetic resonance (NMR) and other experimental methods. c, Distributions of reported resolution for X-ray structures. d, Density distributions of initial PDB release years. e, Prevalence of the displayed PDB functional classes, expressed as percentages of entries within each modality. f, Polymer composition, categorized as homomeric protein, heteromeric protein, protein–oligosaccharide, protein–nucleic acid and other or mixed assemblies. g, Cumulative distributions of deposited non-polymer instance counts per entry, displayed up to 30 instances. h, Alluvial representation linking modality, PDB functional class and polymer composition. Blue, teal and orange distinguish the small-molecule, peptide and protein-binder subsets, respectively, in panels a, c–e and g.
Figure 6: Taxonomic, functional and translational annotation landscape of small-molecule targets in FlexEvo. a, Taxonomic composition of 1,774 targets mapped to reviewed UniProt entries, including 713 human targets (40.2%) and 615 bacterial targets (34.7%). b, Functional target-class composition. In panels a and b, categories containing at least 60 targets are labelled with their counts and percentages. c, Coverage of external database cross-references and UniProt disease annotations among mapped targets, including AlphaFoldDB, ChEMBL, DrugBank, Reactome, Pharos, Open Targets, DrugCentral and Guide to PHARMACOLOGY. d, Corresponding annotation coverage within the human-target subset. e, Annotation coverage stratified by functional target class; darker blue indicates a higher percentage of targets with the corresponding cross-reference or annotation. f, Within-class distributions of receptor-conformation counts per target, grouped into 2–3, 4–7, 8–15, 16–31 and ≥32 conformations, with an additional unmapped category. Database coverage quantifies the availability of target annotations and external cross-references.
Method
Reference ↓
Matched ↓
Divergent ↓
Base
−8.260
−7.440
−6.540
Rerank-only
−8.560
−7.370
−7.750
FlexEvo
−9.740
−8.678
−8.698
Appendix
Table 14: Three-state rerank-only comparison for conventional small-molecule drugs. Vina scores are reported on the common three-arm evaluation set and aggregated across DynamicFlow, FlexSBDD, and YuelDesign. Scores are in kcal/mol; lower values are better. Rerank-only selects an original source candidate without coordinate editing. Bold values indicate the best model-averaged score in each state.
Method
Reference ↑
Matched ↑
Divergent ↑
Base
0.6270
0.3960
0.2760
Rerank-only
0.6630
0.4213
0.3372
FlexEvo
0.8539
0.8614
0.8173
Appendix
Table 15: Three-state rerank-only comparison using DockQ. DockQ scores are reported on the common three-arm evaluation set and aggregated across BoltzGen, Chai-1, and Protenix. Higher values are better. Rerank-only selects an original source candidate without coordinate editing. Bold values indicate the best model-averaged score in each state.
Figure 7: Cross-conformation evaluation of conventional small molecules, nonpeptidic macrocycles and linear peptides. a , Conventional small-molecule drugs. b , Nonpeptidic macrocyclic drugs. c , Linear peptide drugs. In a and b , columns show DynamicFlow, FlexSBDD and YuelDesign from left to right. The horizontal and vertical axes show signed Vina margins under matched and divergent receptor conformations, respectively (kcal mol -1 ; higher is better). In c , columns show BoltzGen, Chai-1 and Protenix from left to right. The horizontal and vertical axes show DockQ margins under reference-complex and divergent-complex conditions, respectively, defined as DockQ−0.49 . Grey open symbols and coloured points denote the original configuration (two candidates) and the FlexEvo-augmented configuration (three candidates), respectively. Diamond markers indicate configuration centroids, with connecting lines showing the displacement between configurations. Dashed lines mark zero margins, and the shaded upper-right quadrant indicates non-negative margins under both conditions.
Figure 8: Cross-conformation evaluation of cyclic peptides, non-antibody proteins and full-length antibodies. a , Cyclic peptide drugs. b , Non-antibody protein drugs. c , Full-length antibody drugs. Columns show BoltzGen, Chai-1 and Protenix from left to right in each panel. The horizontal and vertical axes show DockQ margins under reference-complex and divergent-complex conditions, respectively. Margins are defined as DockQ−0.49 , with higher values indicating better structural agreement with the reference complex under the corresponding condition. Grey open symbols and coloured points denote the original model configuration (two candidates) and the model supplemented with FlexEvo (three candidates), respectively. Diamond markers indicate configuration centroids, with connecting lines showing the displacement between configurations. Dashed lines mark zero margins, corresponding to DockQ=0.49 . The shaded upper-right quadrant indicates DockQ≥0.49 under both conditions.
Figure 9: Cross-conformation evaluation of antibody fragments, nanobodies and nucleic acid drugs. a , Antibody fragments and nanobodies. b , Oligonucleotide drugs. c , Long nucleic acid drugs, including mRNA. Columns show BoltzGen, Chai-1 and Protenix from left to right in each panel. The horizontal and vertical axes show DockQ margins under reference-complex and divergent-complex conditions, respectively. Margins are defined as DockQ−0.49 , with higher values indicating better structural agreement with the reference complex under the corresponding condition. Grey open symbols and coloured points denote the original model configuration (two candidates) and the model supplemented with FlexEvo (three candidates), respectively. Diamond markers indicate configuration centroids, with connecting lines showing the displacement between configurations. Dashed lines mark zero margins, corresponding to DockQ=0.49 . The shaded upper-right quadrant indicates DockQ≥0.49 under both conditions.
Method
HV-AUC ↑
B-pocket top-10 ↑
Best-of- N
0.959
−32.497
FlexEvo
8.482
−31.409
Appendix
Table 16: Best-of- N comparison using the values displayed in Figure 10 . Higher values are preferred for both metrics.
Figure 10: Comparison of FlexEvo and alternative optimization strategies using raw performance metrics. Six methods are compared: Best-of- N , hill climbing, annealed Markov chain Monte Carlo (MCMC), FlexEvo, separable covariance matrix adaptation evolution strategy (sep-CMA-ES) and trust-region Bayesian optimization (BO). a, Hypervolume area under the curve (HV-AUC), displayed on a logarithmic scale; higher values indicate better performance. b, B-pocket top-10 task scores, for which higher (less negative) values indicate better performance. c, Joint comparison of the two metrics, with the upper-right region indicating better performance on both. The open grey diamond denotes the Best-of- N arithmetic-mean summary. d, Rankings of the six displayed values for HV-AUC (blue circles) and B-pocket score (orange diamonds), with rank 1 indicating the highest value. Connecting lines link the two metric-specific ranks for each method. FlexEvo achieves the highest displayed value on both metrics, with an HV-AUC of 8.482 and a B-pocket top-10 score of −31.409 .
Figure 11: End-to-end computational cost of incorporating FlexEvo into generation workflows. Per-sample compute time is shown for DynamicFlow, FlexSBDD, YuelDesign, and the peptide and binder workflows of BoltzGen, Chai-1 and Protenix. Stacked bars separate baseline generation time (grey) from the additional time required for FlexEvo optimization (teal). Labels within the teal segments indicate the added optimization time, and annotations beside each bar report baseline, optimization and total times. Across the nine workflows, FlexEvo adds 1.393–3.141 min per sample, resulting in total compute times of 6.861–12.966 min per sample.
Figure 12: Small-molecule performance across receptor conformations with and without FlexEvo. a, Comparison of DynamicFlow (orange circles), FlexSBDD (purple squares) and YuelDesign (teal diamonds) under reference, matched and divergent receptor conformations. Dashed lines with open markers indicate the original methods; solid lines with filled markers indicate their FlexEvo-optimized counterparts. b, Expanded view of the FlexEvo trajectories on a linear scale, with labels indicating performance values (%). FlexEvo-optimized DynamicFlow, FlexSBDD and YuelDesign retain performance of 81%, 83% and 82%, respectively, under the divergent condition, with decreases of 4–5 percentage points relative to the reference condition.
Figure 13: Peptide performance across receptor conformations with and without FlexEvo. a, Comparison of BoltzGen (blue circles) and Protenix (red squares) under reference, matched and divergent receptor conformations. Dashed lines with open markers indicate the original methods; solid lines with filled markers indicate their FlexEvo-optimized counterparts. b, Expanded view of the FlexEvo trajectories on a linear scale, with labels indicating performance values (%). BoltzGen achieves 86%, 84% and 82%, and Protenix achieves 84%, 82% and 81%, respectively, across the three conditions. FlexEvo reduces the decline in performance associated with changes in receptor conformation.
Figure 14: Protein-binder performance across receptor conformations with and without FlexEvo. a, Comparison of BoltzGen (blue circles), Chai-1 (purple squares) and Protenix (red diamonds) under reference, matched and divergent receptor conformations. Dashed lines with open markers indicate the original methods; solid lines with filled markers indicate their FlexEvo-optimized counterparts. b, Expanded view of the FlexEvo trajectories on a linear scale, with labels indicating performance values (%). Performance decreases from 87% to 83% for BoltzGen, from 85% to 81% for Chai-1 and from 89% to 84% for Protenix between the reference and divergent conditions, showing smaller declines than those observed for the original methods.
Evaluation scenario
Maximum Tc
Median Tc
Antibody fragments and nanobodies
0.222
0.052
Conventional small-molecule drugs
0.159
0.052
Cyclic peptide drugs
0.262
0.055
Full-length antibody drugs
0.245
0.059
Linear peptide drugs
0.226
0.064
Long nucleic acid drugs, including mRNA
0.254
0.071
Appendix
Table 17: Reference-library similarity across nine binder categories. For each category, 10 evaluated candidates are compared with 47 randomly sampled public reference entries. Maximum and median similarity are reported over the resulting 470 pairwise comparisons using the modality-specific similarity definition.
Figure 15: Tanimoto similarity between generated samples and public reference molecules for three drug modalities. a , Conventional small-molecule drugs. b , Nonpeptidic macrocyclic drugs. c , Linear peptide drugs. Each heatmap compares 10 generated samples (rows) with 47 public reference molecules (columns). Each cell represents the Tanimoto similarity of the corresponding pair, with colours indicating values from 0 to 1. Higher values indicate greater similarity, and all panels share the same colour scale.
Figure 16: Tanimoto similarity between generated samples and public reference molecules for cyclic peptides and protein-based drug modalities. a , Cyclic peptide drugs. b , Non-antibody protein drugs. c , Antibody fragments and nanobodies. Each heatmap compares 10 generated samples (rows) with 47 public reference molecules (columns). Each cell represents the Tanimoto similarity of the corresponding pair, with colours indicating values from 0 to 1. Higher values indicate greater similarity, and all panels share the same colour scale.
Figure 17: Tanimoto similarity between generated samples and public reference molecules for antibody and nucleic acid drug modalities. a , Full-length antibody drugs. b , Oligonucleotide drugs. c , Long nucleic acid drugs, such as mRNA. Each heatmap compares 10 generated samples (rows) with 47 public reference molecules (columns). Each cell represents the Tanimoto similarity of the corresponding pair, with colours indicating values from 0 to 1. Higher values indicate greater similarity, and all panels share the same colour scale.
Figure 18: Representative structures of conventional small molecules, nonpeptidic macrocycles and linear peptides. a–c , Conventional small-molecule drugs: imatinib bound to ABL (PDB: 1IEP), gefitinib bound to EGFR (2ITY), and erlotinib bound to EGFR (1M17). d–f , Nonpeptidic macrocyclic drugs: rapamycin in the FKBP12–FRB complex (1FAP), lorlatinib bound to ALK (4CLI), and rifampicin bound to RNA polymerase (1I6V). Panel f shows the local protein environment surrounding rifampicin. g–i , Linear peptide examples: an exendin-4 fragment bound to the GLP-1 receptor extracellular domain (3C5T), GLP-1 bound to the same domain (3IOL), and a long-acting parathyroid hormone analogue bound to PTH1R (6NBF). Target proteins are shown as pale-green surfaces, with bound molecules represented as sticks or cartoons. Grey boxes identify regions reproduced in the enlarged insets. Structures are rendered from deposited coordinates and independently scaled.
Figure 19: Representative structures of cyclic peptides, non-antibody proteins and full-length antibodies. a–c , Cyclic peptide examples: cyclosporin A bound to cyclophilin A (PDB: 1CWA), cyclosporin A bound to an antibody Fab fragment (1IKF), and octreotide bound to somatostatin receptor 2 (7T11). d–f , Non-antibody protein examples: human growth hormone bound to its receptor (3HHR), porcine insulin (4INS), and human erythropoietin bound to its receptor (1EER). g–i , Full-length antibody examples: pembrolizumab IgG4 (5DK3), IgG2a MAb231 (1IGT), and IgG1 MAb61.1.3 (1IGY). The latter two antibodies illustrate intact immunoglobulin architectures. Where present, target proteins are shown as pale-green surfaces. Protein ligands are shown as cartoons, whereas intact antibodies are shown as surfaces with heavy and light chains distinguished by colour. Grey boxes identify regions reproduced in the enlarged insets. Structures are rendered from deposited coordinates and independently scaled.
Figure 20: Representative antibody-fragment and oligonucleotide structures, and long-RNA architectures. a–c , Antibody fragments and nanobodies: a camel single-domain antibody bound to lysozyme (PDB: 1MEL), trastuzumab Fab bound to HER2 (1N8Z), and B38 Fab bound to the SARS-CoV-2 spike receptor-binding domain (7BZ5). d–f , Oligonucleotide examples: a 15-nucleotide DNA aptamer (1HAO), a 26-nucleotide RNA aptamer (3DD2), and a 27-nucleotide DNA aptamer (4I7Y), each bound to thrombin. These aptamers serve as structural representatives of the oligonucleotide modality. Target proteins are shown as pale-green surfaces, with antibody fragments and nucleic acids represented as cartoons or sticks. g–i , Schematic architectures of linear mRNA, circular RNA and self-amplifying RNA, respectively. RNA schematics illustrate molecular organization without specifying sequence, chain length or experimentally determined folding. Grey boxes identify regions reproduced in the enlarged insets. Molecular views are independently scaled.
Figure 21: Physicochemical profiles of selected reference molecules. The figure summarizes seven molecular descriptors for 116 reference structures selected for illustration and divided into two example sets ( n=58 each), shown as blue circles and terracotta triangles. a–g, Exact molecular mass (Da), calculated octanol–water partition coefficient (cLogP), topological polar surface area (TPSA; Å 2 ), hydrogen-bond donor count, hydrogen-bond acceptor count, rotatable bond count and the fraction of carbon atoms with sp3 hybridization, respectively. Grey violins show pooled distributions of continuous descriptors; boxes show the first-to-third-quartile intervals for count descriptors, with whiskers extending to the most extreme observations within 1.5 interquartile ranges (IQRs) of the box boundaries. Open black circles and thick black intervals indicate pooled medians and IQRs. Numerical annotations report median [Q1, Q3], and each coloured marker represents one molecule. Red dashed lines mark commonly used physicochemical guideline upper bounds: molecular mass of 500 Da, cLogP of 5, TPSA of 140 Å 2 , 5 hydrogen-bond donors, 10 hydrogen-bond acceptors and 10 rotatable bonds. For all six descriptors with reference lines, the pooled median and upper quartile lie below the corresponding guideline, although individual molecules exceed these values. The median fraction of sp3 -hybridized carbon atoms is 0.45, with an IQR of 0.31–0.60. h, Key to the graphical encodings.
Figure 22: Performance comparisons for nonpeptidic macrocycles, oligonucleotides and full-length antibodies. Rows show nonpeptidic macrocyclic drugs, oligonucleotide drugs and full-length antibody drugs, respectively. Columns correspond to reference, matched and divergent conformations, from left to right. DynamicFlow, FlexSBDD and YuelDesign are compared with their FlexEvo variants for macrocycles; BoltzGen, Chai1 and Protenix are compared with their FlexEvo variants for oligonucleotides and antibodies. Colors identify model families within each row. Dashed lines with open squares indicate base models, whereas solid lines with filled circles indicate the corresponding FlexEvo variants. Each metric is min–max normalized to [0,1] across all 18 method–conformation combinations within the same drug modality. Metrics marked with downward arrows are reverse-scaled, so outward values follow the indicated preferred direction. Normalization bounds are shared across the three columns within each row but differ between modalities. Means are used where indicated; LRMSD success rates below 2A˚ and median LRMSD are shown separately. The logP axis displays increasing hydrophobicity according to the stated plotting convention. Performance comparisons follow the individual metric axes.
Figure 23: Performance comparisons for antibody fragments, nanobodies and peptide drugs. Rows show antibody fragments and nanobodies, linear peptide drugs and cyclic peptide drugs, respectively. Columns correspond to reference, matched and divergent conformations, from left to right. Each panel compares BoltzGen, Chai1 and Protenix with their corresponding FlexEvo variants. Colors identify model families within each row. Dashed lines with open squares indicate base models, whereas solid lines with filled circles indicate the corresponding FlexEvo variants. Each metric is min–max normalized to [0,1] across all 18 method–conformation combinations within the same drug modality. Metrics marked with downward arrows are reverse-scaled, so outward values follow the indicated preferred direction. Normalization bounds are shared across the three columns within each row but differ between modalities. Means are used where indicated; LRMSD success rates below 2A˚ and median LRMSD are shown separately. Diversity is represented by the median for cyclic peptides and by the mean for the other two modalities. Performance comparisons follow the individual metric axes.
Figure 24: Performance comparisons for long nucleic acids, small molecules and non-antibody proteins. Rows show long nucleic acid drugs such as mRNA, conventional small-molecule drugs and non-antibody protein drugs, respectively. Columns correspond to reference, matched and divergent conformations, from left to right. BoltzGen, Chai1 and Protenix are compared with their FlexEvo variants for nucleic acids and proteins; DynamicFlow, FlexSBDD and YuelDesign are compared with their FlexEvo variants for small molecules. Colors identify model families within each row. Dashed lines with open squares indicate base models, whereas solid lines with filled circles indicate the corresponding FlexEvo variants. Each metric is min–max normalized to [0,1] across all 18 method–conformation combinations within the same drug modality. Metrics marked with downward arrows are reverse-scaled, so outward values follow the indicated preferred direction. Normalization bounds are shared across the three columns within each row but differ between modalities. Means are used where indicated; LRMSD success rates below 2A˚ and median LRMSD are shown separately. The logP axis displays increasing hydrophobicity according to the stated plotting convention. Performance comparisons follow the individual metric axes.
Figure 25: On-target and off-target interaction scores for small-molecule, macrocyclic and linear peptide drug candidates. Rows show a , conventional small-molecule drugs; b , nonpeptidic macrocyclic drugs; and c , linear peptide drugs. Columns retain the original A–C designations. Each point represents a candidate, plotted by its on-target pAE_interaction score and the minimum score across the evaluated off-targets. Reference lines mark an on-target score of 5 and a minimum off-target score of 10. Red boxes highlight the region with on-target scores below 5 and minimum off-target scores above 10. Blue points fall within this region, whereas grey points fall outside it. Annotated percentages indicate the proportion of candidates within the highlighted region.
Figure 26: On-target and off-target interaction scores for cyclic peptide, protein and full-length antibody drug candidates. Rows show a , cyclic peptide drugs; b , non-antibody protein drugs; and c , full-length antibody drugs. Columns retain the original A–C designations. Each point represents a candidate, plotted by its on-target pAE_interaction score and the minimum score across the evaluated off-targets. Reference lines mark an on-target score of 5 and a minimum off-target score of 10. Red boxes highlight the region with on-target scores below 5 and minimum off-target scores above 10. Blue points fall within this region, whereas grey points fall outside it. Annotated percentages indicate the proportion of candidates within the highlighted region.
Figure 27: On-target and off-target interaction scores for antibody fragment and nucleic acid drug candidates. Rows show a , antibody fragments and nanobodies; b , oligonucleotide drugs; and c , long nucleic acid drugs, such as mRNA. Columns retain the original A–C designations. Each point represents a candidate, plotted by its on-target pAE_interaction score and the minimum score across the evaluated off-targets. Reference lines mark an on-target score of 5 and a minimum off-target score of 10. Red boxes highlight the region with on-target scores below 5 and minimum off-target scores above 10. Blue points fall within this region, whereas grey points fall outside it. Annotated percentages indicate the proportion of candidates within the highlighted region.
Figure 28: Optimization trajectories across conventional small-molecule drugs, nonpeptidic macrocyclic drugs, and linear peptide drugs. Panels a–f, g–l, and m–r correspond to the three categories, respectively. For the two small-molecule categories, panels a and g show Vina scores; b and h, QED and SA scores; c and i, Lipinski scores and logP; d and j, high-affinity and PAINS-pass fractions; e and k, auxiliary scoring functions; and f and l, diversity. For linear peptides, panels m–r show DockQ, MM/GBSA and HADDOCK scores, LRMSD success rate, LRMSD, on-target and off-target pAE, and IDDT and diversity, respectively. The horizontal axes indicate cumulative FlexEvo optimization time per sample. The recorded trajectories show metric evolution during adaptation, including local fluctuations and late-stage stabilization. Shaded bands summarize variability across the recorded experimental runs.
Figure 29: Optimization trajectories across cyclic peptide drugs, non-antibody protein drugs, and full-length antibody drugs. Panels a–f, g–l, and m–r correspond to the three categories, respectively. Within each category, the six panels report DockQ, MM/GBSA and HADDOCK scores, LRMSD success rate, LRMSD, on-target and off-target pAE, and IDDT and diversity, in that order. The horizontal axes indicate cumulative FlexEvo optimization time per sample. The recorded trajectories show increasing DockQ and LRMSD success rates, decreasing LRMSD and on-target pAE, and increasing diversity over optimization, with local fluctuations and late-stage stabilization. Differences in turning points and convergence profiles illustrate category-dependent optimization behavior. Energy-based metrics are interpreted according to their metric-specific preferred directions. Panels displaying paired metrics use separate left and right axes, with colors matched to the corresponding legends.
Figure 30: Optimization trajectories across antibody fragments and nanobodies, oligonucleotide drugs, and long nucleic acid drugs such as mRNA. Panels a–f, g–l, and m–r correspond to the three categories, respectively. Within each category, the six panels report DockQ, MM/GBSA and HADDOCK scores, LRMSD success rate, LRMSD, on-target and off-target pAE, and IDDT and diversity, in that order. The horizontal axes indicate cumulative FlexEvo optimization time per sample. The recorded trajectories show higher DockQ and LRMSD success rates, lower LRMSD and on-target pAE, and increased diversity over optimization, together with local fluctuations and category-dependent convergence profiles. Energy and pAE panels use separate left and right axes, with colors identifying the corresponding metrics.
Protein binder design has largely optimized for affinity alone, leaving conformational selectivity unaddressed: for allosteric targets such as kinases, nuclear receptors, and GPCRs, a binder that engages both active and inactive states provides no functional specificity regardless of how tightly it binds. We introduce AlloGen, a modular framework that decouples backbone generation from a learned state-selectivity scorer Qθ, an SE(3)-invariant interface graph transformer trained via a two-phase curriculum that first learns interface geometry before imposing conformational discrimination. Because Qθ is fully differentiable and generator-agnostic, it integrates with any backbone generator as a passive reranker or an active gradient-based guide without retraining. Across a diverse benchmark of proteins spanning multiple families and conformational mechanisms, AlloGen consistently identifies binders that preferentially recognize desired structural states while rejecting alternative conformations. Experimental validation on calmodulin further demonstrates that these computational selectivity signals translate to physical molecules, yielding de novo peptides that bind the desired holo conformation while exhibiting no detectable binding to the apo state. Together, these results establish conformational selectivity as a learnable property and provide a general framework for state-selective protein binder design.
Hanqun Cao, Zachary Quinn, Aastha Pal +4
Department of Computer Science and Engineering, The Chinese University of Hong Kong · Department of Bioengineering, University of Pennsylvania · Department of Computer and Information Science, University of Pennsylvania
The rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in structural biology, by facilitating end-to-end generation via joint sequence-structure modeling or hallucination. However, existing approaches are predominantly implemented under a single-target, single-state assumption, limiting their ability to model multi-target or multi-state interactions required for advanced function-oriented protein design. Here, we introduce Chamaileon, which unifies multi-target and multi-state binder design by formulating the problem as cross-context binding landscape modeling. The framework is underpinned by a training paradigm termed In-Context Complex Co-Design (I3CD) for context-aware sequence-structure co-modeling. During inference, we employ Mixture-of-Paths Sampling (MoPS), a scalable strategy that optimizes a single sequence across contexts while alleviating the scarcity of high-quality multi-conformational paired data. Extensive evaluation on our newly constructed benchmark, CROSS, demonstrates that Chamaileon effectively generates sequences adaptable to diverse conformational landscapes and multi-target requirements. The code is available on https://github.com/caohengyuan/Chamaileon.
Hengyuan Cao, Shizhuo Cheng, Mingxuan Liu +5
Zhejiang University, Hangzhou, China · Fudan University, Shanghai, China · Shanghai Institute for Advanced Study Zhejiang University, Shanghai, China +1
Models from the AlphaFold (AF) family reliably predict one dominant conformation for most well-ordered proteins but struggle to capture biologically relevant alternate states. Several efforts have focused on eliciting greater conformational variability through ad hoc inference-time perturbations of AF models or their inputs. Despite their progress, these approaches remain inefficient and fail to consistently recover major conformational modes. Here, we investigate both the optimal location and manner-of-operation for perturbing latent representations in the AF3 architecture. We distill our findings in ConforNets: channel-wise affine transforms of the pre-Pairformer pair latents. Unlike previous methods, ConforNets globally modulate AF3 representations, making them reusable across proteins. On unsupervised generation of alternate states, ConforNets achieve state-of-the-art success rates on all existing multi-state benchmarks. On the novel supervised task of conformational transfer, ConforNets trained on one source protein can induce a conserved conformational change across a protein family. Collectively, these results introduce a mechanism for conformational control in AF3-based models.
Minji Lee, Colin Kalicki, Minkyu Jeon +3
Department of Computer Science, Columbia University, NY, USA · Department of Systems Biology, Columbia University, NY, USA · Department of Computer Science, Princeton University, Princeton, NJ, USA