Structure-aware Reinforcement Learning for Protein Directed Evolution
Authors: Zikun Nie, Suyuan Zhao, Yizhen Luo, Siqi Fan, Zaiqing Nie
Organizations: Institute of AI Industry Research (AIR), Tsinghua University · Department of Computer Science and Technology, Tsinghua University · Pharmolix Inc.
Protein optimization remains a longstanding goal in life sciences. Existing machine learning-assisted directed evolution (MLDE) methods primarily rely on sequence-only features, overlooking the critical spatial constraints and co-evolutionary interactions encoded in protein structures. However, directly integrating structural information remains challenging due to the scarcity of reliable mutant structures. To address these issues, we propose StructEvo, a novel structure-aware reinforcement learning framework for protein directed evolution. StructEvo employs a delta-structure fusion encoder to approximate mutant structure features via feature differences, enabling dynamic incorporation of spatial knowledge. The vast mutation space is then decomposed into manageable subspaces through a structure-aligned hierarchical action network, while a geometric constraint further stabilizes delta feature learning. Our approach outperforms prior state-of-the-art methods by 9.2% and 16.3% on two challenging optimization benchmarks, and further identifies an experimentally validated epistasis pattern in GFP, highlighting the importance of structural guidance for effective protein directed evolution.
Figures & tables
Figure 1 : (a) Illustration of the active learning pipeline for directed evolution. Starting from an initial pool, candidates are evaluated by an oracle and then update the pool, after which a proxy is iteratively refined to guide candidate proposals. The figure highlights two key challenges for the proposal policy: how to model the fine-grained local fitness landscape, and how to effectively leverage structural knowledge. (b) Performance comparison on GFP-hard benchmark. StructEvo achieves the highest fitness while maintaining comparable diversity to other high-fitness baselines.
Figure 2 : Architecture of the StructEvo framework. A delta-structure fusion encoder approximates mutant structure features via delta representations and integrates sequence and structure information. A structure-aligned hierarchical action network then decomposes action space and enables structure-guided mutation sampling. A geometric constraint loss is calculated to improve robustness. The proxy model remains frozen and provides reward signals during policy updates. A.A.: amino acid.
GB1
PhoQ
Method
Max ↑
Mean ↑
NDCG ↑
Max ↑
Mean ↑
NDCG ↑
CMAES
0.59 ± 0.17
0.13 ± 0.04
0.75 ± 0.02
0.29 ± 0.14
0.05 ± 0.01
0.69 ± 0.02
BO
0.67 ± 0.10
0.17 ± 0.02
0.77 ± 0.01
0.29 ± 0.06
0.05 ± 0.01
0.70 ± 0.01
AdaLead
0.64 ± 0.13
0.30 ± 0.09
0.74 ± 0.02
0.27 ± 0.11
0.14 ± 0.02
0.67 ± 0.02
PEX
0.66 ± 0.03
0.36 ± 0.02
0.76 ± 0.01
0.40 ± 0.03
0.10 ± 0.02
0.71 ± 0.02
MLDE ∗
0.68 ± n.a.
0.20 ± n.a.
0.79 ± n.a.
0.36 ± n.a.
0.10 ± n.a.
0.79 ± n.a.
Table 1 : Results on GB1 and PhoQ benchmarks . We report average results and std across 5 independent runs. ∗ : results are derived from Wang et al. [16] . n.a.: not reported in the original paper.
AAV medium
AAV hard
Method
Mean ↑
Max ↑
Diversity
Novelty
Mean ↑
Max ↑
Diversity
Novelty
CMA-ES
0.05 ± 0.00
0.40 ± 0.01
20.9 ± 0.5
17.2 ± 0.5
0.05 ± 0.00
0.32 ± 0.03
20.4 ± 0.7
18.6 ± 0.5
BO
0.64 ± 0.03
0.72 ± 0.04
8.1 ± 0.2
9.4 ± 0.9
0.62 ± 0.03
0.70 ± 0.03
8.6 ± 0.2
10.3 ± 0.3
PEX
0.65 ± 0.01
0.74 ± 0.00
5.5 ± 0.5
5.3 ± 0.3
0.63 ± 0.02
0.75 ± 0.01
5.7 ± 0.5
6.2 ± 0.4
GGS ∗
0.51 ± 0.01
-
4.0 ± 0.2
5.4 ± 0.5
0.60 ± 0.02
-
4.5 ± 0.5
7.0 ± 0.0
AdaLead
0.74 ± 0.04
0.80 ± 0.05
5.2 ± 0.9
7.6 ± 0.4
0.76 ± 0.03
0.81 ± 0.03
4.0 ± 0.4
8.9 ± 0.6
Table 2 : Performance on AAV and GFP benchmarks . The results include average and standard deviation across 5 independent runs. ∗ : results are derived from Kirjner et al. [30] . † : results are derived from Lee et al. [13] . -: not reported in the original paper.
Figure 3 : Visualization of mean (top row) and maximum (bottom row) fitness across rounds. The curves indicate the average results, and the shaded regions indicate the standard deviation.
Figure 6
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Symbol
Definition
Protein Representation
L
the protein length
X
the protein sequence
S
the protein structure
X
the mutation space
V
the amino acid vocabulary
xi
the amino acid identity of residue i
Appendix
Table 4 : Definition of symbols used in our work.
Name
UniProt ID
Protein Length
Dataset Size
Mutation Sites
<0.3
<0.1
= 0
GB1
P19909
56aa
149,361
V39,D40, G41,V54
99.30%
97.30%
19.74%
PhoQ
P23837
486aa
140,517
A284,V285, S288,T289
99.96%
98.65%
63.70%
Appendix
Table 5 : Description of GB1 and PhoQ benchmarks. The last three columns report the proportion of mutants whose normalized fitness is lower than a given threshold, highlighting the sparsity and challenge of these benchmarks. aa: amino acids.
Level
Range(%)
Gap
Size
Mean Fitness
Medium
20-40th
6
2139
0.376
Hard
<30th
7
3448
0.326
Appendix
Table 6 : Description of initial sets in the AAV benchmark. Mean Fitness is the mean normalized fitness of top 128 sequences in subsets.
Level
Range(%)
Gap
Size
Mean Fitness
Medium
20-40th
6
2828
0.232
Hard
<30th
7
2426
0.092
Appendix
Table 7 : Description of initial sets in the GFP benchmark. Mean Fitness is the mean normalized fitness of top 128 sequences in subsets.
Figure 5 : Kernel density estimation (KDE) of fitness distributions for AAV and GFP benchmarks.
Protein
Source
Structure ID
Chain ID
Position
Length
GB1
PDB
2GI9
A
1–56
56
PhoQ
AFDB
AF-Q8FIB8-F1-v6
A
188–376
189
AAV
PDB
3NG9
A
561–588
28
GFP
PDB
1GFL
A
2–238
237
Appendix
Table 8 : Information of the reference protein structures used in our work.
Round 1
Round 2
Round 3
Method
Max ↑
Mean ↑
NDCG ↑
Max ↑
Mean ↑
NDCG ↑
Max ↑
Mean ↑
NDCG ↑
MLDE
0.650
0.183
0.767
0.680
0.217
0.794
0.684
0.203
0.789
ftMLDE (EVmutation)
0.725
0.233
0.791
0.770
0.280
0.814
0.935
0.414
0.833
ftMLDE (Transformer)
0.761
0.239
0.792
0.814
0.298
0.819
0.932
0.416
0.813
EvoPlay
0.837
0.433
0.826
0.834
0.457
0.839
0.840
0.460
0.847
CLADE
0.785
0.309
0.801
0.802
0.303
0.803
0.835
0.458
0.857
Appendix
Table 9 : Results on GB1 benchmarks across 3 rounds . The best scores bolded and the second best underlined . All baseline results are derived from Wang et al. [16] .
Round 1
Round 2
Round 3
Method
Max ↑
Mean ↑
NDCG ↑
Max ↑
Mean ↑
NDCG ↑
Max ↑
Mean ↑
NDCG ↑
MLDE
0.309
0.069
0.753
0.364
0.087
0.775
0.361
0.095
0.791
ftMLDE (EVmutation)
0.297
0.072
0.754
0.431
0.114
0.807
0.436
0.115
0.804
ftMLDE (Transformer)
0.346
0.073
0.756
0.414
0.108
0.802
0.422
0.117
0.815
EvoPlay
0.443
0.119
0.782
0.444
0.135
0.786
0.474
0.143
0.804
CLADE
0.319
0.070
0.759
0.441
0.089
0.762
0.467
0.089
0.777
Appendix
Table 10 : Results on PhoQ benchmarks across 3 rounds . The best scores bolded and the second best underlined . All baseline results are derived from Wang et al. [16] .
Method
Mean ↑
Max ↑
Diversity
Novelty
AAV medium
StructEvo
0.82 ± 0.01
0.86 ± 0.01
4.5 ± 0.7
8.6 ± 0.2
(1) w/o delta
0.79 ± 0.00
0.83 ± 0.01
5.5 ± 0.8
7.7 ± 0.0
(2) w/o structure
0.77 ± 0.01
0.81 ± 0.02
4.7 ± 1.3
8.1 ± 0.7
(3) w/o reference
0.77 ± 0.01
0.81 ± 0.02
4.8 ± 2.0
7.8 ± 0.0
(4) w/o hierarchy
0.78 ± 0.01
0.80 ± 0.01
5.3 ± 1.3
7.9 ± 0.6
AAV hard
StructEvo
0.83 ± 0.03
0.87 ± 0.04
4.5 ± 0.3
8.4 ± 0.5
Appendix
Table 11 : Full results of ablation studies. w/o: without.
Figure 6 : Correlation between Confidence Scores and Fitness.
Figure 7 : Samples of reference-based difference (left) and pairwise difference (right).
Mean ↑
Max ↑
Diversity
Novelty
Default
1.00 ± 0.04
1.05 ± 0.05
4.9 ± 0.3
11.4 ± 0.6
Oracle-full
1.07 ± 0.01
1.12 ± 0.02
12.9 ± 2.9
13.1 ± 1.2
Oracle-limited
0.41 ± 0.03
0.49 ± 0.04
16.7 ± 4.2
5.8 ± 0.9
Appendix
Table 12 : Proxy robustness analysis on GFP-hard task.
Mean ↑
Max ↑
Diversity
Novelty
λgeo=0
0.984 ± 0.081
1.017 ± 0.072
5.101 ± 1.087
9.648 ± 1.795
λgeo=0.01
0.987 ± 0.052
1.031 ± 0.036
4.878 ± 0.312
11.427 ± 1.134
λgeo=0.05 (default)
0.998 ± 0.041
1.047 ± 0.046
4.868 ± 0.336
11.425 ± 0.602
λgeo=0.1
0.970 ± 0.041
1.010 ± 0.043
5.289 ± 0.358
10.596 ± 0.846
λgeo=0.5
0.925 ± 0.043
0.991 ± 0.031
6.452 ± 1.534
10.677 ± 0.785
λgeo=1
0.931 ± 0.034
0.974 ± 0.027
5.115 ± 0.970
9.706 ± 0.115
Appendix
Table 13 : Hyperparameter analysis of geometric constraint weight on GFP-hard task.
The last five-plus years have seen many protein engineering disciplines transformed by advances in machine learning (ML), but the same cannot be said for directed evolution. Reflecting on a previously co-authored perspective, I discuss why I believe this to be the case, arguing that a disconnect between the goals of machine-learning-assisted directed evolution (MLDE) researchers--"identify an optimal protein"--and the goals of directed evolution more broadly--"identify a sufficient protein given time and resource constraints"--is a principal culprit. As an example, I highlight how nearly all current MLDE methods neglect to account for the cost of DNA synthesis, resulting in strategies that have limited practical applicability regardless of the underlying models' capabilities. I close by discussing recent works that are exceptions to this overarching trend, and emphasize that the last five years of efforts in ML-assisted protein engineering and the prescribed reframe of MLDE objectives need not be mutually exclusive.
Protein inverse folding aims to recover amino acid sequences for a given 3D protein structure, underpinning broad applications such as enzyme engineering and drug discovery.Current methods often follow a serial pipeline, in which a structure encoder predicts a coarse sequence, which is then refined by protein language models (PLMs). However, because PLMs only perform post-hoc sequence edits, the refinement is bounded by the quality of upstream predictions.Thanks to recent multimodal protein language models (MPLMs), we could directly encode structure to generate sequences with pretrained structural knowledge, but we observe that they are not effective for inverse folding. Therefore, we introduce a symmetric dual-path architecture that both leverages PLMs for pretrained sequence evolution knowledge and MPLMs for pretrained structural knowledge to iteratively guide protein sequence generation.Through extensive experiments across standard protein inverse folding benchmarks, our method achieves state-of-the-art performance, surpassing prior approaches, and ablation studies validate the rationale of our symmetric design, revealing a promising direction for the community.
Handong Wang, Jiaxin Qi, Baisheng Lai +1
Computer Network Information Center, Chinese Academy of Sciences · University of Chinese Academy of Sciences
Proteins are shaped by gradual evolution under biophysical and functional constraints. Protein language models learn rich evolutionary constraints from large-scale sequences, and discrete diffusion-based protein language models~(\eg, DPLMs) are promising for both understanding and generation. However, existing DPLMs typically rely on masked diffusion that contradicts a simple biological intuition: proteins evolve through accumulated edits, not by emerging from masks. Consequently, these frameworks lack explicit pretraining objectives for substitution and insertion/deletion (indel) operations, limiting both optimization-style post-editing and flexible guided generation. To address these limitations, we present DPLM-Evo, an evolutionary discrete diffusion framework that explicitly predicts substitution, insertion, and deletion operations during denoising. DPLM-Evo decouples an upsampled-length latent alignment space from the variable-length observed sequence space, which makes indel-aware generation tractable. To better align substitutions with real evolution, we further introduce a contextualized evolutionary noising kernel that produces biologically informed, context-dependent mutation patterns. Across tasks, DPLM-Evo improves sequence understanding and achieves state-of-the-art mutation effect prediction performance on ProteinGym in the single-sequence setting. It also enables variable-length simulated evolution, and post-editing/optimization of existing proteins via explicit edit trajectories.