Structure-aware Reinforcement Learning for Protein Directed Evolution
Authors: Zikun Nie, Suyuan Zhao, Yizhen Luo, Siqi Fan, Zaiqing Nie
Organizations: Institute of AI Industry Research (AIR), Tsinghua University · Department of Computer Science and Technology, Tsinghua University · Pharmolix Inc.
Protein optimization remains a longstanding goal in life sciences. Existing machine learning-assisted directed evolution (MLDE) methods primarily rely on sequence-only features, overlooking the critical spatial constraints and co-evolutionary interactions encoded in protein structures. However, directly integrating structural information remains challenging due to the scarcity of reliable mutant structures. To address these issues, we propose StructEvo, a novel structure-aware reinforcement learning framework for protein directed evolution. StructEvo employs a delta-structure fusion encoder to approximate mutant structure features via feature differences, enabling dynamic incorporation of spatial knowledge. The vast mutation space is then decomposed into manageable subspaces through a structure-aligned hierarchical action network, while a geometric constraint further stabilizes delta feature learning. Our approach outperforms prior state-of-the-art methods by 9.2% and 16.3% on two challenging optimization benchmarks, and further identifies an experimentally validated epistasis pattern in GFP, highlighting the importance of structural guidance for effective protein directed evolution.
Figures & tables
Figure 1 : (a) Illustration of the active learning pipeline for directed evolution. Starting from an initial pool, candidates are evaluated by an oracle and then update the pool, after which a proxy is iteratively refined to guide candidate proposals. The figure highlights two key challenges for the proposal policy: how to model the fine-grained local fitness landscape, and how to effectively leverage structural knowledge. (b) Performance comparison on GFP-hard benchmark. StructEvo achieves the highest fitness while maintaining comparable diversity to other high-fitness baselines.
Figure 2 : Architecture of the StructEvo framework. A delta-structure fusion encoder approximates mutant structure features via delta representations and integrates sequence and structure information. A structure-aligned hierarchical action network then decomposes action space and enables structure-guided mutation sampling. A geometric constraint loss is calculated to improve robustness. The proxy model remains frozen and provides reward signals during policy updates. A.A.: amino acid.
GB1
PhoQ
Method
Max ↑
Mean ↑
NDCG ↑
Max ↑
Mean ↑
NDCG ↑
CMAES
0.59 ± 0.17
0.13 ± 0.04
0.75 ± 0.02
0.29 ± 0.14
0.05 ± 0.01
0.69 ± 0.02
BO
0.67 ± 0.10
0.17 ± 0.02
0.77 ± 0.01
0.29 ± 0.06
0.05 ± 0.01
0.70 ± 0.01
AdaLead
0.64 ± 0.13
0.30 ± 0.09
0.74 ± 0.02
0.27 ± 0.11
0.14 ± 0.02
0.67 ± 0.02
PEX
0.66 ± 0.03
0.36 ± 0.02
0.76 ± 0.01
0.40 ± 0.03
0.10 ± 0.02
0.71 ± 0.02
MLDE ∗
0.68 ± n.a.
0.20 ± n.a.
0.79 ± n.a.
0.36 ± n.a.
0.10 ± n.a.
0.79 ± n.a.
Table 1 : Results on GB1 and PhoQ benchmarks . We report average results and std across 5 independent runs. ∗ : results are derived from Wang et al. [16] . n.a.: not reported in the original paper.
AAV medium
AAV hard
Method
Mean ↑
Max ↑
Diversity
Novelty
Mean ↑
Max ↑
Diversity
Novelty
CMA-ES
0.05 ± 0.00
0.40 ± 0.01
20.9 ± 0.5
17.2 ± 0.5
0.05 ± 0.00
0.32 ± 0.03
20.4 ± 0.7
18.6 ± 0.5
BO
0.64 ± 0.03
0.72 ± 0.04
8.1 ± 0.2
9.4 ± 0.9
0.62 ± 0.03
0.70 ± 0.03
8.6 ± 0.2
10.3 ± 0.3
PEX
0.65 ± 0.01
0.74 ± 0.00
5.5 ± 0.5
5.3 ± 0.3
0.63 ± 0.02
0.75 ± 0.01
5.7 ± 0.5
6.2 ± 0.4
GGS ∗
0.51 ± 0.01
-
4.0 ± 0.2
5.4 ± 0.5
0.60 ± 0.02
-
4.5 ± 0.5
7.0 ± 0.0
AdaLead
0.74 ± 0.04
0.80 ± 0.05
5.2 ± 0.9
7.6 ± 0.4
0.76 ± 0.03
0.81 ± 0.03
4.0 ± 0.4
8.9 ± 0.6
Table 2 : Performance on AAV and GFP benchmarks . The results include average and standard deviation across 5 independent runs. ∗ : results are derived from Kirjner et al. [30] . † : results are derived from Lee et al. [13] . -: not reported in the original paper.
Figure 3 : Visualization of mean (top row) and maximum (bottom row) fitness across rounds. The curves indicate the average results, and the shaded regions indicate the standard deviation.
Figure 6
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Symbol
Definition
Protein Representation
L
the protein length
X
the protein sequence
S
the protein structure
X
the mutation space
V
the amino acid vocabulary
xi
the amino acid identity of residue i
Appendix
Table 4 : Definition of symbols used in our work.
Name
UniProt ID
Protein Length
Dataset Size
Mutation Sites
<0.3
<0.1
= 0
GB1
P19909
56aa
149,361
V39,D40, G41,V54
99.30%
97.30%
19.74%
PhoQ
P23837
486aa
140,517
A284,V285, S288,T289
99.96%
98.65%
63.70%
Appendix
Table 5 : Description of GB1 and PhoQ benchmarks. The last three columns report the proportion of mutants whose normalized fitness is lower than a given threshold, highlighting the sparsity and challenge of these benchmarks. aa: amino acids.
Level
Range(%)
Gap
Size
Mean Fitness
Medium
20-40th
6
2139
0.376
Hard
<30th
7
3448
0.326
Appendix
Table 6 : Description of initial sets in the AAV benchmark. Mean Fitness is the mean normalized fitness of top 128 sequences in subsets.
Level
Range(%)
Gap
Size
Mean Fitness
Medium
20-40th
6
2828
0.232
Hard
<30th
7
2426
0.092
Appendix
Table 7 : Description of initial sets in the GFP benchmark. Mean Fitness is the mean normalized fitness of top 128 sequences in subsets.
Figure 5 : Kernel density estimation (KDE) of fitness distributions for AAV and GFP benchmarks.
Protein
Source
Structure ID
Chain ID
Position
Length
GB1
PDB
2GI9
A
1–56
56
PhoQ
AFDB
AF-Q8FIB8-F1-v6
A
188–376
189
AAV
PDB
3NG9
A
561–588
28
GFP
PDB
1GFL
A
2–238
237
Appendix
Table 8 : Information of the reference protein structures used in our work.
Round 1
Round 2
Round 3
Method
Max ↑
Mean ↑
NDCG ↑
Max ↑
Mean ↑
NDCG ↑
Max ↑
Mean ↑
NDCG ↑
MLDE
0.650
0.183
0.767
0.680
0.217
0.794
0.684
0.203
0.789
ftMLDE (EVmutation)
0.725
0.233
0.791
0.770
0.280
0.814
0.935
0.414
0.833
ftMLDE (Transformer)
0.761
0.239
0.792
0.814
0.298
0.819
0.932
0.416
0.813
EvoPlay
0.837
0.433
0.826
0.834
0.457
0.839
0.840
0.460
0.847
CLADE
0.785
0.309
0.801
0.802
0.303
0.803
0.835
0.458
0.857
Appendix
Table 9 : Results on GB1 benchmarks across 3 rounds . The best scores bolded and the second best underlined . All baseline results are derived from Wang et al. [16] .
Round 1
Round 2
Round 3
Method
Max ↑
Mean ↑
NDCG ↑
Max ↑
Mean ↑
NDCG ↑
Max ↑
Mean ↑
NDCG ↑
MLDE
0.309
0.069
0.753
0.364
0.087
0.775
0.361
0.095
0.791
ftMLDE (EVmutation)
0.297
0.072
0.754
0.431
0.114
0.807
0.436
0.115
0.804
ftMLDE (Transformer)
0.346
0.073
0.756
0.414
0.108
0.802
0.422
0.117
0.815
EvoPlay
0.443
0.119
0.782
0.444
0.135
0.786
0.474
0.143
0.804
CLADE
0.319
0.070
0.759
0.441
0.089
0.762
0.467
0.089
0.777
Appendix
Table 10 : Results on PhoQ benchmarks across 3 rounds . The best scores bolded and the second best underlined . All baseline results are derived from Wang et al. [16] .
Method
Mean ↑
Max ↑
Diversity
Novelty
AAV medium
StructEvo
0.82 ± 0.01
0.86 ± 0.01
4.5 ± 0.7
8.6 ± 0.2
(1) w/o delta
0.79 ± 0.00
0.83 ± 0.01
5.5 ± 0.8
7.7 ± 0.0
(2) w/o structure
0.77 ± 0.01
0.81 ± 0.02
4.7 ± 1.3
8.1 ± 0.7
(3) w/o reference
0.77 ± 0.01
0.81 ± 0.02
4.8 ± 2.0
7.8 ± 0.0
(4) w/o hierarchy
0.78 ± 0.01
0.80 ± 0.01
5.3 ± 1.3
7.9 ± 0.6
AAV hard
StructEvo
0.83 ± 0.03
0.87 ± 0.04
4.5 ± 0.3
8.4 ± 0.5
Appendix
Table 11 : Full results of ablation studies. w/o: without.
Figure 6 : Correlation between Confidence Scores and Fitness.
Figure 7 : Samples of reference-based difference (left) and pairwise difference (right).
Mean ↑
Max ↑
Diversity
Novelty
Default
1.00 ± 0.04
1.05 ± 0.05
4.9 ± 0.3
11.4 ± 0.6
Oracle-full
1.07 ± 0.01
1.12 ± 0.02
12.9 ± 2.9
13.1 ± 1.2
Oracle-limited
0.41 ± 0.03
0.49 ± 0.04
16.7 ± 4.2
5.8 ± 0.9
Appendix
Table 12 : Proxy robustness analysis on GFP-hard task.
Mean ↑
Max ↑
Diversity
Novelty
λgeo=0
0.984 ± 0.081
1.017 ± 0.072
5.101 ± 1.087
9.648 ± 1.795
λgeo=0.01
0.987 ± 0.052
1.031 ± 0.036
4.878 ± 0.312
11.427 ± 1.134
λgeo=0.05 (default)
0.998 ± 0.041
1.047 ± 0.046
4.868 ± 0.336
11.425 ± 0.602
λgeo=0.1
0.970 ± 0.041
1.010 ± 0.043
5.289 ± 0.358
10.596 ± 0.846
λgeo=0.5
0.925 ± 0.043
0.991 ± 0.031
6.452 ± 1.534
10.677 ± 0.785
λgeo=1
0.931 ± 0.034
0.974 ± 0.027
5.115 ± 0.970
9.706 ± 0.115
Appendix
Table 13 : Hyperparameter analysis of geometric constraint weight on GFP-hard task.