Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution
Organizations: College of Computing and Data Science, Nanyang Technological University, Singapore · Sun Yat-sen University, China · Zhejiang University, China · Carnegie Mellon University, USA · Shanghai Institute of Optics and Fine Mechanics, China · Zhongnan Hospital, Wuhan University, China · University of Southern California, USA
Abstract
Model-guided directed evolution seeks to identify high-fitness protein variants under limited oracle budgets. Protein language models (PLMs) provide rich representations for this task, but task-agnostic zero-shot scores can be misaligned with a target assay, while supervised search in high-dimensional embedding spaces can make surrogate modeling and uncertainty estimation sample-inefficient. We propose the Linear Fitness Subspace (LFS) hypothesis: within mutation-induced residue-level representation changes, a compact, assay-specific set of directions makes fitness variation linearly accessible from few labeled variants. This is a local, supervision-recoverable statement rather than a claim that protein fitness landscapes or global PLM geometry are universally linear. Building on this observation, we introduce Subspace-Guided Evolutionary Search (SGES), which estimates an LFS from a small initial sample and performs surrogate modeling, uncertainty estimation, and acquisition in the learned subspace. Across 10 core ProteinGym assays, 87 extended static-validation assays, and an 18-assay budgeted-search evaluation, SGES improves fitness prediction and search efficiency over zero-shot PLMs and recent ML-guided protein optimization baselines. Controlled comparisons with PCA, random projections, label-shuffled PLS, classical mutation features, and acquisition ablations further isolate the benefit of a fitness-aligned site-delta coordinate.
Figures & tables
| Assay | Ridge | Mean-pool | MLP | Ratio |
| SPG1_STRSG | 0.945 | 0.620 | 0.973 | 0.971 |
| BLAT_Firnberg | 0.865 | 0.730 | 0.903 | 0.958 |
| PABP_YEAST | 0.840 | 0.580 | 0.916 | 0.917 |
| P53_HUMAN | 0.810 | 0.650 | 0.874 | 0.927 |
| RASH_HUMAN | 0.755 | 0.590 | 0.825 | 0.915 |
| AMIE_PSEAE | 0.758 | 0.550 | 0.823 | 0.921 |
| Method | Avg. Best Fitness | Avg. Rank |
| SGES (Ours) | 1.471 | 1.38 |
| GP-BO ( Soldát and Kléma, 2024 ) | 1.328 | 2.38 |
| AlphaDE ( Yang et al., 2025 ) | 1.267 | 2.63 |
| LatentDE ( Tran et al., 2025 ) | 1.223 | 3.50 |
| EVOLVEpro ( Jiang et al., 2024 ) | 1.183 | 5.00 |
| BOES ( Soldát and Kléma, 2024 ) | 1.173 | 4.75 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Assay | Ours 50% | Ours 5% | ESM2-650M | ESM2-15B | Tranc.-L | TrEVE-L | EVE | MSA-T | GEMME | ProSST |
| SPG1_STRSG | 0.973 | 0.915 | 0.301 | 0.337 | 0.289 | 0.290 | 0.272 | 0.171 | 0.283 | 0.743 |
| BLAT_Firnberg | 0.903 | 0.785 | 0.737 | 0.424 | 0.637 | 0.736 | 0.729 | 0.741 | 0.683 | 0.766 |
| PABP_YEAST | 0.916 | 0.800 | 0.716 | 0.695 | 0.688 | 0.684 | 0.648 | 0.663 | 0.674 | 0.716 |
| P53_HUMAN | 0.874 | 0.713 | 0.432 | 0.497 | 0.400 | 0.413 | 0.427 | 0.291 | 0.423 | 0.425 |
| RASH_HUMAN | 0.825 | 0.606 | 0.498 | 0.313 | 0.450 | 0.487 | 0.480 | 0.446 | 0.434 | 0.619 |
| AMIE_PSEAE | 0.823 | 0.625 | 0.557 | 0.613 | 0.438 | 0.481 | 0.464 | 0.605 | 0.558 | 0.486 |
| Variant / Feature | Panel | Surrogate | Acquisition | Avg. | Avg. Best | |
| Random projection | A | 8 | Deep ensemble | UCB | 0.512 | 1.08 |
| PCA projection | A | 8 | Deep ensemble | UCB | 0.624 | 1.22 |
| Label-shuffled PLS | A | 8 | Deep ensemble | UCB | 0.485 | 1.02 |
| PLS-LFS projection | A | 8 | Deep ensemble | UCB | 0.804 | 1.45 |
| Full site-delta feature | A | 1280 | Deep ensemble | UCB | 0.763 | 1.32 |
| PLS-LFS + ridge top- | B | 8 | Ridge | Greedy mean | 0.686 | 1.15 |
| Method | Budget | Initial set | Candidate / proposal setting | Representation | Surrogate / acquisition | Seeds |
| Random Search | 500 | Same | Same mutation search space | No learned PLM surrogate | Random | 5 |
| AdaLead | 500 | Same | Local evolutionary proposal | No explicit LFS | Fitness-guided local search | 5 |
| GP-BO | 500 | Same | Same candidate space | Full PLM / high-dimensional | GP Bayesian optimization | 5 |
| EVOLVEpro | 500 | Same | PLM-guided proposal | PLM supervised signal | Method-specific search | 5 |
| AlphaDE | 500 | Same | Directed-evolution proposal | Method-specific | Method-specific acquisition | 5 |
| BOES | 500 | Same | Budgeted search | Embedding-space | BO / evolutionary search | 5 |
| Assay set | Count | Primary role |
| Core assays | 10 | Detailed LFS analysis / ablations |
| Extended assays | 87 | Static large-scale prediction validation |
| Search-statistics assays | 18 | Multi-seed search metrics / Wilcoxon tests |
| Multi-mutant benchmarks | 5 | Combinatorial transfer evaluation |
| Assay | SGES | EVOLVEpro | SGES Best | EVOLVEpro Best |
| SPG1_STRSG | 0.973 | 0.865 | – | – |
| BLAT_Firnberg | 0.903 | 0.812 | 1.691 | 1.410 |
| PABP_YEAST | 0.916 | 0.808 | 1.170 | 1.055 |
| P53_HUMAN | 0.874 | 0.755 | 2.690 | 2.150 |
| RASH_HUMAN | 0.825 | 0.720 | 1.525 | 1.180 |
| AMIE_PSEAE | 0.823 | 0.705 | 0.785 | 0.610 |
| Search Metric | SGES vs GP-BO | SGES vs AdaLead | SGES vs Random | SGES vs EVOLVEpro |
| AUC (Best-so-far) | ||||
| Best Fitness Found | ||||
| Budget-to-Top-5% |
| Assay | SGES | AlphaDE | BOES | LatentDE | TreeNeuralUCB | GP-BO | EVOLVEpro |
| BLAT_Firnberg | |||||||
| P53_HUMAN | |||||||
| PABP_YEAST | |||||||
| RASH_HUMAN | |||||||
| AMIE_PSEAE | |||||||
| DYR_ECOLI |
| Method | Runtime / Seed | GPU Mem. | Avg. Best |
| (min) | (GB) | Fitness | |
| Random Search | 2 | 1.2 | 1.01 |
| AdaLead ( Sinai et al., 2020 ) | 15 | 2.5 | 1.25 |
| EVOLVEpro ( Jiang et al., 2024 ) | 42 | 4.8 | 1.18 |
| SGES (Ours) | 28 | 3.5 | 1.68 |
| GP-BO ( Soldát and Kléma, 2024 ) | 245 | 18.5 | 1.33 |
| Method | Mean | Median | Ratio | vs. ZS |
| Ridge (site-delta) | 0.584 | 0.612 | 35.6% | +0.146 |
| Ridge (mean-pooled) | 0.385 | 0.392 | 8.0% | -0.053 |
| Zero-shot ESM2 | 0.438 | 0.451 | 14.9% | Baseline |
| Group | Mean cos | Max cos | Min cos |
| Same Protein (BLAT F vs BLAT J) | 0.68 | 0.68 | 0.68 |
| Same Function (Enzyme vs Enzyme) | 0.22 | 0.35 | 0.12 |
| Different Function | 0.05 | 0.11 | 0.01 |
| All Pairs | 0.14 | 0.68 | 0.01 |
| Train/Test | BLAT-F | BLAT-J | DYR | AMIE | RASH | P53 | PTEN | SPG1 | PABP | HIS7 |
| BLAT-F | 0.865 | 0.624 | 0.045 | 0.012 | -0.015 | -0.032 | 0.041 | -0.018 | 0.022 | 0.015 |
| BLAT-J | 0.735 | 0.650 | -0.012 | 0.025 | 0.010 | 0.014 | 0.025 | -0.005 | -0.011 | 0.030 |
| DYR | -0.022 | 0.015 | 0.675 | 0.085 | 0.042 | 0.025 | 0.055 | -0.022 | 0.014 | 0.045 |
| AMIE | 0.010 | 0.033 | 0.051 | 0.758 | 0.022 | -0.015 | 0.018 | 0.010 | -0.005 | 0.020 |
| RASH | -0.015 | 0.011 | 0.022 | -0.018 | 0.755 | 0.085 | 0.042 | 0.015 | 0.035 | 0.012 |
| P53 | -0.005 | -0.020 | 0.018 | 0.010 | 0.045 | 0.810 | 0.065 | 0.022 | -0.012 | 0.008 |
| Assay | #Muts | SGES MAE | ZS MAE | SGES | ZS |
| PABP_YEAST | 2–3 | 0.450 | 1.240 | 0.520 | 0.180 |
| SPG1_STRSG | 2–3 | 0.280 | 0.850 | 0.650 | 0.220 |
| GFP | 2–8 | 0.685 | 1.585 | 0.420 | 0.115 |
| GB1 | 2–4 | 0.515 | 1.120 | 0.485 | 0.140 |
| AAV | 2–7 | 0.590 | 1.345 | 0.465 | 0.125 |
| Average | – | 0.504 | 1.228 | 0.508 | 0.156 |
| Assay | SGES | PCNN | SGES Best | PCNN Best |
| BLAT_Firnberg | 0.903 | 0.780 | 1.691 | 1.350 |
| P53_HUMAN | 0.874 | 0.715 | 2.690 | 2.100 |
| RASH_HUMAN | 0.825 | 0.680 | 1.525 | 1.150 |
| PTEN_HUMAN | 0.593 | 0.520 | 1.939 | 1.550 |
| Average | 0.825 | 0.685 | 1.471 | 1.215 |
| Assay | Top Mutations | Predicted Mechanism | Score |
| BLAT_Firnberg | M182T, E104K, G238S | Stabilizing / extended spectrum | 9.2 |
| P53_HUMAN | V143A, Y220C | Core packing restoration | 8.5 |
| AMIE_PSEAE | S293T, F154Y | Active-site volume optimization | 8.8 |