q-bio.QMOct 6, 2026

Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution

Authors: SiYuan Ma, Canran Xiao, Zikai Xiao, Albert Gao, Liang He, Xuan-Yu Wang, Shuying Cao, Xiaojun Jia

Organizations: College of Computing and Data Science, Nanyang Technological University, Singapore · Sun Yat-sen University, China · Zhejiang University, China · Carnegie Mellon University, USA · Shanghai Institute of Optics and Fine Mechanics, China · Zhongnan Hospital, Wuhan University, China · University of Southern California, USA

Abstract

Model-guided directed evolution seeks to identify high-fitness protein variants under limited oracle budgets. Protein language models (PLMs) provide rich representations for this task, but task-agnostic zero-shot scores can be misaligned with a target assay, while supervised search in high-dimensional embedding spaces can make surrogate modeling and uncertainty estimation sample-inefficient. We propose the Linear Fitness Subspace (LFS) hypothesis: within mutation-induced residue-level representation changes, a compact, assay-specific set of directions makes fitness variation linearly accessible from few labeled variants. This is a local, supervision-recoverable statement rather than a claim that protein fitness landscapes or global PLM geometry are universally linear. Building on this observation, we introduce Subspace-Guided Evolutionary Search (SGES), which estimates an LFS from a small initial sample and performs surrogate modeling, uncertainty estimation, and acquisition in the learned subspace. Across 10 core ProteinGym assays, 87 extended static-validation assays, and an 18-assay budgeted-search evaluation, SGES improves fitness prediction and search efficiency over zero-shot PLMs and recent ML-guided protein optimization baselines. Controlled comparisons with PCA, random projections, label-shuffled PLS, classical mutation features, and acquisition ablations further isolate the benefit of a fitness-aligned site-delta coordinate.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing

    Sep 7, 2026Mingrui Li, Sixian Shen, Minzhang Li +4Language Modeling

  2. Towards A Generative Protein Evolution Machine with DPLM-Evo

    Apr 30, 2026Xinyou Wang, Liang Hong, Jiasheng Ye +4AlphaevolveEvolutionary Optimization Methods

  3. PFArena: Benchmarking Language Models for Protein Modification

    Sep 24, 2026Yawen Ouyang, Xinbo Zhang, Ziyuan Ma +15ProteinQuality-Driven Selective Mutation