PFArena: Benchmarking Language Models for Protein Modification
Organizations: Shanghai Artificial Intelligence Laboratory · Generative Symbolic Intelligence Lab (GenSI), Tsinghua University · Institute for AI Industry Research (AIR), Tsinghua University · School of Pharmaceutical Sciences, Tsinghua University
Abstract
Protein modification requires navigating an immense sequence space, yet wet-lab validation remains low-throughput and costly. Although computational paradigms including protein language models (PLMs), large language models (LLMs), and LLM-based agents have shown promise in protein modification, their relative efficacy across realistic experimental decision-making settings remains unclear. To bridge this gap, we introduce PFArena, a benchmark comprising four controlled task interfaces that cover single-mutant generation and multi-mutant ranking. By providing varying levels of mutation fitness data, PFArena reflects four representative research scenarios characterized by differing degrees of prior experimental context. We assess six PLMs, six LLMs, and five LLM-based agents using complementary metrics to measure both peak and overall protein modification performance. Our evaluation reveals that model performance shifts systematically with the availability of target-specific experimental evidence: PLMs demonstrate proficiency in open-ended single-mutant generation by leveraging protein-specific priors, whereas LLMs and agents perform strongly in multi-mutant ranking, particularly when target-specific fitness data are available. Nevertheless, all model families face fundamental challenges with increasing search-space size and mutation depth. We release our code and benchmark suite to facilitate reproducible research in model-assisted protein modification.
Figures & tables
| Task component | Specification |
| Definition | Budgeted generation of legal single substitutions without target-specific mutation measurements. |
| Input | Wild-type sequence , assay context and proposed budget . |
| Output | Proposed mutation set of at most single substitutions. |
| Instance | Input: Wild-type sequence: MSSSLGKE...AMV Assay context: - Primary task class: activity function. - Fitness type: organismal or cellular fitness. - Readout subclass: growth or selection proxy. Proposed budget : 40 Output: Proposed mutation set: P557L , K369Y , N434D , ... . |
| Statistics | Assays: 123; measured single mutants: 594,695. |
| Task component | Specification |
| Definition | Ranking multi-mutant candidates without target-specific mutation measurements. |
| Input | Wild-type sequence , assay context , and a score-blind candidate pool. |
| Output | Complete ranking of the supplied candidate pool. |
| Instance | Input: Wild-type sequence: MLEGKVKW...KEA Assay context: - Primary task class: stability. - Fitness type: stability. - Readout subclass: folding free energy stability readout. Candidate pool: { V28C+V63L , V28Q+V63P , V28M+V63F , ... }. Output: Ranked candidate set: V28C+V63L , V28M+V63F , V28Q+V63P , ... . |
| Statistics | Assays: 74; candidate number: 6,288. |
| Task component | Specification |
| Definition | Ranking strict successors of a measured anchor mutant. |
| Input | Wild-type sequence , assay context , a measured anchor mutant with score , and a score-blind successor pool . |
| Output | Complete ranking of the supplied candidate pool. |
| Instance | Input: Wild-type sequence: QVQLVQSG...VSS Assay context: - Primary task class: binding. - Fitness type: binding. - Readout subclass: binding. Anchor: T28P with score -1.586. Successor pool: { T28P+S30R+N59K+T76A , T28P+S30R+N59K+Q62P+S75F , T28P+S30R+L104V , ... }. Output: Ranked candidate set: T28P+S30R+N59K+Q62P+S75F , T28P+S30R+N59K+T76A , T28P+S30R+L104V , ... . |
| Statistics | Assays: 67; candidate number: 3,648. |
| Task component | Specification |
| Definition | Multi-mutant ranking with measured context for every component substitution. |
| Input | Wild-type sequence , assay context , a measured single-mutant context map , and a score-blind multi-mutant candidate pool . |
| Output | Complete ranking of the supplied candidate pool. |
| Instance | Input: Wild-type sequence: EVKLDETG...EIK Assay context: - Primary task class: stability. - Fitness type: abundance or expression. - Readout subclass: cellular abundance stability proxy. Single-mutant context: { W108E : -0.354, G109P : -1.062, M34Q : 0.605, N35S : 1.154, ... }. Candidate pool: { W108E+G109P , Y102P+M105K , M34Q+N35S , ... }. Output: Ranked candidate set: M34Q+N35S , Y102P+M105K , W108E+G109P , ... . |
| Statistics | Assays: 29; candidate number: 2,638; visible single-mutant context rows: 28,609. |
| Objective | Metric | T1 | T2 | T3 | T4 |
| Global ranking | Spearman Correlation | ||||
| Top-weighted ranking | NDCG | ||||
| Peak discovery | Normalized Maximum Score@ | ||||
| Peak coverage | Recall@ |
| Category | Model | NMS@ | Recall@ |
| Statistic | Random | ||
| PLM | VenusREM | 0.7932 | 0.0427 |
| S3F-MSA | |||
| ProSST-2048 | 0.0407 | ||
| S3F | |||
| ESM-2 |
| Category | Model | Spearman | NDCG | NMS@ | Recall@ |
| Statistic | Random | ||||
| PLM | VenusREM | ||||
| S3F-MSA | |||||
| ProSST-2048 | |||||
| S3F | |||||
| ESM-2 |
| Category | Model | Spearman | NDCG | NMS@ | Recall@ |
| Statistic | Random | ||||
| PLM | VenusREM | ||||
| S3F-MSA | |||||
| ProSST-2048 | |||||
| S3F | |||||
| ESM-2 |
| Setting | Method | NMS@ | Recall@ | Invalid per Assay |
| Vanilla | Sample Avg | |||
| RRF | 0.7847 | 0.0199 | 1.4472 | |
| Clean | Sample Avg | 1.0531 | ||
| RRF | 0.7845 | 0.0197 | ||
| In-Assay | Sample Avg | |||
| RRF | 0.7842 | 0.0191 |
| Model | Inputs and Checkpoint | Core Configuration | Candidate Score |
| ESM-2 [ 1 ] | Wild-type sequence; esm2_t33_650M_UR50D | Masked inference with a 1,024-token context (at most 1,022 residues); deterministic mutation-centered windows for longer chains | Sum of wild-type-context masked-marginal log odds, , over substituted sites |
| ProGen2-base [ 2 ] | Complete wild-type and mutant sequences; progen2-base | Forward and reversed causal sequence scoring with a 2,048-token context; all evaluated chains were processed at full length | Difference between the bidirectional sequence scores of the complete mutant and wild-type chains |
| ProSST-2048 [ 3 ] | Wild-type sequence and AlphaFold 3 structure; AI4Protein/ProSST-2048 | Official GVP quantizer with a 2,048-code structural vocabulary; all evaluated chains were processed at full length within the 2,046-residue limit | Sum of structure-conditioned wild-type-context marginal log odds over substituted sites |
| S3F [ 4 ] | Wild-type sequence, AlphaFold 3 structure, and the corresponding molecular surface; released S3F checkpoint | Mutation-centered windows of at most 1,022 residues; sequence logits replace structure-conditioned logits where AF3 pLDDT is below 70 | Sum of masked-marginal log-odds contributions over substituted sites |
| VenusREM [ 5 ] | Wild-type sequence, ProSST-2048 structural tokens, and a UniRef100/MMseqs2 MSA | aa_seq_aln retrieval; , sampling ratio , and one sampling pass; full-chain inference for all evaluated contexts | Sum of mutant-minus-wild-type values from the fused sequence, structure, and MSA representation |
| S3F-MSA [ 4 , 6 ] | S3F score and an EVE ensemble trained on the corresponding wild-type-chain MSA | Five EVE seeds; 400,000 optimization steps, batch size 256, learning rate ; 20,000 Monte Carlo samples per seed at scoring time | Equal average of query-wise standardized S3F and EVE scores, where the EVE score is the negative mean evolution index across seeds |
| Category | Model | Baseline | Adaptive | |||
| T1 | T2 | T3 | T4 | |||
| Statistic | Total | 123 | 74 | 67 | 29 | 123 |
| LLM | GPT-6 Astra | 7 | 0 | 0 | 0 | 5 |
| Claude Opus 5 | 12 | 6 | 6 | 1 | 9 | |
| Gemini 3.1 Pro | 0 | 0 | 0 | 0 | 0 | |
| Kimi K3 | 2 | 3 | 2 | 1 | — | |