cs.AISep 24, 2026

PFArena: Benchmarking Language Models for Protein Modification

Authors: Yawen Ouyang, Xinbo Zhang, Ziyuan Ma, Yixin Wu, Wenbin Liao, Feiran Zhang, Wenjie Li, Lihao Wang, +10 more

Organizations: Shanghai Artificial Intelligence Laboratory · Generative Symbolic Intelligence Lab (GenSI), Tsinghua University · Institute for AI Industry Research (AIR), Tsinghua University · School of Pharmaceutical Sciences, Tsinghua University

Abstract

Protein modification requires navigating an immense sequence space, yet wet-lab validation remains low-throughput and costly. Although computational paradigms including protein language models (PLMs), large language models (LLMs), and LLM-based agents have shown promise in protein modification, their relative efficacy across realistic experimental decision-making settings remains unclear. To bridge this gap, we introduce PFArena, a benchmark comprising four controlled task interfaces that cover single-mutant generation and multi-mutant ranking. By providing varying levels of mutation fitness data, PFArena reflects four representative research scenarios characterized by differing degrees of prior experimental context. We assess six PLMs, six LLMs, and five LLM-based agents using complementary metrics to measure both peak and overall protein modification performance. Our evaluation reveals that model performance shifts systematically with the availability of target-specific experimental evidence: PLMs demonstrate proficiency in open-ended single-mutant generation by leveraging protein-specific priors, whereas LLMs and agents perform strongly in multi-mutant ranking, particularly when target-specific fitness data are available. Nevertheless, all model families face fundamental challenges with increasing search-space size and mutation depth. We release our code and benchmark suite to facilitate reproducible research in model-assisted protein modification.

Figures & tables

Explore similar work

CardsList
  1. AgentPLM: Agentic Protein Language Models with Reasoning-Augmented Decoding for Protein Sequence Design

    Jun 1, 2026Sahil Rahman, Maxx Richard RahmanProtein DesignAntibody

  2. AMix-2: Establishing Protein as a Native Modality in Large Language Models

    May 29, 2026Keyue Qiu, Yixin Wu, Lihao Wang +19ProteinLanguage Modeling

  3. Task- and dataset-specific information in protein language models

    Aug 12, 2026Roman Joeres, Ilya Senatorov, Anastasia Kolchina +2Language Modeling