cs.CLMar 11, 2025

PUMA: Learning a Mutation-Aware Vocabulary of Protein Units

Authors: Burak Suyunu, Özdeniz Dolu, Ibukunoluwa Abigail Olaosebikan, Hacer Karatas Bristow, Arzucan Özgür

Organizations: Department of Computer Engineering, Boaziçi University, 34342, Bebek, Istanbul, Turkey · C. Eugene Bennett Department of Chemistry, West Virginia University, Morgantown, West Virginia 26505, United States · Department of Pharmaceutical Chemistry, Istanbul Medipol University, Faculty of Pharmacy, 34815 Beykoz, Istanbul, Turkey · Istanbul Medipol University, Research Institute for Health Sciences and Technologies (SABITA), 34810 Beykoz, Istanbul, Turkey

Abstract

Modeling protein sequences as a language has made language models a powerful tool in computational biology, yet the language itself remains poorly understood. A key step toward understanding it is identifying its constituent units. In natural languages, morphemes can occur in multiple forms; similarly, in proteins, mutations can give rise to variations of a unit that persist through evolution, forming families of related units. We introduce PUMA (Protein Units via Mutation-Aware Merging), an algorithm that learns protein units from sequence and explores their mutational variants using substitution matrices, forming a genealogy of unit families. Our results show that mutations remaining within a PUMA family are more often benign than the substitution matrix alone predicts, and that PUMA genealogy improves molecular function representations compared to treating units independently. A case study of a unit family demonstrates relatedness beyond homology. PUMA achieves competitive performance on downstream tasks when used as a protein language model tokenizer. Moreover, collapsing units into families results in a smaller embedding table and faster training. Together, these results support PUMA as a biologically grounded protein vocabulary that organizes protein units into plausible families of mutational variants. The source code is available at https://github.com/boun-tabi-lifelu/PUMA.

Figures & tables

Explore similar work

CardsList
  1. ProtSent: Protein Sentence Transformers

    May 7, 2026Dan Ofer, Oriel Perets, Michal Linial +1Language ModelingTransformer Architectures

  2. Better Protein Function Prediction by Modeling Survivorship Bias

    May 7, 2026Zhongmou Chao, Poompol Buathong, Ekaterina Selivanovitch +2Survival Predictions

  3. ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing

    Sep 7, 2026Mingrui Li, Sixian Shen, Minzhang Li +4Language Modeling