cs.LGSep 30, 2026

Analysis of Quantized and Efficiently Adapted Protein Language Models

Authors: Ilan Yaniv Zeisler, Sebastian Clancy, Pouriya Bayat, Saaim Raad, Ivan Kraskov, Matthew Xie, Vivian White, Spencer Perkins, +3 more

Organizations: Leslie Dan Faculty of Pharmacy, University of Toronto, Toronto, ON, Canada. · Department of Electrical and Computer Engineering, University of Toronto, Toronto, Ontario, Canada. · Department of Computer Science, University of Toronto, Toronto, Ontario, Canada. · Department of Mechanical and Industrial Engineering, University of Toronto, Toronto, Ontario, Canada. · Department of Computing and Software, McMaster University, Hamilton, ON, Canada.

Abstract

Background: Protein language models (PLMs) are increasingly used for sequence generation and property prediction, but their size makes fine-tuning and deployment expensive. The effects of quantization and parameter efficient fine-tuning on performance, representations and generation remain insufficiently characterized. Results: We evaluated 4-bit quantization and low-rank adapter fine-tuning (QLoRA) across ESM-2, ESMC, ProtBERT, ProtT5, Ankh, Ankh3 and Profluent-E1. Across protein prediction tasks, many model-task pairs retained more than 90% of full fine-tuning performance. Peak GPU memory savings approached 90% for the largest models, although performance and efficiency varied by model, dataset and training configuration. QLoRA often preserved early-layer representations while inducing task-specific adaptations in middle and late layers, resembling full fine-tuning with smaller representational changes. Training speed and power effects were more varied. For unconditional generation with ProLLaMA, ProtGPT2, ProGen2, ProteinGLM and ESM3, 4-bit quantization largely preserved predicted structural and sequence-level properties, but token-level analysis revealed model-dependent shifts in autoregressive output distributions. Conclusion: QLoRA and 4-bit quantization reduce PLM computational requirements, particularly GPU memory usage. Our results support QLoRA as a first-pass strategy for memory limited adaptation, reserving full fine-tuning for challenging tasks, unstable architectures or low validation recovery. For generative PLMs, sequence-level and structural metrics should be complemented with distributional analysis, since downstream predictions alone may miss quantization-induced shifts. These approaches can broaden access to large-scale protein modelling while requiring model- and task-specific validation.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression

    May 30, 2026Wenya Yu, Chao Zhang, Li Wang +2Large Language Model QuantizationPost-Training Quantization

  2. Loopy: Low-Bit Quantization Framework for Looped Language Models

    Oct 4, 2026Zeyu LI, Yipu ZHANG, Jintao Chen +2Large Language Model Quantization

  3. On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs

    Sep 22, 2025Rongguang Ye, Ming Tang, Edith C. H. NgaiLarge Language Model QuantizationTraining-Side Stage-Aware Low-Rank Adaptation