Analysis of Quantized and Efficiently Adapted Protein Language Models
Organizations: Leslie Dan Faculty of Pharmacy, University of Toronto, Toronto, ON, Canada. · Department of Electrical and Computer Engineering, University of Toronto, Toronto, Ontario, Canada. · Department of Computer Science, University of Toronto, Toronto, Ontario, Canada. · Department of Mechanical and Industrial Engineering, University of Toronto, Toronto, Ontario, Canada. · Department of Computing and Software, McMaster University, Hamilton, ON, Canada.
Abstract
Background: Protein language models (PLMs) are increasingly used for sequence generation and property prediction, but their size makes fine-tuning and deployment expensive. The effects of quantization and parameter efficient fine-tuning on performance, representations and generation remain insufficiently characterized. Results: We evaluated 4-bit quantization and low-rank adapter fine-tuning (QLoRA) across ESM-2, ESMC, ProtBERT, ProtT5, Ankh, Ankh3 and Profluent-E1. Across protein prediction tasks, many model-task pairs retained more than 90% of full fine-tuning performance. Peak GPU memory savings approached 90% for the largest models, although performance and efficiency varied by model, dataset and training configuration. QLoRA often preserved early-layer representations while inducing task-specific adaptations in middle and late layers, resembling full fine-tuning with smaller representational changes. Training speed and power effects were more varied. For unconditional generation with ProLLaMA, ProtGPT2, ProGen2, ProteinGLM and ESM3, 4-bit quantization largely preserved predicted structural and sequence-level properties, but token-level analysis revealed model-dependent shifts in autoregressive output distributions. Conclusion: QLoRA and 4-bit quantization reduce PLM computational requirements, particularly GPU memory usage. Our results support QLoRA as a first-pass strategy for memory limited adaptation, reserving full fine-tuning for challenging tasks, unstable architectures or low validation recovery. For generative PLMs, sequence-level and structural metrics should be complemented with distributional analysis, since downstream predictions alone may miss quantization-induced shifts. These approaches can broaden access to large-scale protein modelling while requiring model- and task-specific validation.
Figures & tables
| Use case | Recommended starting point | Validation check | Signals for caution |
|---|---|---|---|
| Supervised protein property prediction | Start with QLoRA (especially when GPU memory is limiting) | Compare QLoRA against full fine-tuning on representative validation subset (ratio and uncertainty) | QLoRA recovery below practical threshold, typically 90-95%*, large uncertainty introduced or latency degrades |
| Low signal or high variance tasks | Use QLoRA as initial screen, not as a default | Run multiple seeds and compare absolute performance and variance, not just QLoRA/full ratio | Baseline full fine-tuning performance is weak/highly variable |
| Hyperparameter sensitive models | Perform a small learning rate and batch size sweep before comparison | Check convergence stability across seeds | Results are unstable across random seeds, learning rates or batch sizes |
| Unconditional generation | Use 4-bit quantization for exploratory memory efficient generation | Compare sequence-level and structural metrics between full and quantized models | Token-level or entropy stratified KL shifts substantially; consider residue frequencies, length distribution, diversity metrics |
| Downstream experimental design | Consider quantization as a bonus, not as the only decision factor | Confirm quantized outputs preserve relevant design objectives | Generated candidates will be experimentally validated |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Workflow | Checklist item | Purpose/what to report |
|---|---|---|
| Supervised fine-tuning | Report pretrained model checkpoint and parameter scale | Enables comparison across models, parameter sizes |
| Report quantization configuration | Bit width, quantization method, target modules, compute dtype | |
| Report LoRA configuration | Rank, alpha, dropout, target modules | |
| Report fine-tuning hyperparameters | Learning rate, optimizer, batch size, epoch count/early stopping criterion, random seeds | |
| Compare QLoRA with full fine-tuning when feasible | Use same train/validation/test split | |
| Report absolute performances | Avoid relying only on QLoRA/full performance ratio |
| Fluorescence | Stability | Secondary Structure | ||||||||||
| Full | QLoRA | Full | QLoRA | Full | QLoRA | |||||||
| Model | BS | LR | BS | LR | BS | LR | BS | LR | BS | LR | BS | LR |
| ESM-2 8M | 32 | 5E-05 | 32 | 5E-05 | 32 | 5E-05 | 32 | 5E-05 | 32 | 5E-05 | 32 | 5E-05 |
| ESM-2 35M | 32 | 5E-05 | 32 | 5E-05 | 32 | 5E-05 | 32 | 5E-05 | 32 | 5E-05 | 32 | 5E-05 |
| ESM-2 150M | 32 | 5E-05 | 32 | 5E-05 | 32 | 5E-05 | 32 | 5E-05 | 32 | 5E-06 | 32 | 5E-05 |
| ESM-2 650M | 16 | 5E-05 | 16 | 5E-05 | 16 | 5E-05 | 16 | 5E-05 | 16 | 5E-06 | 16 | 5E-05 |