Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models
Organizations: School of Computer Science and Statistics, Trinity College Dublin, Dublin, Ireland and Lero the Research Ireland Centre for Software, Ireland · School of Computer Science, University College Dublin, Dublin, Ireland
Abstract
Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is especially severe when fine-tuned on small datasets due to the inherent tendency toward miscalibration. In this work, we introduce Bayesian-LoRA, which reformulates the deterministic LoRA update as a probabilistic low-rank representation inspired by Sparse Gaussian Processes. We identify a structural isomorphism between LoRA's factorization and Kronecker-factored SGP posteriors, and show that LoRA emerges as a limiting case when posterior uncertainty collapses. We conduct extensive experiments on various LLM architectures across commonsense reasoning benchmarks. With only approximately 0.42M additional parameters and training cost relative to standard LoRA, Bayesian-LoRA significantly improves calibration across models up to 30B, achieving up to 84% ECE reduction and 76% NLL reduction while maintaining competitive accuracy for both in-distribution and out-of-distribution (OoD) evaluations.
Figures & tables
| All tokens | Top 5% entropy tokens | |||||||||||
| Validation | Test | Validation | Test | |||||||||
| Method | NLL | Brier | ECE | NLL | Brier | ECE | NLL | Brier | ECE | NLL | Brier | ECE |
| LoRA (MAP) | 1.76 | 0.52 | 1.68 | 1.75 | 0.52 | 1.48 | 5.16 | 0.97 | 0.82 | 5.16 | 0.97 | 1.60 |
| LoRA + Temp | 1.78 | 0.51 | 1.62 | 1.77 | 0.50 | 1.54 | 5.22 | 0.92 | 0.81 | 5.25 | 0.92 | 1.49 |
| Dropout ( ) | 1.77 | 0.50 | 1.54 | 1.70 | 0.50 | 1.59 | 5.21 | 0.94 | 0.79 | 5.19 | 0.95 | 1.51 |
| Bayesian-LoRA (N = 1) | 1.73 | 0.51 | 1.46 | 1.71 | 0.51 | 1.51 | 5.14 | 0.95 | 0.80 | 5.14 | 0.92 | 1.31 |
| Model (Zero-shot) | Method | CoT-NLL | CoT-ECE | Answer Acc. |
| Qwen2.5-14B-Instruct | Baseline FT | 2.165 | 12.2 | 49.8 |
| Dropout ( Gal and Ghahramani, 2016 ) | 2.103 | 11.9 | 50.0 | |
| Temp ( Guo et al., 2017 ) | 1.96 | 10.7 | 49.9 | |
| LA (post-hoc) ( Yang et al., 2024 ) | 0.81 | 7.12 | 49.8 | |
| BLoB ( ) ( Wang et al., 2024a ) | 1.21 | 8.41 | 47.2 | |
| Bayesian-LoRA ( ) | 0.513 | 5.81 | 51.1 |
| Method | Trainable Params | Train time ( MAP) | Peak mem. ( MAP) | Inference ( MAP) | Samples (Val) |
| MAP (LoRA) ( Hu et al., 2022 ) | 4.48M | 1.00 | 1.00 | 1.00 | 1 |
| Dropout ( Gal and Ghahramani, 2016 ) | 4.48M | 4 | |||
| Ckpt Ens (3) ( Huang et al., 2017 ) | MAP | 3 | |||
| Deep Ens (3) ( Lakshminarayanan et al., 2017 ) | MAP | 3 | |||
| BBB ( Blundell et al., 2015 ) | MAP | 4 | |||
| BLoB ( ) ( Wang et al., 2024a ) | 1.5 MAP | 4 |
| Flow depth | ACC | ECE | NLL | Efficiency ( MAP) | ||||
| value | vs. | value | vs. | value | vs. | train time | peak mem. | |
| 0 (pure SGP) | 79.0 0.21 | -2.6 | 5.8 0.13 | +0.1 | 0.58 0.08 | +0.09 | 1.19 | 1.002 |
| 1 | 81.6 0.10 | 0.0 | 5.7 0.20 | 0.0 | 0.49 0.02 | 0.00 | 1.23 | 1.003 |
| 2 | 80.8 0.14 | -0.8 | 5.6 0.09 | -0.1 | 0.52 0.06 | +0.03 | 1.30 | 1.008 |
| 4 | 80.9 0.08 | -0.7 | 4.9 0.03 | -0.8 | 0.48 0.13 | -0.01 | 1.38 | 1.010 |
| ID | Smaller Distribution Shift | Larger Distribution Shift | ||||||
| Metrics | Methods | OBQA | ARC-C | ARC-E | CS | Eng | Law | Health |
| ACC | MAP ( Hu et al., 2022 ) | |||||||
| Dropout ( Gal and Ghahramani, 2016 ) | ||||||||
| Ckpt Ens ( Huang et al., 2017 ) | ||||||||
| Temp ( Guo et al., 2017 ) | ||||||||
| BBB ( Blundell et al., 2015 ) | ||||||||
| Method | ACC | ECE | NLL |
| LoRA (MAP) | |||
| Degenerate Bayesian-LoRA | |||
| Bayesian-LoRA (full) |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Task type | Candidates | Primary metric(s) |
| WinoGrande-S/M (WG-S/M) | Cloze coreference (2-way) | 2 | ACC, ECE, NLL |
| ARC-Challenge (ARC-C) | Multiple-choice science QA | 3–5 (mostly 4) | ACC, ECE, NLL |
| ARC-Easy (ARC-E) | Multiple-choice science QA | 3–5 (mostly 4) | ACC, ECE, NLL |
| OpenBookQA (OBQA) | Multiple-choice science QA | 4 | ACC, ECE, NLL |
| BoolQ | Yes/No reading comprehension | 2 | ACC, ECE, NLL |
| Hyperparameter | Value |
| Optimizer | AdamW |
| Learning rate | |
| Betas | |
| Epsilon ( ) | |
| Weight decay | |
| Scheduler | MultiStepLR (milestones = [4, 6], ) |
| Hyperparameter | Value |
| inducing_rows | 9 |
| inducing_cols | 9 |
| whitened_u | True |
| q_inducing | diagonal |
| learn_lambda | True |
| init_lambda | 0.001 |
| Subject | Tasks |
| Computer Science (CS) | college computer science, computer security, |
| high school computer science, machine learning | |
| Engineering (Eng) | electrical engineering |
| Law | international law, jurisprudence, professional law |
| Health | anatomy, clinical knowledge, college medicine, human aging, |
| nutrition, professional medicine, virology |
| Method | ACC | ECE | NLL |
| Qwen3-14B base (4-shot) | 76.7 | 9.07 | 0.2856 |
| Standard LoRA | 80.0 | 11.26 | 0.2982 |
| Bayesian-LoRA (epoch 5) | 80.0 | 5.26 | 0.1144 |
| Metric | Method | ARC-C | ARC-E |
| ACC | BLoB-Mean + TFB | 83.33 0.19 | 91.76 0.48 |
| Bayesian-LoRA | 83.01 | 91.12 | |
| ECE | BLoB-Mean + TFB | 6.48 0.36 | 2.44 0.50 |
| Bayesian-LoRA | 3.23 | 1.83 | |
| NLL | BLoB-Mean + TFB | 0.530 0.04 | 0.230 0.02 |
| Bayesian-LoRA | 0.503 | 0.217 |
| Seed | ACC | ECE | NLL |
| 42 | 94.59 | 0.76 | 0.1347 |
| 123 | 94.57 | 0.88 | 0.1345 |
| 456 | 94.57 | 0.88 | 0.1348 |
| 789 | 94.57 | 0.81 | 0.1349 |
| 2024 | 94.57 | 0.76 | 0.1347 |
| mean std | 94.57 0.01 | 0.82 0.06 | 0.1347 0.0001 |
| Placement | Target modules | ACC | ECE | NLL |
| A (default) | q, k, lm_head | 94.59 | 0.76 | 0.1347 |
| B (LoRA-style) | q, v, lm_head | 94.57 | 0.74 | 0.1347 |
| C (Attention) | q, k, v, o, lm_head | 94.57 | 0.71 | 0.1348 |
| D (MLP) | gate, up, down, lm_head | 94.49 | 0.93 | 0.1348 |
| E (All) | all linear + lm_head | 94.53 | 0.74 | 0.1348 |
| Configuration | ACC | ECE | NLL | ||
| (default) | 9 | 9 | 94.59 | 0.76 | 0.1347 |
| 4 | 9 | 94.57 | 0.80 | 0.1347 | |
| 9 | 4 | 94.57 | 0.72 | 0.1349 | |
| (larger) | 9 | 16 | 94.53 | 0.82 | 0.1349 |
| (larger) | 16 | 9 | 94.53 | 0.83 | 0.1349 |
| Category | Hyperparameter (value) |
| Model & tokenizer | |
| Model name | Qwen/Qwen2.5-14B-Instruct , Qwen/Qwen3-30B-A3B-Instruct-2507 |
| Max input length | 1024 |
| BF16 precision | True |
| FP16 precision | False |
| Load in 8-bit | False |