A Riemannian Geometry for Low-rank Adaptation
Organizations: NTT, Inc.
Abstract
Low-rank adaptation (LoRA) is widely used as a parameter-efficient fine-tuning technique for pre-trained deep neural networks, which approximates the weight update via full fine-tuning by a low-rank matrix . This parameterization leads to the equivalence relation for any invertible matrix because and thus both pairs yield the same loss value. This relation induces a quotient manifold where matrices for all are identified, eliminating redundant directions along which the loss value remains unchanged. To respect the geometry of this manifold, the original search space is endowed with a Riemannian metric that is invariant under the equivalence relation. Such a metric induces preconditioning at each gradient step and ensures that each weight update via LoRA changes the loss value, leading to efficient optimization. In this paper, we propose a new Riemannian metric that is specifically tailored to LoRA to close the gap to full fine-tuning at the weight level. We theoretically show that LoRA with our preconditioning induced by this metric satisfies the following two properties at each iteration: (i) The weight update follows the direction closest to the gradient of full fine-tuning within the subspace of first-order weight changes allowed by the LoRA parameterization. (ii) The updated weight matrix is closer in Frobenius norm to that of full fine-tuning than the updated weight matrices of LoRA with conventional preconditioning and without preconditioning. These theoretical insights suggest that our preconditioning makes LoRA better approximate full fine-tuning, thereby leading to more efficient optimization. Experiments show the effectiveness and efficiency of our preconditioning for LoRA on fine-tuning tasks with language and vision domains.
Figures & tables
| Method | BLEU | NIST | METEOR | ROUGE-L | CIDEr |
|---|---|---|---|---|---|
| SGD | 0.6698 | 8.5505 | 0.4479 | 0.6908 | 2.2996 |
| w/ 4 | 0.6852 | 8.6371 | 0.4641 | 0.7097 | 2.4367 |
| w/ 5 (Ours) | 0.6944 | 8.7449 | 0.4632 | 0.7127 | 2.5082 |
| AdamW | 0.7020 | 8.8433 | 0.4676 | 0.7193 | 2.5500 |
| w/ 4 | 0.7003 | 8.8170 | 0.4671 | 0.7181 | 2.5312 |
| w/ 5 (Ours) | 0.7047 | 8.8527 | 0.4697 | 0.7209 | 2.5454 |
| Method | Cars | DTD | EuroSAT | GTRSB | RESISC45 | SUN397 | SVHN | Ave. |
|---|---|---|---|---|---|---|---|---|
| AdamW | 71.82 | 70.53 | 98.35 | 97.84 | 94.48 | 72.04 | 96.79 | 85.98 |
| w/ 4 | 72.14 | 71.28 | 98.19 | 98.32 | 94.49 | 72.64 | 96.83 | 86.27 |
| w/ 5 (Ours) | 72.03 | 71.86 | 98.52 | 98.54 | 94.70 | 72.57 | 96.92 | 86.45 |
| Method | BoolQ | PIQA | SIQA | HellaS | WinoG | ARC-e | ARC-c | OBQA | Ave. |
|---|---|---|---|---|---|---|---|---|---|
| AdamW | 73.33 | 83.41 | 79.38 | 92.36 | 83.74 | 85.77 | 72.87 | 82.60 | 81.68 |
| w/ 4 | 74.31 | 84.17 | 77.99 | 93.12 | 83.43 | 86.28 | 74.23 | 82.20 | 81.97 |
| w/ 5 (Ours) | 73.55 | 83.62 | 79.02 | 92.95 | 85.48 | 86.41 | 74.32 | 83.40 | 82.34 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| # of Samples | Cars | DTD | EuroSAT | GTRSB | RESISC45 | SUN397 | SVHN |
|---|---|---|---|---|---|---|---|
| Train | 7,329 | 1,692 | 16,200 | 23,976 | 18,900 | 17,865 | 65,931 |
| Validation | 815 | 188 | 2,700 | 2,664 | 6,300 | 1,985 | 7,326 |
| Test | 8,041 | 1,880 | 8,100 | 12,630 | 6,300 | 19,850 | 26,032 |
| # of Samples | BoolQ | PIQA | SIQA | HellaS | WinoG | ARC-e | ARC-c | OBQA |
|---|---|---|---|---|---|---|---|---|
| Test | 1172 | 2376 | 3270 | 10042 | 500 | 1838 | 1954 | 1267 |
| Method | Learning rate |
|---|---|
| SGD | 0.08 |
| w/ 4 | 0.06 |
| w/ 5 (Ours) | 0.02 |
| AdamW | 0.0028 |
| w/ 4 | 0.0028 |
| w/ 5 (Ours) | 0.0030 |
| # of Samples | Cars | DTD | EuroSAT | GTRSB | RESISC45 | SUN397 | SVHN |
|---|---|---|---|---|---|---|---|
| AdamW | 0.0008 | 0.0006 | 0.001 | 0.0007 | 0.001 | 0.0005 | 0.001 |
| w/ 4 | 0.0009 | 0.0008 | 0.0009 | 0.0009 | 0.001 | 0.0007 | 0.001 |
| w/ 5 (Ours) | 0.0009 | 0.001 | 0.001 | 0.001 | 0.001 | 0.0007 | 0.001 |
| Method | Learning rate |
|---|---|
| AdamW | 0.0001 |
| w/ 4 | 0.0001 |
| w/ 5 (Ours) | 0.0001 |
| Learning rate | 0.02 | 0.06 | 0.08 |
|---|---|---|---|
| 10647.35 | 16542.16 | 20334.76 | |
| 10646.05 | 16534.73 | 20324.08 | |
| 10646.04 | 16534.69 | 20324.02 |