Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT), yet a performance gap can remain relative to full fine-tuning (FFT). Many LoRA variants improve the initialization or optimization of low-rank factors. At each training step, however, their first-order weight-space directions are constrained by the current parameterization. We characterize the corresponding LoRA-accessible gradient space and show that it coincides with the tangent space induced by the current LoRA parameterization. This characterization yields an orthogonal decomposition of the full weight gradient at the current model parameters. We term the component orthogonal to this space the normal gradient. Based on this decomposition, we propose GDLoRA (Gradient-Decomposed Low-Rank Adaptation). GDLoRA reconstructs the full weight gradient from forward activations and backward signals, extracts its normal component, and directly updates the base weights with this component, while retaining standard AdamW optimization for the LoRA factors. GDLoRA incorporates complementary normal gradients without increasing standard LoRA's optimizer-state memory budget under matched adapter and optimizer configurations. Experiments on natural language understanding, mathematical reasoning, commonsense reasoning, and image classification show that GDLoRA consistently improves over LoRA and narrows the performance gap to FFT. The code is available at https://anonymous.4open.science/r/GDLoRA.
Figures & tables
Figure 1: Overview of GDLoRA. (a) GDLoRA reconstructs the full weight gradient at the current model parameters from forward activations and backward signals, and obtains G as defined in Section 3.1 . (b) The gradient is decomposed with respect to the current parameterization-induced LoRA tangent space, and its orthogonal component G⊥ is the normal gradient . (c) The adapters follow standard AdamW optimization, while G⊥ updates Wbase without additional base-weight momentum or variance buffers. The two branches are complementary at first order at the current iterate.
Model
Method
SST-2
MRPC
CoLA
QNLI
RTE
STS-B
Avg.
FFT
94.42 ± 0.18
89.24 ± 0.61
64.53 ± 1.08
92.81 ± 0.22
78.96 ± 0.84
91.24 ± 0.19
85.20
LoRA
93.04 ± 0.27
87.12 ± 0.74
62.89 ± 1.31
92.31 ± 0.33
75.81 ± 1.12
89.49 ± 0.24
83.44
PiSSA
93.15 ± 0.41
87.09 ± 0.58
60.98 ± 0.73
92.02 ± 0.68
77.98 ± 0.97
90.32 ± 0.17
83.59
DoRA
93.03 ± 0.34
87.25 ± 0.81
62.77 ± 1.46
92.38 ± 0.19
76.53 ± 1.28
89.54 ± 0.31
83.58
rsLoRA
93.06 ± 0.16
87.73 ± 0.49
61.62 ± 1.12
92.77 ± 0.37
77.62 ± 0.88
89.87 ± 0.22
83.78
Fira
92.51 ± 0.29
86.92 ± 0.67
60.64 ± 1.24
92.62 ± 0.26
77.34 ± 1.05
89.99 ± 0.18
83.34
Table 1: Results on six GLUE tasks with RoBERTa-Base and RoBERTa-Large. All LoRA-based methods use rank 8. For each model, the best and second-best scores among all methods, including FFT, are shown in bold and underlined , respectively.
Method
AddSub
MultiArith
SingleEq
SVAMP
GSM8K
AQuA
Avg.
FFT
86.78 ± 1.37
90.21 ± 1.83
93.21 ± 0.84
70.87 ± 1.26
56.72 ± 0.75
25.96 ± 1.54
70.63
LoRA
81.52 ± 1.54
87.28 ± 1.11
91.54 ± 0.69
66.93 ± 1.59
55.34 ± 0.66
22.44 ± 0.60
67.51
PiSSA
82.78 ± 1.58
86.89 ± 1.44
91.27 ± 0.16
66.40 ± 0.79
54.11 ± 0.57
25.33 ± 0.68
67.80
DoRA
80.51 ± 1.98
88.78 ± 0.86
91.08 ± 1.40
67.00 ± 1.61
55.45 ± 0.80
23.75 ± 1.27
67.76
rsLoRA
81.34 ± 1.23
86.15 ± 1.09
90.23 ± 0.78
67.24 ± 1.24
54.21 ± 0.68
24.72 ± 0.95
67.32
Fira
80.64 ± 1.47
86.69 ± 1.16
90.39 ± 0.96
67.64 ± 1.12
54.18 ± 0.77
24.57 ± 1.25
67.35
Table 2: Results on six mathematical reasoning tasks with Llama-3-8B. All LoRA-based methods use rank 8. The best and second-best scores among all methods, including FFT, are shown in bold and underlined , respectively.
Method
OBQA
ARC-c
WinoGrande
PIQA
SIQA
ARC-e
BoolQ
HellaSwag
Avg.
FFT
86.77 ± 1.46
84.93 ± 1.66
88.05 ± 1.25
89.11 ± 1.07
80.65 ± 1.43
93.06 ± 1.28
70.12 ± 0.24
92.89 ± 0.84
85.70
LoRA
81.62 ± 0.82
82.74 ± 0.89
85.08 ± 0.88
87.54 ± 0.65
77.79 ± 1.14
92.67 ± 1.01
68.84 ± 0.25
90.95 ± 0.65
83.40
PiSSA
81.24 ± 0.57
81.06 ± 1.06
81.69 ± 0.73
86.24 ± 0.48
75.74 ± 1.28
91.12 ± 0.84
67.43 ± 0.31
89.13 ± 0.79
81.71
DoRA
82.53 ± 0.94
81.95 ± 0.68
83.58 ± 1.02
86.82 ± 0.72
76.93 ± 0.96
91.16 ± 1.12
68.45 ± 0.22
83.56 ± 0.58
81.87
rsLoRA
82.86 ± 0.71
81.83 ± 1.14
84.50 ± 0.79
87.63 ± 0.54
77.12 ± 1.21
92.13 ± 0.93
68.09 ± 0.37
92.23 ± 0.83
83.30
Fira
83.12 ± 1.03
81.66 ± 0.74
84.16 ± 0.91
87.92 ± 0.60
75.84 ± 1.35
91.26 ± 0.76
68.71 ± 0.29
90.42 ± 0.52
82.89
Table 3: Results on eight commonsense reasoning tasks with Gemma-7B. All LoRA-based methods use rank 8. The best and second-best scores among all methods, including FFT, are shown in bold and underlined , respectively.
Model
Method
OxfordPets
StanfordCars
CIFAR-10
DTD
EuroSAT
FGVC
RESISC45
CIFAR-100
Avg.
FFT
96.13 ± 0.21
76.81 ± 0.94
98.96 ± 0.04
77.75 ± 0.88
98.92 ± 0.19
54.64 ± 0.97
96.12 ± 0.34
92.11 ± 0.18
86.43
LoRA
95.92 ± 0.37
74.93 ± 0.58
97.26 ± 0.06
75.03 ± 0.54
98.57 ± 0.12
49.17 ± 0.71
95.43 ± 0.29
91.44 ± 0.22
84.72
PiSSA
95.46 ± 0.18
74.24 ± 0.83
97.92 ± 0.03
74.82 ± 0.97
98.58 ± 0.05
51.07 ± 0.52
95.12 ± 0.17
90.14 ± 0.13
84.67
DoRA
95.62 ± 0.29
74.69 ± 0.41
97.64 ± 0.07
75.77 ± 0.63
98.76 ± 0.11
49.81 ± 0.86
95.59 ± 0.23
91.32 ± 0.20
84.90
rsLoRA
95.51 ± 0.33
74.12 ± 0.69
97.66 ± 0.05
75.87 ± 0.45
98.49 ± 0.08
50.78 ± 0.61
95.48 ± 0.12
91.08 ± 0.16
84.87
Fira
95.06 ± 0.16
75.28 ± 1.07
98.54 ± 0.02
76.19 ± 0.76
98.87 ± 0.06
51.22 ± 0.74
95.08 ± 0.31
90.95 ± 0.24
85.15
Table 4: Image classification results with ViT-Base and ViT-Large. For each model, the best and second-best scores among all methods are shown in bold and underlined , respectively.
Figure 2: Ablation studies and analysis of GDLoRA. (a) Comparison of gradient components on six GLUE tasks with RoBERTa-Base and RoBERTa-Large. (b) Sensitivity to the base-update learning rate when fine-tuning Llama-3-8B on Math10K. Shaded bands indicate ±1 standard deviation. (c) Average mathematical reasoning accuracy and memory usage with Llama-3-8B on Math10K.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Gradient subspace drift and projection refresh costs. (a) Pairwise similarity between the leading rank-8 gradient subspaces of model.layers.15.self_attn.v_proj over 50 optimizer steps. (b) Accuracy and training time of native Fira with projection refresh intervals of 200, 150, 100, and 50 steps on Math10K with Llama-3-8B. GDLoRA is shown as a fixed reference; its repeated values correspond to the same configuration.
Figure 4: Gradient subspace variation across network depths and attention projections. Each heatmap shows pairwise similarity between leading rank-8 gradient subspaces over 50 optimizer steps during Llama-3-8B fine-tuning on Math10K. From left to right, the top row shows the layer-0 key, layer-0 value, and layer-15 query projections; the bottom row shows the layer-15 value, layer-31 query, and layer-31 value projections. Layer numbers follow the model’s module indices. The layer-15 value projection is also shown in Figure 3 (a). All panels use the same color scale.
Operation
Fira
GDLoRA
Full-weight gradient
2Ndodi
2Ndodi
Gradient projection and reconstruction
6dodir
8dodir
Basis construction
CSVD/T
CQR
Adapter forward and backward
0
6Nr(do+di)
Appendix
Table 5: Leading FLOPs per layer and optimization step. Fira’s SVD cost is amortized over its refresh interval T . GDLoRA recomputes both bases every step. Projection counts assume rank- r bases and efficient matrix-product ordering.
Method
SST-2
MRPC
CoLA
QNLI
RTE
STS-B
FFT
10.16
1.64
5.78
19.56
2.21
3.64
LoRA
6.92
0.58
3.54
13.24
1.06
2.50
PiSSA
13.06
1.08
4.92
16.78
1.47
2.98
DoRA
9.84
0.65
4.03
13.78
1.19
2.52
rsLoRA
9.98
0.57
3.51
14.23
1.03
2.50
Fira
16.18
1.53
6.31
19.65
2.34
4.35
Appendix
Table 6: Training time (hours) on six GLUE tasks with RoBERTa-Base.
Model
Hyperparameter
SST-2
MRPC
CoLA
QNLI
RTE
STS-B
Optimizer
AdamW
Warmup ratio
0.06
Learning-rate schedule
Linear
RoBERTa-Base
# GPUs
1
Epochs
60
30
80
25
160
80
Learning rate (LoRA)
1e-4
Appendix
Table 7: Hyperparameters used for GDLoRA in the natural language understanding experiments on the GLUE benchmark.
Method
AddSub
MultiArith
SingleEq
SVAMP
GSM8K
AQuA
Avg.
LoRA ( r=8 )
81.52
87.28
91.54
66.93
55.34
22.44
67.51
AdaLoRA
80.57
84.22
89.47
64.80
52.12
25.44
66.10
IA 3
86.84
83.50
86.42
64.50
50.64
26.38
66.38
FourierFT
85.06
88.67
88.98
64.30
45.40
27.95
66.73
MiSS
84.67
86.93
88.71
64.62
55.21
25.35
67.58
GDLoRA
84.81
90.69
92.52
69.61
56.03
26.38
70.01
Appendix
Table 8: Additional comparison on six mathematical reasoning tasks with Llama-3-8B. Scores are accuracy (%); Avg. is the arithmetic mean across the six tasks. The best mean score in each column is shown in bold .
Method
AddSub
MultiArith
SingleEq
SVAMP
GSM8K
AQuA
Avg.
LoRA ( r=8 )
81.52
87.28
91.54
66.93
55.34
22.44
67.51
LoRA ( r=16 )
83.56
91.34
92.43
68.81
56.97
24.89
69.67
LoRA ( r=32 )
85.14
91.87
93.45
69.72
57.66
26.95
70.80
GDLoRA ( r=8 )
84.81
90.69
92.52
69.61
56.03
26.38
70.01
GDLoRA ( r=16 )
88.15
92.16
93.78
70.85
57.94
28.89
71.96
GDLoRA ( r=32 )
90.36
94.48
94.23
71.35
58.47
27.94
72.81
Appendix
Table 9: LoRA and GDLoRA across adapter ranks on six mathematical reasoning tasks with Llama-3-8B. Scores are accuracy (%); Avg. is the arithmetic mean across tasks. GDLoRA rows are shaded blue, and the best score in each column is shown in bold .
Hyperparameter
Setting
Optimizer
AdamW
Learning rate (LoRA)
1e-4
Normal-gradient learning rate
0.5
Weight decay
0.0
Warmup ratio
0.06
Learning-rate schedule
Linear
Appendix
Table 10: Training and evaluation settings for mathematical reasoning with Llama-3-8B fine-tuned on Math10K.
Hyperparameter
Setting
Optimizer
AdamW
Learning rate (LoRA)
3e-5
Normal-gradient learning rate
0.5
Weight decay
0.0
Warmup ratio
0.06
Learning-rate schedule
Linear
Appendix
Table 11: Training and evaluation settings for commonsense reasoning with Gemma-7B fine-tuned on Commonsense170K.
Model
Hyperparameter
Oxford Pets
Stanford Cars
CIFAR-10
DTD
EuroSAT
FGVC
RESISC45
CIFAR-100
Optimizer
AdamW
Weight decay
0.01
Learning-rate schedule
Linear
Epochs
20
ViT-Base
# GPUs
1
Learning rate (head)
5e-3
5e-2
5e-2
5e-2
5e-2
5e-2
5e-2
1e-2
Appendix
Table 12: Hyperparameters used for GDLoRA in the image classification experiments with ViT-Base and ViT-Large. The normal-gradient learning rate controls the direct updates to the base weights.
Low-rank adaptation (LoRA) fine-tunes large pretrained models at a fraction of the cost of full fine-tuning, but its performance depends strongly on how the adapters are initialized. Recent schemes initialize the adapters from the downstream loss gradient: some project the raw gradient onto its top directions, while others first whiten it with an estimate of the loss curvature. We show that these seemingly distinct methods are points on a single continuum: a two-parameter family of preconditioned gradient initializations, which we call Unified LoRA (ULoRA), governed by a spectral whitening exponent and an Adam-like diagonal exponent. Sweeping this family under a full learning-rate search, we find that no single fixed preconditioning strength dominates: the best operating point is task-dependent and frequently lies strictly inside the family, away from the published endpoints. Treated as an upper bound of this family, a tuned ULoRA configuration matches or exceeds full fine-tuning on all five GLUE tasks with RoBERTa-base and is competitive with the strongest baselines on GSM8K with LLaMA-2-7B. Our deployable, search-free variant, ULoRA-Auto, selects per-layer exponents from measured spectral statistics, approaches this upper bound at no additional search cost, and ranks at or near the top among deployable LoRA methods. Our results show that a principled design space for LoRA initialization and curvature preconditioning should be treated as a tunable dimension rather than a fixed design decision.
Low-Rank Adaptation (LoRA) enables efficient adaptation of large pre-trained models to downstream tasks by parameterizing weight updates with low-rank matrices. In this paper, we investigate the limitations of the LoRA parameterization from a geometric perspective. Specifically, we show that when a full fine-tuning gradient is backpropagated to the low-rank matrices, it undergoes anisotropic scaling driven by their singular values. We argue that this phenomenon is undesirable because it distorts the full fine-tuning gradient by skewing it toward dominant singular directions while suppressing others. Our analyses demonstrate that anisotropic gradient scaling reduces the effective rank of the low-rank matrices' gradients and results in suboptimal alignment between the full fine-tuning gradient and its low-rank approximation in LoRA, thereby exacerbating the gap to full fine-tuning. To address these limitations, we propose a new low-rank parameterization, SDS-LoRA, which structurally decouples singular values from the backward pass. Our method ensures that the full fine-tuning gradient backpropagates only through the orthonormal bases of the low-rank matrices' subspaces, independent of their scales. Convergence analysis demonstrates that while LoRA's convergence rate degrades with the condition number of the low-rank matrices, SDS-LoRA remains independent of it. Experimental results across natural language and vision benchmarks show that SDS-LoRA improves loss convergence and reduces the gap to full fine-tuning, significantly enhancing adaptation performance.
Junghun Oh, Sungyong Baik, Kyoung Mu Lee
Dept. of ECE, ASRI Seoul National University · Dept. of Artificial Intelligence Dept. of Data Science Hanyang University
Low-Rank Adaptation (LoRA) is a prominent fine-tuning method for large models, achieving competitive performance with reduced memory overhead. However, a persistent performance gap remains between LoRA and full fine-tuning. Recent studies have sought to narrow this gap by employing one-step gradient approximations of pretrained weights to align LoRA updates with the principal directions or intrinsic dimensionalities of full fine-tuning updates. Nevertheless, these approaches fail to capture the full dynamics of the gradients. In this paper, we propose LoRA-GA2, an effective fine-tuning algorithm that fully leverages multi-step gradient information. Specifically, we introduce a lightweight probe for multi-step gradients of pretrained weights that incurs no additional GPU memory cost and only marginal time overhead. We further employ a spectrum-aware, importance-based rank allocation and optimal initialization derived from multi-step gradients. Extensive experimental results demonstrate that LoRA-GA2 consistently outperforms existing LoRA variants while preserving the efficiency advantages of vanilla LoRA. For instance, LoRA-GA2 surpasses the leading baseline by an average of 0.66 points on the GLUE benchmark, and outperforms the strongest baseline by 1.03 points on GSM8K and 0.87 points on HumanEval, respectively.
Haonan He, Xinyue Fan
University of Science and Technology of China · Independent Researcher