Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT), yet a performance gap can remain relative to full fine-tuning (FFT). Many LoRA variants improve the initialization or optimization of low-rank factors. At each training step, however, their first-order weight-space directions are constrained by the current parameterization. We characterize the corresponding LoRA-accessible gradient space and show that it coincides with the tangent space induced by the current LoRA parameterization. This characterization yields an orthogonal decomposition of the full weight gradient at the current model parameters. We term the component orthogonal to this space the normal gradient. Based on this decomposition, we propose GDLoRA (Gradient-Decomposed Low-Rank Adaptation). GDLoRA reconstructs the full weight gradient from forward activations and backward signals, extracts its normal component, and directly updates the base weights with this component, while retaining standard AdamW optimization for the LoRA factors. GDLoRA incorporates complementary normal gradients without increasing standard LoRA's optimizer-state memory budget under matched adapter and optimizer configurations. Experiments on natural language understanding, mathematical reasoning, commonsense reasoning, and image classification show that GDLoRA consistently improves over LoRA and narrows the performance gap to FFT. The code is available at https://anonymous.4open.science/r/GDLoRA.
Figures & tables
Figure 1: Overview of GDLoRA. (a) GDLoRA reconstructs the full weight gradient at the current model parameters from forward activations and backward signals, and obtains G as defined in Section 3.1 . (b) The gradient is decomposed with respect to the current parameterization-induced LoRA tangent space, and its orthogonal component G⊥ is the normal gradient . (c) The adapters follow standard AdamW optimization, while G⊥ updates Wbase without additional base-weight momentum or variance buffers. The two branches are complementary at first order at the current iterate.
Model
Method
SST-2
MRPC
CoLA
QNLI
RTE
STS-B
Avg.
FFT
94.42 ± 0.18
89.24 ± 0.61
64.53 ± 1.08
92.81 ± 0.22
78.96 ± 0.84
91.24 ± 0.19
85.20
LoRA
93.04 ± 0.27
87.12 ± 0.74
62.89 ± 1.31
92.31 ± 0.33
75.81 ± 1.12
89.49 ± 0.24
83.44
PiSSA
93.15 ± 0.41
87.09 ± 0.58
60.98 ± 0.73
92.02 ± 0.68
77.98 ± 0.97
90.32 ± 0.17
83.59
DoRA
93.03 ± 0.34
87.25 ± 0.81
62.77 ± 1.46
92.38 ± 0.19
76.53 ± 1.28
89.54 ± 0.31
83.58
rsLoRA
93.06 ± 0.16
87.73 ± 0.49
61.62 ± 1.12
92.77 ± 0.37
77.62 ± 0.88
89.87 ± 0.22
83.78
Fira
92.51 ± 0.29
86.92 ± 0.67
60.64 ± 1.24
92.62 ± 0.26
77.34 ± 1.05
89.99 ± 0.18
83.34
Table 1: Results on six GLUE tasks with RoBERTa-Base and RoBERTa-Large. All LoRA-based methods use rank 8. For each model, the best and second-best scores among all methods, including FFT, are shown in bold and underlined , respectively.
Method
AddSub
MultiArith
SingleEq
SVAMP
GSM8K
AQuA
Avg.
FFT
86.78 ± 1.37
90.21 ± 1.83
93.21 ± 0.84
70.87 ± 1.26
56.72 ± 0.75
25.96 ± 1.54
70.63
LoRA
81.52 ± 1.54
87.28 ± 1.11
91.54 ± 0.69
66.93 ± 1.59
55.34 ± 0.66
22.44 ± 0.60
67.51
PiSSA
82.78 ± 1.58
86.89 ± 1.44
91.27 ± 0.16
66.40 ± 0.79
54.11 ± 0.57
25.33 ± 0.68
67.80
DoRA
80.51 ± 1.98
88.78 ± 0.86
91.08 ± 1.40
67.00 ± 1.61
55.45 ± 0.80
23.75 ± 1.27
67.76
rsLoRA
81.34 ± 1.23
86.15 ± 1.09
90.23 ± 0.78
67.24 ± 1.24
54.21 ± 0.68
24.72 ± 0.95
67.32
Fira
80.64 ± 1.47
86.69 ± 1.16
90.39 ± 0.96
67.64 ± 1.12
54.18 ± 0.77
24.57 ± 1.25
67.35
Table 2: Results on six mathematical reasoning tasks with Llama-3-8B. All LoRA-based methods use rank 8. The best and second-best scores among all methods, including FFT, are shown in bold and underlined , respectively.
Method
OBQA
ARC-c
WinoGrande
PIQA
SIQA
ARC-e
BoolQ
HellaSwag
Avg.
FFT
86.77 ± 1.46
84.93 ± 1.66
88.05 ± 1.25
89.11 ± 1.07
80.65 ± 1.43
93.06 ± 1.28
70.12 ± 0.24
92.89 ± 0.84
85.70
LoRA
81.62 ± 0.82
82.74 ± 0.89
85.08 ± 0.88
87.54 ± 0.65
77.79 ± 1.14
92.67 ± 1.01
68.84 ± 0.25
90.95 ± 0.65
83.40
PiSSA
81.24 ± 0.57
81.06 ± 1.06
81.69 ± 0.73
86.24 ± 0.48
75.74 ± 1.28
91.12 ± 0.84
67.43 ± 0.31
89.13 ± 0.79
81.71
DoRA
82.53 ± 0.94
81.95 ± 0.68
83.58 ± 1.02
86.82 ± 0.72
76.93 ± 0.96
91.16 ± 1.12
68.45 ± 0.22
83.56 ± 0.58
81.87
rsLoRA
82.86 ± 0.71
81.83 ± 1.14
84.50 ± 0.79
87.63 ± 0.54
77.12 ± 1.21
92.13 ± 0.93
68.09 ± 0.37
92.23 ± 0.83
83.30
Fira
83.12 ± 1.03
81.66 ± 0.74
84.16 ± 0.91
87.92 ± 0.60
75.84 ± 1.35
91.26 ± 0.76
68.71 ± 0.29
90.42 ± 0.52
82.89
Table 3: Results on eight commonsense reasoning tasks with Gemma-7B. All LoRA-based methods use rank 8. The best and second-best scores among all methods, including FFT, are shown in bold and underlined , respectively.
Model
Method
OxfordPets
StanfordCars
CIFAR-10
DTD
EuroSAT
FGVC
RESISC45
CIFAR-100
Avg.
FFT
96.13 ± 0.21
76.81 ± 0.94
98.96 ± 0.04
77.75 ± 0.88
98.92 ± 0.19
54.64 ± 0.97
96.12 ± 0.34
92.11 ± 0.18
86.43
LoRA
95.92 ± 0.37
74.93 ± 0.58
97.26 ± 0.06
75.03 ± 0.54
98.57 ± 0.12
49.17 ± 0.71
95.43 ± 0.29
91.44 ± 0.22
84.72
PiSSA
95.46 ± 0.18
74.24 ± 0.83
97.92 ± 0.03
74.82 ± 0.97
98.58 ± 0.05
51.07 ± 0.52
95.12 ± 0.17
90.14 ± 0.13
84.67
DoRA
95.62 ± 0.29
74.69 ± 0.41
97.64 ± 0.07
75.77 ± 0.63
98.76 ± 0.11
49.81 ± 0.86
95.59 ± 0.23
91.32 ± 0.20
84.90
rsLoRA
95.51 ± 0.33
74.12 ± 0.69
97.66 ± 0.05
75.87 ± 0.45
98.49 ± 0.08
50.78 ± 0.61
95.48 ± 0.12
91.08 ± 0.16
84.87
Fira
95.06 ± 0.16
75.28 ± 1.07
98.54 ± 0.02
76.19 ± 0.76
98.87 ± 0.06
51.22 ± 0.74
95.08 ± 0.31
90.95 ± 0.24
85.15
Table 4: Image classification results with ViT-Base and ViT-Large. For each model, the best and second-best scores among all methods are shown in bold and underlined , respectively.
Figure 2: Ablation studies and analysis of GDLoRA. (a) Comparison of gradient components on six GLUE tasks with RoBERTa-Base and RoBERTa-Large. (b) Sensitivity to the base-update learning rate when fine-tuning Llama-3-8B on Math10K. Shaded bands indicate ±1 standard deviation. (c) Average mathematical reasoning accuracy and memory usage with Llama-3-8B on Math10K.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Gradient subspace drift and projection refresh costs. (a) Pairwise similarity between the leading rank-8 gradient subspaces of model.layers.15.self_attn.v_proj over 50 optimizer steps. (b) Accuracy and training time of native Fira with projection refresh intervals of 200, 150, 100, and 50 steps on Math10K with Llama-3-8B. GDLoRA is shown as a fixed reference; its repeated values correspond to the same configuration.
Figure 4: Gradient subspace variation across network depths and attention projections. Each heatmap shows pairwise similarity between leading rank-8 gradient subspaces over 50 optimizer steps during Llama-3-8B fine-tuning on Math10K. From left to right, the top row shows the layer-0 key, layer-0 value, and layer-15 query projections; the bottom row shows the layer-15 value, layer-31 query, and layer-31 value projections. Layer numbers follow the model’s module indices. The layer-15 value projection is also shown in Figure 3 (a). All panels use the same color scale.
Operation
Fira
GDLoRA
Full-weight gradient
2Ndodi
2Ndodi
Gradient projection and reconstruction
6dodir
8dodir
Basis construction
CSVD/T
CQR
Adapter forward and backward
0
6Nr(do+di)
Appendix
Table 5: Leading FLOPs per layer and optimization step. Fira’s SVD cost is amortized over its refresh interval T . GDLoRA recomputes both bases every step. Projection counts assume rank- r bases and efficient matrix-product ordering.
Method
SST-2
MRPC
CoLA
QNLI
RTE
STS-B
FFT
10.16
1.64
5.78
19.56
2.21
3.64
LoRA
6.92
0.58
3.54
13.24
1.06
2.50
PiSSA
13.06
1.08
4.92
16.78
1.47
2.98
DoRA
9.84
0.65
4.03
13.78
1.19
2.52
rsLoRA
9.98
0.57
3.51
14.23
1.03
2.50
Fira
16.18
1.53
6.31
19.65
2.34
4.35
Appendix
Table 6: Training time (hours) on six GLUE tasks with RoBERTa-Base.
Model
Hyperparameter
SST-2
MRPC
CoLA
QNLI
RTE
STS-B
Optimizer
AdamW
Warmup ratio
0.06
Learning-rate schedule
Linear
RoBERTa-Base
# GPUs
1
Epochs
60
30
80
25
160
80
Learning rate (LoRA)
1e-4
Appendix
Table 7: Hyperparameters used for GDLoRA in the natural language understanding experiments on the GLUE benchmark.
Method
AddSub
MultiArith
SingleEq
SVAMP
GSM8K
AQuA
Avg.
LoRA ( r=8 )
81.52
87.28
91.54
66.93
55.34
22.44
67.51
AdaLoRA
80.57
84.22
89.47
64.80
52.12
25.44
66.10
IA 3
86.84
83.50
86.42
64.50
50.64
26.38
66.38
FourierFT
85.06
88.67
88.98
64.30
45.40
27.95
66.73
MiSS
84.67
86.93
88.71
64.62
55.21
25.35
67.58
GDLoRA
84.81
90.69
92.52
69.61
56.03
26.38
70.01
Appendix
Table 8: Additional comparison on six mathematical reasoning tasks with Llama-3-8B. Scores are accuracy (%); Avg. is the arithmetic mean across the six tasks. The best mean score in each column is shown in bold .
Method
AddSub
MultiArith
SingleEq
SVAMP
GSM8K
AQuA
Avg.
LoRA ( r=8 )
81.52
87.28
91.54
66.93
55.34
22.44
67.51
LoRA ( r=16 )
83.56
91.34
92.43
68.81
56.97
24.89
69.67
LoRA ( r=32 )
85.14
91.87
93.45
69.72
57.66
26.95
70.80
GDLoRA ( r=8 )
84.81
90.69
92.52
69.61
56.03
26.38
70.01
GDLoRA ( r=16 )
88.15
92.16
93.78
70.85
57.94
28.89
71.96
GDLoRA ( r=32 )
90.36
94.48
94.23
71.35
58.47
27.94
72.81
Appendix
Table 9: LoRA and GDLoRA across adapter ranks on six mathematical reasoning tasks with Llama-3-8B. Scores are accuracy (%); Avg. is the arithmetic mean across tasks. GDLoRA rows are shaded blue, and the best score in each column is shown in bold .
Hyperparameter
Setting
Optimizer
AdamW
Learning rate (LoRA)
1e-4
Normal-gradient learning rate
0.5
Weight decay
0.0
Warmup ratio
0.06
Learning-rate schedule
Linear
Appendix
Table 10: Training and evaluation settings for mathematical reasoning with Llama-3-8B fine-tuned on Math10K.
Hyperparameter
Setting
Optimizer
AdamW
Learning rate (LoRA)
3e-5
Normal-gradient learning rate
0.5
Weight decay
0.0
Warmup ratio
0.06
Learning-rate schedule
Linear
Appendix
Table 11: Training and evaluation settings for commonsense reasoning with Gemma-7B fine-tuned on Commonsense170K.
Model
Hyperparameter
Oxford Pets
Stanford Cars
CIFAR-10
DTD
EuroSAT
FGVC
RESISC45
CIFAR-100
Optimizer
AdamW
Weight decay
0.01
Learning-rate schedule
Linear
Epochs
20
ViT-Base
# GPUs
1
Learning rate (head)
5e-3
5e-2
5e-2
5e-2
5e-2
5e-2
5e-2
1e-2
Appendix
Table 12: Hyperparameters used for GDLoRA in the image classification experiments with ViT-Base and ViT-Large. The normal-gradient learning rate controls the direct updates to the base weights.