Organizations: School of Computer Science and Technology, Beijing Jiaotong University, Beijing 100044, China · Key Laboratory of Big Data & Artificial Intelligence in Transportation (Beijing Jiaotong University), Ministry of Education.
Low-Rank Adaptation (LoRA) achieves parameter-efficient fine-tuning by constraining model updates to a low-rank subspace and has been widely used in practice. However, LoRA typically employs a shared low-rank update across tokens, which limits its ability to fully exploit the adaptation subspace for tokens from different sequences. To address this issue, we propose an adaptive utilization of Low-Rank Adaptation (U-LoRA), which employs conditioned gating to explicitly learn effective token-level utilization of the limited low-rank adaptation subspace. Specifically, U-LoRA generates utilization coefficients along low-rank directions for each token and jointly coordinates and constrains them using sequence-level contextual information, thereby inducing more consistent adaptive patterns within a sentence. To further enhance training stability, we introduce a bias-corrected exponential moving average (EMA) historical prior that calibrates utilization signals across optimization steps, suppressing noise caused by batch-to-batch fluctuations. The effectiveness of our method arises from a better utilization of the existing low-rank subspace via input-conditioned strategies, rather than from expanding the subspace. Experiments on mathematical reasoning and natural language understanding benchmarks demonstrate that U-LoRA achieves competitive performance under comparable parameter budgets when with strong LoRA baselines and recent variants.
Figures & tables
Figure 1 : The same token may require different low-rank update directions under different task contexts.Vanilla LoRA reuses similar directions due to shared low-rank bases, while input-aware variants diversify directions but may lack task-consistent utilization.
Figure 2 : U-LoRA augments standard LoRA with token-wise, input-conditioned gating to modulate low-rank updates. Token features are projected into a rank-aligned space and combined with sequence-level context to coordinate utilization. During training, a lightweight historical prior stabilizes the gating signals across optimization steps, while the pretrained backbone remains frozen.
Model
Method
#Params
AddSub
MultiArith
SingleEq
GSM8K
AQuA
SVAMP
Avg
Gemma-7B
LoRA ( r=8 )
4.82M
87.59
90.33
89.76
56.10
29.13
75.70
71.44
LoRA ( r=16 )
9.63M
86.84
92.83
89.57
58.15
30.71
74.90
72.17
LoRA ( r=32 )
19.30M
86.58
91.50
91.93
58.45
32.28
75.50
72.71
DoRA ( r=16 )
9.98M
87.59
94.17
91.34
58.68
27.95
75.80
72.59
MELoRA ( r=16 )
9.63M
87.26
92.22
91.01
59.49
32.15
74.97
72.85
HydraLoRA ( r=8 )
11.10M
87.34
92.67
90.16
58.83
27.95
75.50
72.08
Table 1 : Mathematical reasoning benchmark results of different parameter-efficient fine-tuning methods on multiple LLMs. The highest average precision is bolded, and the second-highest is underlined.
Model
Method
#Params
RTE
MRPC
STS-B
CoLA
SST-2
QNLI
MNLI
QQP
Avg
RoBERTa-Base
LoRA ( r=8 )
0.29M
72.56
87.25
87.12
56.10
93.46
91.58
84.89
87.46
82.55
LoRA ( r=16 )
0.59M
72.92
87.99
87.46
55.21
93.81
91.89
85.52
87.79
82.82
LoRA ( r=32 )
1.18M
75.09
89.22
88.01
58.58
93.58
90.12
85.84
88.37
83.60
DoRA ( r=16 )
0.61M
74.85
87.83
88.48
56.46
93.39
91.86
85.25
87.89
83.25
MELoRA ( r=16 )
0.59M
75.45
88.73
87.27
54.43
93.00
91.51
84.93
87.53
82.86
HydraLoRA ( r=8 )
0.65M
73.65
89.46
88.53
57.03
93.23
91.89
85.52
87.57
83.36
Table 2 : GLUE benchmark results of different parameter-efficient fine-tuning methods on RoBERTa-Base and RoBERTa-Large models. The highest average precision is bolded, and the second-highest is underlined.
Figure 3 : Scalability analysis on mathematical reasoning tasks using Qwen3-8B. (a) Comparison across different LoRA ranks. (b) Impact of target modules (tuning granularity) with rank r=8 .
Method
Avg
Δ
U-LoRA (Full)
83.97
+0.00
w/o Tok
78.95
-5.02
w/o Seq
79.16
-4.81
w/o Ema
79.66
-4.31
w/o Tok, Seq
78.90
-5.07
w/o Tok, Ema
80.32
-3.65
Table 3 : Component ablations of U-LoRA on Qwen3-8B for mathematical reasoning. We report the unweighted mean accuracy ( Avg ) across benchmarks; Δ denotes the change relative to Full . Components: token-wise utilization ( Tok ), sequence-level aggregation ( Seq ), and EMA historical prior ( Ema ). Details of all results are provided in Appendix C .
Figure 4 : PCA visualization of low-rank parameter updates ΔQ for RoBERTa-Large during multi-task fine-tuning. Implementation details are provided in the Appendix B .
Figure 5 : Depth-wise differences in low-rank utilization on RoBERTa-Large. Heatmaps report v^1(l)−v^2(l) for v^(l)=v(l)/∥v(l)∥2 (rank representation before B ). Left: layer–rank difference heatmaps for two controlled comparisons: ( Top ) the same token book under verb vs. noun contexts; ( Bottom ) different tokens ( a vs. book ) within the same sentence. Right: the corresponding rank-wise cosine similarity across layers.
Table 5 : Statistics of the arithmetic reasoning benchmark datasets.
Corpus
Train
Valid
Test
Metric
RTE
2.5k
277
3k
Accuracy
MRPC
3.7k
408
1.7k
Accuracy
STS-B
5.7k
1.5k
1.4k
Pearson corr.
CoLA
8.5k
1,043
1,063
Matthews corr.
SST-2
67k
872
1.8k
Accuracy
QNLI
105k
5.5k
5.5k
Accuracy
Appendix
Table 6 : Statistics of the GLUE benchmark datasets.
Method
AddSub
MultiArith
SingleEq
GSM8K
AQuA
SVAMP
Avg
U-LoRA (Full)
91.39
98.50
97.64
85.29
43.31
87.70
83.97
w/o Tok
89.87
91.67
90.94
76.57
40.55
84.10
78.95
w/o Seq
90.13
97.67
94.29
74.60
31.89
86.40
79.16
w/o Ema
90.89
95.00
91.73
76.04
37.01
87.30
79.66
w/o Tok, Seq
91.90
97.00
92.72
74.37
31.50
85.90
78.90
w/o Tok, Ema
93.16
90.67
91.93
75.66
42.91
87.60
80.32
Appendix
Table 7 : Component ablations of U-LoRA on Qwen3-8B for mathematical reasoning. We report accuracies on each benchmark; Avg denotes the unweighted mean. Components: (1) token-wise utilization, (2) sequence-level aggregation, and (3) EMA historical prior.
Target
Method
#Params
AddSub
MultiArith
SingleEq
GSM8K
AQuA
SVAMP
Avg
Q
LoRA
2.36M
89.87
96.33
92.13
73.92
31.89
83.00
77.86
U-LoRA
3.54M
89.37
97.83
93.50
77.10
43.31
83.40
80.75
QK
LoRA
3.83M
86.84
97.67
90.94
75.89
36.22
81.50
78.18
U-LoRA
6.20M
87.85
98.83
97.05
83.85
40.16
80.10
81.31
QKV
LoRA
5.31M
90.38
96.83
93.70
74.00
30.71
85.50
78.52
U-LoRA
8.86M
91.39
98.50
97.64
85.29
43.31
87.70
83.97
Appendix
Table 8 : The accuracy of LoRA and U-LoRA with varying target modules on mathematical reasoning tasks using Qwen3-8B.
Rank
Method
#Params
AddSub
MultiArith
SingleEq
GSM8K
AQuA
SVAMP
Avg
r=2
LoRA
1.33M
90.13
95.83
92.91
74.45
28.35
83.90
77.60
U-LoRA
2.21M
90.38
98.67
93.31
79.83
44.49
85.70
82.06
r=4
LoRA
2.65M
87.34
97.83
92.91
73.92
31.10
85.00
78.02
U-LoRA
4.43M
93.42
99.00
96.06
84.23
38.98
87.50
83.20
r=8
LoRA
5.31M
90.38
96.83
93.70
74.00
30.71
85.50
78.52
U-LoRA
8.86M
91.39
98.50
97.64
85.29
43.31
87.70
83.97
Appendix
Table 9 : The accuracy of LoRA and U-LoRA with varying ranks on mathematical reasoning tasks using Qwen3-8B.
Method
50 Steps
100 Steps
200 Steps
300 Steps
400 Steps
Final Loss
Avg
LoRA
0.3657
0.2759
0.2788
0.2711
0.2615
0.2710
78.52
DoRA
0.3549
0.2713
0.2747
0.2662
0.2559
0.2652
79.79
TopLoRA
0.3494
0.2707
0.2746
0.2638
0.2505
0.2572
80.75
U-LoRA
0.3418
0.2694
0.2745
0.2592
0.2470
0.2543
83.97
Appendix
Table 10 : Training Loss and Accuracy Comparison Across Methods (Qwen3-8B).
Model
Method
#Params
AddSub
MultiArith
SingleEq
GSM8K
AQuA
SVAMP
Avg
Qwen-72B (Base)
Base
—
90.63
96.67
90.75
65.13
24.80
89.50
76.25
LoRA ( r=16 )
44.56M
93.16
98.83
94.49
81.73
41.73
87.50
82.91
TopLoRA ( r=8 )
38.01M
93.42
98.83
96.65
84.08
40.55
90.30
83.97
U-LoRA ( r=8 )
38.05M
93.16
99.50
97.05
85.29
44.49
90.60
85.08
Qwen-72B (Instruct)
Base
—
89.37
97.50
91.73
71.72
22.83
88.50
76.94
LoRA ( r=16 )
44.56M
88.35
98.17
94.69
85.06
39.76
85.30
81.89
Appendix
Table 11 : Scaling results on Qwen2.5 72B and instruction-tuned models. Qwen denotes Qwen2.5.
Figure 6 : Training Loss Curves of LoRA and U-LoRA on Mathematical Reasoning Tasks.
AddSub
MultiArith
SingleEq
GSM8K
AQuA
SVAMP
Avg
β=0.9
92.66
98.67
97.44
83.40
37.40
89.80
83.23
β=0.95
92.66
99.00
97.24
83.17
38.58
89.70
83.39
β=0.98
93.16
99.00
96.85
82.79
37.40
90.20
83.23
β=0.99
91.39
98.50
97.64
85.29
43.31
87.70
83.97
β=0.995
92.91
99.00
97.05
83.55
41.73
90.80
84.17
β=0.998
93.67
99.00
96.85
84.08
40.55
90.70
84.14
Appendix
Table 12 : The effect of different β values in U-LoRA on mathematical reasoning benchmarks on the Qwen3-8B model.
Model
Method
RTE
MRPC
STS-B
CoLA
SST-2
QNLI
MNLI
QQP
RoBERTa-Base
LoRA ( r=8 )
0.51
0.46
0.27
0.57
0.11
0.15
0.11
0.10
LoRA ( r=16 )
0.84
0.40
0.34
0.64
0.29
0.11
0.04
0.08
LoRA ( r=32 )
0.78
0.35
0.39
0.56
0.14
1.00
0.11
0.03
DoRA ( r=16 )
0.95
0.46
0.31
0.57
0.14
0.02
0.24
0.11
MELoRA ( r=16 )
0.96
0.64
0.62
0.52
0.29
0.14
0.14
0.03
HydraLoRA ( r=8 )
0.59
0.87
0.75
0.61
0.34
0.17
0.21
0.08
Appendix
Table 13 : The standard deviation of different methods on the GLUE benchmark.
Method
Gemma-7B
LLaMA-3-8B
Qwen3-8B
Qwen2.5-14B
LoRA ( r=8 )
0.19
0.88
0.42
0.38
LoRA ( r=16 )
0.61
0.62
0.36
0.34
LoRA ( r=32 )
0.48
0.80
0.41
0.39
DoRA ( r=16 )
0.58
0.25
0.22
0.21
MELoRA ( r=16 )
0.31
0.67
0.28
0.26
HydraLoRA ( r=8 )
0.63
0.39
0.24
0.20
Appendix
Table 14 : The standard deviation of the average accuracy on mathematical reasoning tasks.
Low-Rank Adaptation (LoRA) is the most widely adopted method for fine-tuning large language models. Notably, LoRA is inherently overparameterized: multiple pairs of low-rank factors can yield the same adapted weight matrix. We show--both theoretically and empirically--that these pairs exhibit significantly different condition numbers. As a result, converging to different loss minimizers directly impacts the convergence rate of LoRA. Building on this observation, we introduce Balanced Low-Rank Adaptation (BaLoRA), a variant of LoRA that projects iterates onto a balanced manifold. This manifold improves the conditioning of the loss landscape while preserving the adapted matrix. The projection step is computationally lightweight and integrates seamlessly into existing fine-tuning pipelines. Empirically, BaLoRA converges faster than standard LoRA and achieves superior performance across a range of fine-tuning tasks.
Valérie Castin, Kimia Nadjahi, Pierre Ablin +1
1 ´Ecole Normale Sup´erieure PSL, Paris, France · 2CNRS · 3Apple, Paris, France.
Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT), yet a performance gap can remain relative to full fine-tuning (FFT). Many LoRA variants improve the initialization or optimization of low-rank factors. At each training step, however, their first-order weight-space directions are constrained by the current parameterization. We characterize the corresponding LoRA-accessible gradient space and show that it coincides with the tangent space induced by the current LoRA parameterization. This characterization yields an orthogonal decomposition of the full weight gradient at the current model parameters. We term the component orthogonal to this space the normal gradient. Based on this decomposition, we propose GDLoRA (Gradient-Decomposed Low-Rank Adaptation). GDLoRA reconstructs the full weight gradient from forward activations and backward signals, extracts its normal component, and directly updates the base weights with this component, while retaining standard AdamW optimization for the LoRA factors. GDLoRA incorporates complementary normal gradients without increasing standard LoRA's optimizer-state memory budget under matched adapter and optimizer configurations. Experiments on natural language understanding, mathematical reasoning, commonsense reasoning, and image classification show that GDLoRA consistently improves over LoRA and narrows the performance gap to FFT. The code is available at https://anonymous.4open.science/r/GDLoRA.
Yihao Ouyang, Shiwei Li, Haozhao Wang +5
Huazhong University of Science and Technology, Wuhan, China · Hebei University of Technology, Tianjin, China
Low-Rank Adaptation (LoRA) has become a widely adopted parameter-efficient fine-tuning method for large language models, with its effectiveness largely influenced by the allocation of ranks and scaling factors, as well as initialization. Existing LoRA variants typically address only one of these factors, often at the cost of increased training complexity or reduced practical efficiency. In this work, we present Task-aware Low-Rank Adaptation (TLoRA), a unified framework that jointly optimizes initialization and resource allocation at the outset of training. TLoRA introduces a data-driven initialization strategy that aligns the LoRA A matrix with task-relevant subspaces by performing singular value decomposition on the product of pre-trained weights and input activation covariance. After this, the A matrix is frozen, and only the B matrix is trained. Furthermore, TLoRA employs a sensitivity-based importance metric to adaptively allocate ranks and scaling factors across layers under a fixed parameter budget. We conduct extensive experiments that demonstrate TLoRA consistently performs excellently across various tasks, including natural language understanding, commonsense reasoning, math reasoning, code generation, and chat generation, while significantly reducing the number of trainable parameters.
Weicheng Lin, Yi Zhang, Jiawei Dang +1
College of Computer Science and Software Engineering, Shenzhen University, China