Organizations: School of Computer Science and Technology, Beijing Jiaotong University, Beijing 100044, China · Key Laboratory of Big Data & Artificial Intelligence in Transportation (Beijing Jiaotong University), Ministry of Education.
Low-Rank Adaptation (LoRA) achieves parameter-efficient fine-tuning by constraining model updates to a low-rank subspace and has been widely used in practice. However, LoRA typically employs a shared low-rank update across tokens, which limits its ability to fully exploit the adaptation subspace for tokens from different sequences. To address this issue, we propose an adaptive utilization of Low-Rank Adaptation (U-LoRA), which employs conditioned gating to explicitly learn effective token-level utilization of the limited low-rank adaptation subspace. Specifically, U-LoRA generates utilization coefficients along low-rank directions for each token and jointly coordinates and constrains them using sequence-level contextual information, thereby inducing more consistent adaptive patterns within a sentence. To further enhance training stability, we introduce a bias-corrected exponential moving average (EMA) historical prior that calibrates utilization signals across optimization steps, suppressing noise caused by batch-to-batch fluctuations. The effectiveness of our method arises from a better utilization of the existing low-rank subspace via input-conditioned strategies, rather than from expanding the subspace. Experiments on mathematical reasoning and natural language understanding benchmarks demonstrate that U-LoRA achieves competitive performance under comparable parameter budgets when with strong LoRA baselines and recent variants.
Figures & tables
Figure 1 : The same token may require different low-rank update directions under different task contexts.Vanilla LoRA reuses similar directions due to shared low-rank bases, while input-aware variants diversify directions but may lack task-consistent utilization.
Figure 2 : U-LoRA augments standard LoRA with token-wise, input-conditioned gating to modulate low-rank updates. Token features are projected into a rank-aligned space and combined with sequence-level context to coordinate utilization. During training, a lightweight historical prior stabilizes the gating signals across optimization steps, while the pretrained backbone remains frozen.
Model
Method
#Params
AddSub
MultiArith
SingleEq
GSM8K
AQuA
SVAMP
Avg
Gemma-7B
LoRA ( r=8 )
4.82M
87.59
90.33
89.76
56.10
29.13
75.70
71.44
LoRA ( r=16 )
9.63M
86.84
92.83
89.57
58.15
30.71
74.90
72.17
LoRA ( r=32 )
19.30M
86.58
91.50
91.93
58.45
32.28
75.50
72.71
DoRA ( r=16 )
9.98M
87.59
94.17
91.34
58.68
27.95
75.80
72.59
MELoRA ( r=16 )
9.63M
87.26
92.22
91.01
59.49
32.15
74.97
72.85
HydraLoRA ( r=8 )
11.10M
87.34
92.67
90.16
58.83
27.95
75.50
72.08
Table 1 : Mathematical reasoning benchmark results of different parameter-efficient fine-tuning methods on multiple LLMs. The highest average precision is bolded, and the second-highest is underlined.
Model
Method
#Params
RTE
MRPC
STS-B
CoLA
SST-2
QNLI
MNLI
QQP
Avg
RoBERTa-Base
LoRA ( r=8 )
0.29M
72.56
87.25
87.12
56.10
93.46
91.58
84.89
87.46
82.55
LoRA ( r=16 )
0.59M
72.92
87.99
87.46
55.21
93.81
91.89
85.52
87.79
82.82
LoRA ( r=32 )
1.18M
75.09
89.22
88.01
58.58
93.58
90.12
85.84
88.37
83.60
DoRA ( r=16 )
0.61M
74.85
87.83
88.48
56.46
93.39
91.86
85.25
87.89
83.25
MELoRA ( r=16 )
0.59M
75.45
88.73
87.27
54.43
93.00
91.51
84.93
87.53
82.86
HydraLoRA ( r=8 )
0.65M
73.65
89.46
88.53
57.03
93.23
91.89
85.52
87.57
83.36
Table 2 : GLUE benchmark results of different parameter-efficient fine-tuning methods on RoBERTa-Base and RoBERTa-Large models. The highest average precision is bolded, and the second-highest is underlined.
Figure 3 : Scalability analysis on mathematical reasoning tasks using Qwen3-8B. (a) Comparison across different LoRA ranks. (b) Impact of target modules (tuning granularity) with rank r=8 .
Method
Avg
Δ
U-LoRA (Full)
83.97
+0.00
w/o Tok
78.95
-5.02
w/o Seq
79.16
-4.81
w/o Ema
79.66
-4.31
w/o Tok, Seq
78.90
-5.07
w/o Tok, Ema
80.32
-3.65
Table 3 : Component ablations of U-LoRA on Qwen3-8B for mathematical reasoning. We report the unweighted mean accuracy ( Avg ) across benchmarks; Δ denotes the change relative to Full . Components: token-wise utilization ( Tok ), sequence-level aggregation ( Seq ), and EMA historical prior ( Ema ). Details of all results are provided in Appendix C .
Figure 4 : PCA visualization of low-rank parameter updates ΔQ for RoBERTa-Large during multi-task fine-tuning. Implementation details are provided in the Appendix B .
Figure 5 : Depth-wise differences in low-rank utilization on RoBERTa-Large. Heatmaps report v^1(l)−v^2(l) for v^(l)=v(l)/∥v(l)∥2 (rank representation before B ). Left: layer–rank difference heatmaps for two controlled comparisons: ( Top ) the same token book under verb vs. noun contexts; ( Bottom ) different tokens ( a vs. book ) within the same sentence. Right: the corresponding rank-wise cosine similarity across layers.
Table 5 : Statistics of the arithmetic reasoning benchmark datasets.
Corpus
Train
Valid
Test
Metric
RTE
2.5k
277
3k
Accuracy
MRPC
3.7k
408
1.7k
Accuracy
STS-B
5.7k
1.5k
1.4k
Pearson corr.
CoLA
8.5k
1,043
1,063
Matthews corr.
SST-2
67k
872
1.8k
Accuracy
QNLI
105k
5.5k
5.5k
Accuracy
Appendix
Table 6 : Statistics of the GLUE benchmark datasets.
Method
AddSub
MultiArith
SingleEq
GSM8K
AQuA
SVAMP
Avg
U-LoRA (Full)
91.39
98.50
97.64
85.29
43.31
87.70
83.97
w/o Tok
89.87
91.67
90.94
76.57
40.55
84.10
78.95
w/o Seq
90.13
97.67
94.29
74.60
31.89
86.40
79.16
w/o Ema
90.89
95.00
91.73
76.04
37.01
87.30
79.66
w/o Tok, Seq
91.90
97.00
92.72
74.37
31.50
85.90
78.90
w/o Tok, Ema
93.16
90.67
91.93
75.66
42.91
87.60
80.32
Appendix
Table 7 : Component ablations of U-LoRA on Qwen3-8B for mathematical reasoning. We report accuracies on each benchmark; Avg denotes the unweighted mean. Components: (1) token-wise utilization, (2) sequence-level aggregation, and (3) EMA historical prior.
Target
Method
#Params
AddSub
MultiArith
SingleEq
GSM8K
AQuA
SVAMP
Avg
Q
LoRA
2.36M
89.87
96.33
92.13
73.92
31.89
83.00
77.86
U-LoRA
3.54M
89.37
97.83
93.50
77.10
43.31
83.40
80.75
QK
LoRA
3.83M
86.84
97.67
90.94
75.89
36.22
81.50
78.18
U-LoRA
6.20M
87.85
98.83
97.05
83.85
40.16
80.10
81.31
QKV
LoRA
5.31M
90.38
96.83
93.70
74.00
30.71
85.50
78.52
U-LoRA
8.86M
91.39
98.50
97.64
85.29
43.31
87.70
83.97
Appendix
Table 8 : The accuracy of LoRA and U-LoRA with varying target modules on mathematical reasoning tasks using Qwen3-8B.
Rank
Method
#Params
AddSub
MultiArith
SingleEq
GSM8K
AQuA
SVAMP
Avg
r=2
LoRA
1.33M
90.13
95.83
92.91
74.45
28.35
83.90
77.60
U-LoRA
2.21M
90.38
98.67
93.31
79.83
44.49
85.70
82.06
r=4
LoRA
2.65M
87.34
97.83
92.91
73.92
31.10
85.00
78.02
U-LoRA
4.43M
93.42
99.00
96.06
84.23
38.98
87.50
83.20
r=8
LoRA
5.31M
90.38
96.83
93.70
74.00
30.71
85.50
78.52
U-LoRA
8.86M
91.39
98.50
97.64
85.29
43.31
87.70
83.97
Appendix
Table 9 : The accuracy of LoRA and U-LoRA with varying ranks on mathematical reasoning tasks using Qwen3-8B.
Method
50 Steps
100 Steps
200 Steps
300 Steps
400 Steps
Final Loss
Avg
LoRA
0.3657
0.2759
0.2788
0.2711
0.2615
0.2710
78.52
DoRA
0.3549
0.2713
0.2747
0.2662
0.2559
0.2652
79.79
TopLoRA
0.3494
0.2707
0.2746
0.2638
0.2505
0.2572
80.75
U-LoRA
0.3418
0.2694
0.2745
0.2592
0.2470
0.2543
83.97
Appendix
Table 10 : Training Loss and Accuracy Comparison Across Methods (Qwen3-8B).
Model
Method
#Params
AddSub
MultiArith
SingleEq
GSM8K
AQuA
SVAMP
Avg
Qwen-72B (Base)
Base
—
90.63
96.67
90.75
65.13
24.80
89.50
76.25
LoRA ( r=16 )
44.56M
93.16
98.83
94.49
81.73
41.73
87.50
82.91
TopLoRA ( r=8 )
38.01M
93.42
98.83
96.65
84.08
40.55
90.30
83.97
U-LoRA ( r=8 )
38.05M
93.16
99.50
97.05
85.29
44.49
90.60
85.08
Qwen-72B (Instruct)
Base
—
89.37
97.50
91.73
71.72
22.83
88.50
76.94
LoRA ( r=16 )
44.56M
88.35
98.17
94.69
85.06
39.76
85.30
81.89
Appendix
Table 11 : Scaling results on Qwen2.5 72B and instruction-tuned models. Qwen denotes Qwen2.5.
Figure 6 : Training Loss Curves of LoRA and U-LoRA on Mathematical Reasoning Tasks.
AddSub
MultiArith
SingleEq
GSM8K
AQuA
SVAMP
Avg
β=0.9
92.66
98.67
97.44
83.40
37.40
89.80
83.23
β=0.95
92.66
99.00
97.24
83.17
38.58
89.70
83.39
β=0.98
93.16
99.00
96.85
82.79
37.40
90.20
83.23
β=0.99
91.39
98.50
97.64
85.29
43.31
87.70
83.97
β=0.995
92.91
99.00
97.05
83.55
41.73
90.80
84.17
β=0.998
93.67
99.00
96.85
84.08
40.55
90.70
84.14
Appendix
Table 12 : The effect of different β values in U-LoRA on mathematical reasoning benchmarks on the Qwen3-8B model.
Model
Method
RTE
MRPC
STS-B
CoLA
SST-2
QNLI
MNLI
QQP
RoBERTa-Base
LoRA ( r=8 )
0.51
0.46
0.27
0.57
0.11
0.15
0.11
0.10
LoRA ( r=16 )
0.84
0.40
0.34
0.64
0.29
0.11
0.04
0.08
LoRA ( r=32 )
0.78
0.35
0.39
0.56
0.14
1.00
0.11
0.03
DoRA ( r=16 )
0.95
0.46
0.31
0.57
0.14
0.02
0.24
0.11
MELoRA ( r=16 )
0.96
0.64
0.62
0.52
0.29
0.14
0.14
0.03
HydraLoRA ( r=8 )
0.59
0.87
0.75
0.61
0.34
0.17
0.21
0.08
Appendix
Table 13 : The standard deviation of different methods on the GLUE benchmark.
Method
Gemma-7B
LLaMA-3-8B
Qwen3-8B
Qwen2.5-14B
LoRA ( r=8 )
0.19
0.88
0.42
0.38
LoRA ( r=16 )
0.61
0.62
0.36
0.34
LoRA ( r=32 )
0.48
0.80
0.41
0.39
DoRA ( r=16 )
0.58
0.25
0.22
0.21
MELoRA ( r=16 )
0.31
0.67
0.28
0.26
HydraLoRA ( r=8 )
0.63
0.39
0.24
0.20
Appendix
Table 14 : The standard deviation of the average accuracy on mathematical reasoning tasks.