Low-Rank Adaptation (LoRA) has become a standard approach for parameter-efficient fine-tuning, yet a fundamental practical question remains unresolved: how should the adapter rank be chosen? An overly small rank may lead to a poorly conditioned optimization landscape, whereas an unnecessarily large rank sacrifices the efficiency that motivates LoRA in the first place. Existing theoretical analyses provide only limited guidance on this trade-off, and their guarantees are typically established under restrictive theoretical settings. We address this gap by developing a substantially sharper landscape theory for LoRA, building on modern results from nonconvex low-rank matrix sensing. Our central insight is that the appropriate adapter rank should depend on the quality of the data-induced optimization geometry, rather than on the model alone. To formalize this connection, we introduce LoRA-RIP, a data-dependent restricted-isometry metric that characterizes the conditioning of the cross-entropy (CE) objective along LoRA-relevant low-rank directions. We prove that sufficient rank over-parameterization, with the required rank explicitly determined by the LoRA-RIP constant, eliminates spurious local minima, thereby extending existing RIP-based guarantees beyond the classical 1/3 regime. This characterization further enables principled data selection under a fixed rank budget. Experiments across language and vision tasks support these theoretical predictions, showing that rank and data quality are two coupled resources that should be jointly considered for more efficient and reliable LoRA fine-tuning.
Figures & tables
Figure 1: Rank-sweep validation CE. Left: MNLI with MPNet. Right: CIFAR-100 with ViT-B/16.
Rank
MNLI
Rank
CIFAR-100
Val acc
Test acc
Val CE
Val acc
Test acc
Val CE
4
0.6199
0.6167
0.8501
4
0.7392
0.7267
1.2299
16
0.6360
0.6324
0.8318
16
0.7400
0.7251
1.2291
64
0.6547
0.6506
0.8018
64
0.7414
0.7279
1.2254
256
0.6696
0.6634
0.7752
256
0.7422
0.7318
1.2244
Table 1: Representative rank-sweep summary at the best validation checkpoint.
Figure 2: Validation-accuracy trajectories for the MNLI gradient-coverage intervention.
Figure 3: Validation-accuracy trajectories for the MNLI saturation-band intervention.
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
Experiment
Dataset / model
LoRA target
Main variable
NLP rank sweep
MNLI / MPNet-base
layer 10 attention query
rank r∈{1,4,16,32,64,128,256,512}
CV rank sweep
CIFAR-100 / ViT-B/16
query projection in block 10 attention
rank r∈{1,4,16,32,64,128,256,512}
NLP coverage
MNLI / MPNet-base
layer 10 attention query
gradient-group sampling mass and rank
NLP saturation
MNLI / MPNet-base
layer 10 attention query
saturation band under fixed gradient-group quotas and rank
Appendix
Table D.1: Protocols for the rank sweeps and MNLI data interventions. Diagnostic and robustness studies are reported in Appendix E .
Experiment group
Runs
Batch size / epochs
NLP rank sweep
1 linear probe + 8 LoRA ranks
LP: 32 / 10; LoRA: 32 / 10
CV rank sweep
1 linear probe + 8 LoRA ranks
LP: 64 / 20; LoRA: 64 / 30
NLP gradient coverage
4 data conditions × 4 ranks
LoRA: 32 / 50; subset size 17,526
NLP saturation bands
3 bands × 4 ranks
LoRA: 32 / 50; subset size 17,526
Appendix
Table D.2: Hardware and training configurations for the rank sweeps and MNLI data interventions.
Rank
Best val acc
Test acc
Best val CE
Best train CE
Final train CE
Best epoch
1
0.6047
0.6036
0.8737
0.9506
0.9506
9
4
0.6199
0.6167
0.8501
0.9325
0.9325
10
16
0.6360
0.6324
0.8318
0.9184
0.9184
10
32
0.6476
0.6456
0.8130
0.9032
0.9032
9
64
0.6547
0.6506
0.8018
0.8939
0.8939
9
128
0.6630
0.6570
0.7883
0.8829
0.8829
7
Appendix
Table D.3: MNLI rank sweep. The test accuracy is reported at the best validation checkpoint.
Rank
Best val acc
Test acc
Best val CE
Best train CE
Final train CE
Best epoch
1
0.7384
0.7243
1.2393
0.1137
0.1137
2
4
0.7392
0.7267
1.2299
0.1101
0.1101
28
16
0.7400
0.7251
1.2291
0.1057
0.1057
8
32
0.7398
0.7265
1.2323
0.1013
0.1013
4
64
0.7414
0.7279
1.2254
0.0940
0.0940
16
128
0.7414
0.7296
1.2236
0.0840
0.0840
12
Appendix
Table D.4: CIFAR-100 rank sweep. The train loss improves with rank, while held-out accuracy is much flatter than in MNLI.
Figure D.1: Sampling mass over label-conditioned gradient groups. The heavy long-tail intervention collapses most mass onto the first few groups in every label, while the balanced condition spreads mass uniformly across the gradient groups.
Condition
Subset size
Unique ratio
Mass min
Mass median
Mass max
Max/min
Label counts
Balanced
17,526
1.0
0.1249
0.1250
0.1251
1.0
5847/5788/5891
Long-tail Light
17,526
1.0
0.0371
0.1025
0.2779
7.5
5847/5788/5891
Long-tail Heavy
17,526
1.0
0.0020
0.0364
0.5510
270.5
5847/5788/5891
Appendix
Table D.5: Group-mass checks for the MNLI gradient-coverage construction. Mass statistics are computed within label-conditioned gradient groups.
Rank
Condition
Best val acc
Test acc
Final train CE
Best epoch
Δ val / test
4
Balanced
0.5939
0.5954
0.9547
37
0.0000 / 0.0000
4
Long-tail Light
0.5868
0.5875
0.9601
48
-0.0071 / -0.0079
4
Long-tail Heavy
0.5853
0.5838
0.9350
48
-0.0086 / -0.0116
16
Balanced
0.6046
0.6033
0.9353
43
0.0000 / 0.0000
16
Long-tail Light
0.5955
0.5940
0.9436
42
-0.0091 / -0.0093
16
Long-tail Heavy
0.5961
0.5922
0.9127
44
-0.0085 / -0.0111
Appendix
Table D.6: MNLI gradient-coverage results. Deltas are relative to the balanced condition at the same rank.
Figure D.2: Validation cross-entropy trajectories for the MNLI gradient-coverage intervention. Each subplot fixes the LoRA rank and compares balanced group coverage with light and heavy long-tail group-mass distributions.
Condition
Quantile band
Score mean
Score p10
Score p50
Score p90
Subset size
Easy-band
[0.10,0.45]
0.2210
0.2044
0.2234
0.2344
17,526
Middle-band
[0.33,0.67]
0.2389
0.2320
0.2401
0.2445
17,526
Boundary-band
[0.62,0.95]
0.2477
0.2457
0.2481
0.2494
17,526
Appendix
Table D.7: Saturation-band checks for the MNLI experiment. The selected subsets all have size 17,526 , unique ratio 1.0 , no fallback groups, and fixed label-conditioned gradient-group quotas.
Figure D.3: Selected pairwise softmax-curvature scores for the three MNLI saturation bands. The boundary band has the highest average score, while the middle band avoids both very easy examples and the most boundary-concentrated tail.
Rank
Condition
Best val acc
Test acc
Best val CE
Best train CE
Best epoch
4
Easy-band
0.5853
0.5868
0.9010
0.9509
47
4
Middle-band
0.5962
0.5984
0.8843
0.9481
42
4
Boundary-band
0.5945
0.5947
0.8860
0.9545
42
16
Easy-band
0.5967
0.5959
0.8837
0.9326
46
16
Middle-band
0.6091
0.6074
0.8668
0.9290
39
16
Boundary-band
0.6056
0.6058
0.8676
0.9335
44
Appendix
Table D.8: MNLI saturation-band results. All conditions use the same subset size, label budget, and label-conditioned gradient-group quotas.
Figure D.4: Validation cross-entropy trajectories for the MNLI saturation-band intervention. The middle band is consistently strongest, while easy-band examples have lower softmax-side curvature leverage under matched gradient-group coverage.
Diagnostic
Dataset/model
Adapted matrix
Seeds/checkpoints
Main quantities
Rank/checkpoint
MNLI/MPNet-base
layer-10 query
r=4,64,128,256 ; 3 seeds; 0/5/10
projected CE curvature
Data condition
MNLI/MPNet-base
layer-10 query
3 diagnostic seeds
GN condition, stable rank, residual
Saturation condition
MNLI/MPNet-base
layer-10 query
3 diagnostic seeds
softmax weight and GN spectrum
Ambient reference
MNLI, CIFAR-100
layer-10 query
two LoRA warm starts/task
Fλ and spectral ranks
Multi-initialization
MNLI, CIFAR-100
layer-10 query
3 runs/rank
objective and held-out stability
Appendix
Table E.1: Landscape diagnostic protocols.
Rank
Checkpoint
CE
99% energy rank
GN floor
GN ceiling
GN condition
Full-H min eig.
Neg. eigs.
4
Start
0.959 ± 0.000
0.0 ± 0.0
4.911×10−7
1.606×10−5
32.70 ± 0.00
−9.012×10−6
8.00 ± 0.00
4
Middle
0.843 ± 0.001
2.0 ± 0.0
4.136×10−7
1.764×10−5
42.62 ± 2.07
−1.914×10−7
1.33 ± 0.58
4
Final
0.839 ± 0.001
2.0 ± 0.0
3.967×10−7
1.754×10−5
44.22 ± 2.30
−4.915×10−7
1.67 ± 0.58
64
Start
0.959 ± 0.000
0.0 ± 0.0
4.911×10−7
1.606×10−5
32.70 ± 0.00
−9.012×10−6
8.00 ± 0.00
64
Middle
0.800 ± 0.002
4.33 ± 0.58
3.445×10−7
1.966×10−5
56.83 ± 9.85
3.596×10−7
0.00 ± 0.00
64
Final
0.793 ± 0.002
5.0 ± 0.0
3.448×10−7
1.874×10−5
54.35 ± 10.00
4.796×10−7
0.00 ± 0.00
Appendix
Table E.2: MNLI projected curvature over training checkpoints. Values are mean ± standard deviation over three training seeds. GN floor and ceiling are shown in absolute units.
Figure E.1: Projected GN and CE-Hessian diagnostics over rank and checkpoint.
Task
Warm start
Final Fλ
Threshold rank
r90
r95
r99
Rel. prox. map.
MNLI
LoRA rank 256
0.790509
44
4
6
8
5.00×10−5
MNLI
LoRA rank 64
0.809354
35
3
4
5
7.26×10−5
CIFAR-100
LoRA rank 256
0.061640
406
24
33
62
4.30×10−4
CIFAR-100
LoRA rank 64
0.085553
457
9
13
25
2.95×10−3
Appendix
Table E.3: Ambient matrix-space proximal diagnostics from two LoRA warm starts. Energy ranks summarize spectral concentration, and the relative mapping is the normalized proximal-gradient mapping used to monitor the optimization residual.
Task
Rank 4
Rank 64
Rank 256
MNLI
0.06450
0.02130
0.00261
CIFAR-100
0.04666
0.02545
0.00038
Appendix
Table E.4: Best observed LoRA objective gaps to the corresponding ambient reference.
Task
Rank
Best validation accuracy
Final factor objective
MNLI
4
0.6199±0.0009
0.855593±0.000605
MNLI
64
0.6556±0.0007
0.812421±0.000531
MNLI
128
0.6626±0.0014
0.803074±0.001195
MNLI
256
0.6695±0.0009
0.793288±0.000174
CIFAR-100
4
0.7389±0.0009
0.108684±0.000127
CIFAR-100
64
0.7407±0.0005
0.087717±0.000223
Appendix
Table E.5: Three-run multi-initialization diagnostic. Accuracy is reported as a fraction; values are mean ± standard deviation.
Figure E.2: Multi-initialization performance and objective diagnostics.
Condition
GN floor ( ×10−7 )
GN ceiling ( ×10−5 )
GN condition
Stable rank
Residual ( ×10−5 )
Balanced
4.66±0.14
1.59±0.01
34.19±1.17
4.81±0.02
2.11±0.16
Light long-tail
4.88±0.14
1.73±0.02
35.37±0.97
4.47±0.04
3.27±0.17
Heavy long-tail
4.93±0.10
1.90±0.04
38.60±1.32
4.04±0.05
4.19±0.43
Appendix
Table E.6: Exact projected data-condition diagnostics at the shared zero-update model. Means and standard deviations are over three diagnostic seeds; training-trajectory results are reported separately in Table E.7 .
Figure E.3: Projected Gauss–Newton condition numbers and restricted curvature spectra for the gradient-coverage subsets.
Condition
Checkpoint
GN condition
Full-H min eig.
Negative eigs.
Residual/GN ceiling
Balanced
Start
32.70±0.00
−9.012e−6
8.00±0.00
1.194±0.000
Balanced
Middle
57.54±1.88
1.816e−7
0.00±0.00
0.755±0.059
Balanced
Final
59.60±1.52
2.096e−7
0.00±0.00
0.756±0.020
Light
Middle
41.92±3.85
−1.437e−6
3.67±0.58
0.721±0.217
Light
Final
40.16±2.69
−1.440e−6
3.00±1.00
0.799±0.178
Heavy
Middle
44.44±0.58
−2.543e−6
5.00±0.00
0.594±0.033
Appendix
Table E.7: Complete training-trajectory measurements. Reported uncertainties are standard deviations over three training seeds; start is one shared state. The accompanying figure displays the two projected CE-Hessian sign diagnostics.
Figure E.4: Projected CE-Hessian geometry along MNLI training trajectories. Left: minimum eigenvalue. Right: number of negative eigenvalues. All conditions use the same diagnostic set and projection basis; curves and shaded bands summarize three training seeds.
Band
Pairwise score s
GN trace ( ×10−5 )
GN condition
Easy
0.2210
7.43±0.03
32.40±0.55
Middle
0.2389
7.71±0.10
33.16±1.59
Boundary
0.2477
7.73±0.07
32.59±1.21
Appendix
Table E.8: Confidence-band curvature summaries at the shared reference model. GN statistics are means ± standard deviations over three diagnostic seeds.
Rank
Easy
Middle
Boundary
4
58.83±0.47
59.58±0.04
59.43±0.04
16
60.10±0.60
60.89±0.07
60.59±0.02
32
60.95±0.55
61.36±0.12
61.36±0.06
64
61.65±0.33
62.04±0.19
62.13±0.05
Appendix
Table E.9: Three-seed MNLI saturation-band training results. Entries are validation accuracy in percent, reported as mean ± standard deviation.
Rank
Condition
Validation accuracy (%)
Test accuracy (%)
16
Balanced
44.54±0.06
45.04±0.06
16
Light long-tail
43.42±0.07
43.76±0.11
16
Heavy long-tail
42.25±0.03
42.86±0.03
64
Balanced
45.12±0.09
45.16±0.18
64
Light long-tail
44.10±0.16
44.38±0.03
64
Heavy long-tail
42.92±0.11
43.54±0.16
Appendix
Table E.10: DeBERTa-v3-base on MNLI: gradient-coverage results over seeds 0,1,2 at the best validation-CE checkpoint of each run. Accuracy is in percent; bold marks the best condition at each rank.
Rank
Condition
Validation accuracy (%)
Test accuracy (%)
16
Balanced
71.73±0.27
72.16±0.33
16
Light long-tail
70.63±0.32
70.38±0.26
16
Heavy long-tail
68.15±0.33
66.66±0.19
64
Balanced
75.01±0.29
76.06±0.33
64
Light long-tail
73.83±0.27
73.71±0.32
64
Heavy long-tail
71.02±0.21
69.27±0.20
Appendix
Table E.11: MPNet-base on QNLI: three-condition gradient-coverage experiment with 5,125 examples. Results summarize seeds 0,1,2 at each run’s best validation-CE checkpoint. Accuracy is in percent; bold marks the best condition at each rank.
Adapter location
Balanced gain
Layer-5 query
+0.890
Layer-7 query
+1.063
Layer-8 query
+1.939
All-layer query/value
+1.124
Appendix
Table E.12: Adapter-location robustness on MNLI. Gains are Balanced minus Heavy test accuracy in percentage points.
Low-Rank Adaptation (LoRA) has become the standard mechanism for fine-tuning large pretrained models, yet its statistical properties remain only partially understood. Existing generalization results provide upper bounds of the form O~(sqrt(rd/n)) or O~(rd/n), but a matching lower bound is missing, and the question of how to choose the LoRA rank r has no formal answer. Both gaps are closed here. A local Rademacher argument establishes an upper bound of O~(rd/n) on the excess risk of the empirical risk minimizer over rank-r LoRA, whenever the target adaptation has rank at most r. A matching minimax lower bound of Omega(rd/n) is then proved via a Fano-type packing of the rank-r subspace of R^{d x d}; the bound applies to any estimator whose output lies in the rank-r LoRA class. Combining the two yields a rank-selection dichotomy. For the constrained empirical risk minimizer, the optimal rank equals the intrinsic rank r*, and over-ranking strictly hurts. For adaptive estimators of the nuclear-norm-then-truncate type, over-ranking is harmless and the rate saturates at Theta~(r* d / n) regardless of r. Taken together, the three results characterize the statistical complexity of LoRA fine-tuning within the well-specified locally quadratic regime, and identify the empirically observed over-parameterization penalty as a property of unregularized empirical risk minimization rather than of the LoRA class itself. Predictions of the theory are verified on a synthetic trace-regression benchmark and on real LoRA fine-tuning across three (model, task) configurations covering DistilBERT and RoBERTa on SST-2 and MRPC. All configurations exhibit the predicted U-shape in validation loss, with two showing statistically significant loss inflation at large ranks (paired permutation p = 0.016).
Fine-tuning adapts a pre-trained model to downstream tasks using a small amount of labeled data. Low-Rank Adaptation (LoRA) is an efficient fine-tuning method that reduces memory and computation costs while often achieving performance close to full fine-tuning. Despite its widespread use, the theoretical behavior of LoRA is not yet well understood. In this paper, we study LoRA in a simple linear regression setting and compare its excess risk with that of full fine-tuning. Our analysis identifies regimes in which LoRA achieves lower excess risk than full fine-tuning in both overdetermined and underdetermined settings. Specifically, our theory predicts that LoRA can outperform full fine-tuning when the difference between the pretraining and the downstream tasks is effectively low-rank. We further show how the choice of LoRA rank affects generalization performance, explaining why using a very small rank can improve test accuracy in certain settings, even though it limits model expressivity. Finally, we support our theoretical results with experiments on practical tasks, suggesting that the identified tradeoffs and insights extend beyond linear regression.
Ali Zindari, Rotem Mulayoff, Sebastian U. Stich
Universität des Saarlandes · CISPA Helmholtz Center for Information Security
Low-Rank Adaptation (LoRA) is an effective approach for adapting large pretrained models by learning low-rank weight updates. In practice, the LoRA rank is used to control an adapter's parameter budget and representational capacity. We show that this view is incomplete: while the nominal rank determines the representational capacity, the optimizer shapes how much of that capacity is used in the induced weight-space updates. In a case study of GPT-2 adaptation with LoRA, we observe a strong rank-dependent optimizer effect. Despite using the same nominal rank, AdamW often produces per-step updates with concentrated singular spectra and low effective rank, whereas Muon uses a richer set of directions and benefits more consistently from increasing LoRA rank. These observations motivate ISO-LoRA, an optimizer that couples the LoRA factor updates through spectral descent on the induced tangent perturbation in weight space. ISO-LoRA promotes updates that distribute energy more evenly across singular directions, improving rank utilization while preserving compatibility with the LoRA parameterization. We complement this design with theoretical guarantees showing that ISO-LoRA can achieve higher effective rank than standard factor-wise optimizers through a one-step analysis under a stylized spiked-gradient model. We validate this design on language-model adaptation across 0.1B-7B-parameter models, where ISO-LoRA improves effective rank and downstream performance, with the strongest gains at moderate-to-large LoRA ranks. Our results highlight rank utilization as a key factor in LoRA optimization and suggest that optimizer design offers an important path toward stronger parameter-efficient adaptation.
Zihan Zhu, Zhehang Du, Xuyang Chen +5
University of Pennsylvania · DRW Associates LLC · University of Chicago