cs.LGSep 28, 2026
SaveRotated Manifold Optimization for Low-Rank Adaptation
Organizations: ETH Zurich · Microsoft Research Cambridge
Abstract
We propose a novel optimizer for low-rank adaptation (LoRA) that explicitly incorporates the gauge symmetry of low-rank factorization. Our optimizer extends recent matrix optimizers for full-parameter training to the manifold of fixed-rank matrices by interpreting them as normalization under a rotated basis. We show how rotation and normalization can be integrated with the fixed-rank manifold efficiently. Our optimizer converges faster to lower held-out loss and achieves better or comparable downstream performance on both supervised finetuning and reinforcement learning tasks.
Figures & tables
Figure 1: Validation loss (mean cross-entropy across response tokens) over the last 100 training steps on the math SFT task at rank (see Appendix C for the case and additional results across different ranks).
| Llama-3.2-3B | Llama-3.1-8B | |||||
|---|---|---|---|---|---|---|
| MetaMathQA val | GSM8K | MATH | MetaMathQA val | GSM8K | MATH | |
| Pretrained | 728.23 | 18.1 | 7.3 | 652.22 | 41.9 | 15.5 |
| RoM (ours) | 182.69 0.21 | 65.4 0.1 | 18.2 0.3 | 157.60 0.06 | 77.7 1.0 | 29.2 0.3 |
| Adam | 184.11 0.14 | 64.9 0.9 | 18.0 0.1 | 158.70 0.26 | 77.2 0.6 | 28.4 0.4 |
| Muon | 185.42 0.19 | 64.3 0.9 | 17.8 0.4 | 159.45 0.13 | 77.3 0.9 | 28.0 0.5 |
| Riemannion | 185.79 0.22 | 64.0 0.7 | 17.7 0.4 | 159.91 0.13 | 76.9 0.7 | 29.0 0.6 |
Table 1: Math SFT results for the Llama models at rank : final MetaMathQA validation loss ( ) and test accuracies (%). Results are mean standard deviation over three seeds; the pretrained models are evaluated once. Bold marks the best entry and entries whose intervals overlap it; no entry is bold if all intervals overlap.
| Math SFT val | GSM8K | MATH | |
|---|---|---|---|
| RoM (ours) | 126.19 0.13 | 84.2 0.8 | 43.4 0.3 |
| Adam | 126.55 0.21 | 84.1 0.6 | 42.2 0.5 |
| Muon | 127.02 0.29 | 83.8 0.7 | 42.6 0.4 |
| Riemannion | 127.85 0.20 | 83.8 0.9 | 42.2 0.6 |
| LoRA-RITE | 129.71 0.13 | 83.7 0.5 | 42.7 0.2 |
| BaLoRA | 129.05 0.27 | 84.6 0.9 | 43.3 0.1 |
Table 2: Math SFT results for Gemma-2-27B at rank 16. Bold follows the convention in Table 1 .
Figure 2: Evaluation accuracy (%) during GRPO finetuning of Qwen2.5-3B with rank , on each of the three held-out benchmarks. The curves are smoothed using a 5-point centered moving average for visualization; higher accuracy is better.
| MATH-500 | GSM8K | Olympiad | Time | |
|---|---|---|---|---|
| Pretrained | 55.5 | 56.8 | 22.0 | – |
| Adam | 62.9 0.3 | 82.7 0.4 | 25.9 0.8 | |
| RoM (ours) | 64.9 0.6 | 84.2 0.4 | 29.0 0.8 |
Table 4: Final downstream accuracy (%) after GRPO finetuning over three independent runs. Results are mean standard deviation. “Time” denotes the median end-to-end wall clock time per optimizer step, normalized to Adam.
Figure 4: Llama-3.2-3B at rank on the math SFT task. (a) Validation loss over the last 100 training steps, removing parts of RoM . (b) Validation loss at different ratios of best learning rate.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Validation loss (mean cross-entropy across response tokens) over the last 100 training steps on the math SFT task at rank , the counterpart of Figure 1 .
Figure 6: Final validation loss for RoM and Adam across LoRA ranks on math SFT, using Llama-3.2-3B. Learning rate is independently tuned for each optimizer and rank.
Explore similar work
While low-rank adaptation (LoRA) is widely used for parameter-efficient model adaptation, how to regularize its training dynamics for stable and effective optimization remains underexplored. Because LoRA initializes the up-projection to zero, its early optimization dynamics are largely governed by the down-projection. Building on this observation, we introduce Normalized Low-Rank Adaptation (NoRA), a simple yet effective method that normalizes the down-projection matrices during training. We further show that the same normalization can be applied only at initialization, improving standard LoRA without requiring repeated normalization throughout training. Across pretraining, supervised finetuning, and reinforcement learning, NoRA consistently accelerates convergence, improves performance and training stability, and mitigates catastrophic forgetting. These benefits require neither additional trainable parameters nor inference-time computation, making NoRA a simple and broadly applicable enhancement to LoRA.
Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence
Low-Rank Adaptation (LoRA) is the most widely adopted method for fine-tuning large language models. Notably, LoRA is inherently overparameterized: multiple pairs of low-rank factors can yield the same adapted weight matrix. We show--both theoretically and empirically--that these pairs exhibit significantly different condition numbers. As a result, converging to different loss minimizers directly impacts the convergence rate of LoRA. Building on this observation, we introduce Balanced Low-Rank Adaptation (BaLoRA), a variant of LoRA that projects iterates onto a balanced manifold. This manifold improves the conditioning of the loss landscape while preserving the adapted matrix. The projection step is computationally lightweight and integrates seamlessly into existing fine-tuning pipelines. Empirically, BaLoRA converges faster than standard LoRA and achieves superior performance across a range of fine-tuning tasks.
Rank-Efficient LoRA via Joint Tangent-Space Optimization under Isotropic Curvature
Low-Rank Adaptation (LoRA) is an effective approach for adapting large pretrained models by learning low-rank weight updates. In practice, the LoRA rank is used to control an adapter's parameter budget and representational capacity. We show that this view is incomplete: while the nominal rank determines the representational capacity, the optimizer shapes how much of that capacity is used in the induced weight-space updates. In a case study of GPT-2 adaptation with LoRA, we observe a strong rank-dependent optimizer effect. Despite using the same nominal rank, AdamW often produces per-step updates with concentrated singular spectra and low effective rank, whereas Muon uses a richer set of directions and benefits more consistently from increasing LoRA rank. These observations motivate ISO-LoRA, an optimizer that couples the LoRA factor updates through spectral descent on the induced tangent perturbation in weight space. ISO-LoRA promotes updates that distribute energy more evenly across singular directions, improving rank utilization while preserving compatibility with the LoRA parameterization. We complement this design with theoretical guarantees showing that ISO-LoRA can achieve higher effective rank than standard factor-wise optimizers through a one-step analysis under a stylized spiked-gradient model. We validate this design on language-model adaptation across 0.1B-7B-parameter models, where ISO-LoRA improves effective rank and downstream performance, with the strongest gains at moderate-to-large LoRA ranks. Our results highlight rank utilization as a key factor in LoRA optimization and suggest that optimizer design offers an important path toward stronger parameter-efficient adaptation.