Parameter-efficient fine-tuning (PEFT) reduces the cost of adapting foundation models by focusing training on a small parameter subset. Complementary to this idea, we introduce RoSA (Rotational Sparse Adaptation), which narrows adaptation to a subset of layers at a time. RoSA freezes lower layers close to the input throughout training and rotates a trainable block over later layers, progressively increasing the number of frozen layers close to the input. This design reduces optimizer-state memory, shortens backpropagation, and even forward propagation if activations at the last frozen layer are cached. Because RoSA is orthogonal to the choice of trainable parameterization, it can be combined with PEFT methods or sparse optimizers within each active block. Experiments across multiple LLM architectures and tasks show that RoSA reduces peak memory while maintaining strong fine-tuning performance.
Figures & tables
Figure 1: RoSA training schedule. RoSA freezes a large input-side prefix, keeps an output-side anchor active, and rotates one consecutive trainable group through the later layers. Freezing the prefix removes optimizer states and parameter gradients for early layers, shortens the backward path, and enables future activation caching for the fixed prefix.
Table 2
Figure 2: Peak VRAM usage. Full fine-tuning with AdamW and full-model sparse optimization remain above the 40 GB budget, while RoSA fits within it and preserves strong performance.
Table 4Table 5Table 6
Figure 3: Layer-wise gradient norm analysis of Llama-2-7B during full fine-tuning. Left: Step 100, where activity is dominated by the input and head. Middle: Step 1100, where a “reasoning hump” emerges in Layers 8–18 while Layers 2–7 remain flat. Right: the average gradient profile over the first epoch, which highlights inefficient updates in deeper Layers 20–29. RoSA exploits this structure by freezing stable early Layers 0–7 and rotating updates through the more active middle layers.
Method
Ep.
Acc.
Peak
Later
Runtime
(%)
VRAM
VRAM
DoRA
3
81.58
16.42
–
15h 28m
RoSA-DoRA
6
81.79
12.87
11.04/9.44
23h 40m
RoSA-DoRA (3 seeds)
6
81.83 ± 0.32
12.87
–
–
Table 9: RoSA-DoRA on Qwen-2.5-3B commonsense reasoning. RoSA-DoRA maintains similar accuracy while reducing VRAM across rotation stages. The last row reports the mean and standard deviation over three seeds.
Task
Full
Core
Δ
OpenBookQA
79.2
79.4
+0.2
PIQA
82.9
83.1
+0.2
WinoGrande
82.2
82.2
0.0
ARC-Easy
83.8
83.6
-0.2
HellaSwag
89.3
89.0
-0.3
BoolQ
71.0
70.7
-0.3
Table 10: Core-only RoSA ablation. Restricting rotation to Layers 8–23 gives results almost identical to the full RoSA schedule. The average is computed over the full 8-task suite.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Baseline ρ=0.2,Δt=600
Low Freq. ρ=0.1,Δt=600
High Freq. ρ=0.1,Δt=100
ARC-C
69.9
68.3
67.5
ARC-E
83.8
82.8
83.2
BoolQ
71.0
69.7
68.8
HellaS
89.3
74.9
87.9
OBQA
79.2
77.8
78.0
PIQA
82.9
82.5
82.9
Appendix
Table 11: Task accuracy (%) under different sparsity-update settings. Lower density hurts performance when the sparse mask is refreshed slowly, but much of this loss can be recovered by refreshing the mask more often.
Optimizer states (NanoAdam, ∼ 495.5M tracked parameters)
2.52
Activations and other allocations
13.91
Peak allocated tensor memory (PyTorch)
32.62
CUDA context and caching allocator
∼ 2.3
Appendix
Table 12: Breakdown of the peak memory of RoSA + NanoAdam on Llama-2-7B commonsense reasoning ( ρ=0.2 , group size 8, micro-batch size 1, cutoff length 256).
Figure 4: Average per-layer ℓ2 gradient norms from short dense fine-tuning probes. Axis scales differ between panels.
Setting
Run
Avg Acc.
Llama-2-7B, RoSA + NanoAdam
original run (Table 2 )
79.61
seed 44
79.49
Qwen-2.5-3B, RoSA-DoRA
seed 42 (Table 9 )
81.79
seed 43
82.17
seed 44
81.53
mean ± std
81.83±0.32
Appendix
Table 13: Average commonsense reasoning accuracy (%) across random seeds. For RoSA-DoRA we report the mean and standard deviation over three seeds.
Figure 5: Training loss of the default RoSA run on Llama-2-7B (two cycles, 6 epochs, 79.61% average accuracy). Each step-down coincides with an epoch boundary, where the active group changes; the third marks the return to the first group at the start of the second cycle.
Parameter-efficient fine-tuning(PEFT) has largely focused on LoRA and its accuracy-oriented variants, leaving the original goal of reducing trainable parameters has receivedcomparatively little attention. We introduce FoRA, which revisits this goal by reducing the number of adapted layers rather than adapter rank. FoRA selects task-informative layers via a single-pass diagonal Fisher score (under 1% of training cost) and trains the LoRA down-projection at selected layers on the Stiefel manifold, preserving column orthonormality and effective rank. FoRA consistently outperforms LoRA and DoRA at half their parameter budget, and falls within 0.7-0.8 accuracy points of AdaLoRA at one-quarter its parameter count, across five LLaMA-family backbones. Cross-architecture experiments on twelve backbones from the LLaMA, Qwen3, and Gemma families confirm consistent gains from 270M to 32B parameters. The two components combine super-additively: Fisher selection alone matches rank reduction at the same budget, while the Stiefel constraint provides the decisive additional gain.
Both full fine-tuning (Full FT) and parameter-efficient fine-tuning methods such as LoRA introduce weight updates without accounting for the spectral structure established during pretraining. As a result, noisy gradients from limited fine-tuning data can perturb robust pretrained features. We identify spectral preconditioning as the missing ingredient: reparameterizing each weight matrix through its full-rank singular value decomposition (SVD) and freezing one singular basis constrains updates to the pretrained column space, yielding a preconditioned optimization scheme that outperforms unconstrained Full FT at the same trainable parameter count. Building on this insight, we propose FuRA (Full-Rank Adaptation), an efficient full-rank adaptation framework based on a block tensor-train factorization W = LSR, where the large core L is fixed to the pretrained block-wise SVD basis, while only the compact core R and the block-wise singular values S are optimized. This design simultaneously provides full-rank spectral preconditioning, preserves full-rank update expressivity, and achieves parameter, memory, and step-time efficiency comparable to LoRA. FuRA consistently outperforms Full FT across multiple settings, including LLM fine-tuning (+1.37 on LLaMA-3-8B commonsense reasoning), LLM reinforcement learning for mathematical reasoning, and visual instruction tuning for VLMs. Furthermore, the 4-bit quantized variant, QFuRA, also surpasses QLoRA. Code is available at https://github.com/olokevin/FuRA-NIPS
Yequan Zhao, Ruijie Zhang, Liyan Tan +3
University of California at Santa Barbara · Amazon Lab126
As the scale of large pre-trained models continues to grow, fine-tuning them under limited memory budgets has become increasingly challenging. Low-Rank Adaptation (LoRA), currently one of the most widely adopted parameter-efficient fine-tuning (PEFT) methods, mitigates this challenge by optimizing only low-rank adaptation matrices, thereby greatly reducing the number of trainable parameters. With the parameter overhead substantially reduced, the activations retained for backpropagation have emerged as the primary remaining memory bottleneck during LoRA fine-tuning. To address this, we propose CARE-LoRA, a data-aware Compressed Activation REconstruction framework. By exploiting the inherent projection structure of LoRA, CARE-LoRA replaces the full input activation with the low-rank compressed activation naturally produced by the LoRA branch. It further computes a lightweight reconstruction matrix during the forward pass with negligible additional computation cost, which is used during backpropagation to reconstruct the gradient signal, thereby keeping LoRA matrices fully trainable. Extensive experiments across diverse models and downstream tasks demonstrate that, while substantially reducing the overall memory footprint, CARE-LoRA achieves competitive or even superior performance compared with standard LoRA and representative LoRA variants. Our code is publicly available at https://github.com/fishandyu/CARE-LoRA .
Gengyu Zhang, Haiyin Ran, Zhengbao He +4
Institute of Image Processing and Pattern Recognition, Shanghai Jiao Tong University, Shanghai 200240, China