Parameter-efficient fine-tuning (PEFT) reduces the cost of adapting foundation models by focusing training on a small parameter subset. Complementary to this idea, we introduce RoSA (Rotational Sparse Adaptation), which narrows adaptation to a subset of layers at a time. RoSA freezes lower layers close to the input throughout training and rotates a trainable block over later layers, progressively increasing the number of frozen layers close to the input. This design reduces optimizer-state memory, shortens backpropagation, and even forward propagation if activations at the last frozen layer are cached. Because RoSA is orthogonal to the choice of trainable parameterization, it can be combined with PEFT methods or sparse optimizers within each active block. Experiments across multiple LLM architectures and tasks show that RoSA reduces peak memory while maintaining strong fine-tuning performance.
Figures & tables
Figure 1: RoSA training schedule. RoSA freezes a large input-side prefix, keeps an output-side anchor active, and rotates one consecutive trainable group through the later layers. Freezing the prefix removes optimizer states and parameter gradients for early layers, shortens the backward path, and enables future activation caching for the fixed prefix.
Table 2
Figure 2: Peak VRAM usage. Full fine-tuning with AdamW and full-model sparse optimization remain above the 40 GB budget, while RoSA fits within it and preserves strong performance.
Table 4Table 5Table 6
Figure 3: Layer-wise gradient norm analysis of Llama-2-7B during full fine-tuning. Left: Step 100, where activity is dominated by the input and head. Middle: Step 1100, where a “reasoning hump” emerges in Layers 8–18 while Layers 2–7 remain flat. Right: the average gradient profile over the first epoch, which highlights inefficient updates in deeper Layers 20–29. RoSA exploits this structure by freezing stable early Layers 0–7 and rotating updates through the more active middle layers.
Method
Ep.
Acc.
Peak
Later
Runtime
(%)
VRAM
VRAM
DoRA
3
81.58
16.42
–
15h 28m
RoSA-DoRA
6
81.79
12.87
11.04/9.44
23h 40m
RoSA-DoRA (3 seeds)
6
81.83 ± 0.32
12.87
–
–
Table 9: RoSA-DoRA on Qwen-2.5-3B commonsense reasoning. RoSA-DoRA maintains similar accuracy while reducing VRAM across rotation stages. The last row reports the mean and standard deviation over three seeds.
Task
Full
Core
Δ
OpenBookQA
79.2
79.4
+0.2
PIQA
82.9
83.1
+0.2
WinoGrande
82.2
82.2
0.0
ARC-Easy
83.8
83.6
-0.2
HellaSwag
89.3
89.0
-0.3
BoolQ
71.0
70.7
-0.3
Table 10: Core-only RoSA ablation. Restricting rotation to Layers 8–23 gives results almost identical to the full RoSA schedule. The average is computed over the full 8-task suite.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Baseline ρ=0.2,Δt=600
Low Freq. ρ=0.1,Δt=600
High Freq. ρ=0.1,Δt=100
ARC-C
69.9
68.3
67.5
ARC-E
83.8
82.8
83.2
BoolQ
71.0
69.7
68.8
HellaS
89.3
74.9
87.9
OBQA
79.2
77.8
78.0
PIQA
82.9
82.5
82.9
Appendix
Table 11: Task accuracy (%) under different sparsity-update settings. Lower density hurts performance when the sparse mask is refreshed slowly, but much of this loss can be recovered by refreshing the mask more often.
Optimizer states (NanoAdam, ∼ 495.5M tracked parameters)
2.52
Activations and other allocations
13.91
Peak allocated tensor memory (PyTorch)
32.62
CUDA context and caching allocator
∼ 2.3
Appendix
Table 12: Breakdown of the peak memory of RoSA + NanoAdam on Llama-2-7B commonsense reasoning ( ρ=0.2 , group size 8, micro-batch size 1, cutoff length 256).
Figure 4: Average per-layer ℓ2 gradient norms from short dense fine-tuning probes. Axis scales differ between panels.
Setting
Run
Avg Acc.
Llama-2-7B, RoSA + NanoAdam
original run (Table 2 )
79.61
seed 44
79.49
Qwen-2.5-3B, RoSA-DoRA
seed 42 (Table 9 )
81.79
seed 43
82.17
seed 44
81.53
mean ± std
81.83±0.32
Appendix
Table 13: Average commonsense reasoning accuracy (%) across random seeds. For RoSA-DoRA we report the mean and standard deviation over three seeds.
Figure 5: Training loss of the default RoSA run on Llama-2-7B (two cycles, 6 epochs, 79.61% average accuracy). Each step-down coincides with an epoch boundary, where the active group changes; the third marks the return to the first group at the start of the second cycle.