Are Parameter-Efficient Fine-tuning Methods Really Different?
Organizations: Department of Statistics, The University of Warwick, UK. · Key Laboratory of Interdisciplinary Research of Computation and Economics, Shanghai University of Finance and Economics, China. · School of Mathematical Sciences, Institute of Natural Sciences and MOE-LSC, Shanghai Jiao Tong University, China.
Abstract
Parameter-efficient fine-tuning (PEFT) offers many parameterizations, yet their methodological and functional differences remain unclear. We compare six methods in language and diffusion models to examine how their parameterizations relate to task performance, forgetting, and changes in pretrained weight geometry. Motivated by the spectrum-preserving design of orthogonal fine-tuning (OFT), we first ask whether spectral preservation is itself important for adaptation and retention. We find that the selected LoRA-family methods also approximately preserve pretrained geometry, and that restoring their slightly drifted singular-value spectra largely preserves task performance, questioning the necessity of explicit geometric preservation. Beyond this, we observe that some methods exhibit distinct adaptation--retention trade-offs that vary across settings: LoRA most consistently limits forgetting at competitive performance, DoRA achieves higher mean task scores than LoRA in most comparisons, while PiSSA often incurs greater retention costs. Further intervention experiments suggest that while performance gains from different PEFT methods can be attributed to modifications in different groups of spectral components, we consistently find that restoring dominant rather than intermediate or trailing components produces the largest mean reduction in general-text NLL or base-image drift. Together, these results motivate evaluating geometric constraints through their functional consequences rather than preservation alone. Code is available at https://github.com/Kuaaannn/PEFT_methods.
Figures & tables
| Method | Adapted matrix | Initialization or structure |
|---|---|---|
| LoRA | Zero initial update | |
| DoRA | Learned magnitudes and directions | |
| PiSSA | Leading singular components | |
| MiLoRA | Trailing singular components | |
| OFT | ||
| HRA |
| Qwen2.5-7B | Llama-3.1-8B | |||||
|---|---|---|---|---|---|---|
| Method | S | M | L | S | M | L |
| LoRA | ||||||
| OFT | ||||||
| DoRA | ||||||
| PiSSA | ||||||
| MiLoRA | ||||||
| Budget | Metric | LoRA | OFT | DoRA | PiSSA | MiLoRA | HRA |
|---|---|---|---|---|---|---|---|
| r4/b32 | DINO | ||||||
| Drift | |||||||
| CLIP | |||||||
| r8/b64 | DINO | ||||||
| Drift | |||||||
| CLIP |
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Backbone | Adaptation | Primary retention view |
|---|---|---|
| Qwen2.5-7B | MetaMath / six methods | General-text NLL |
| Llama-3.1-8B | MetaMath / six methods | General-text NLL |
| Qwen2.5-7B | CodeFeedback / six methods | General-text NLL |
| Llama-3.1-8B | CodeFeedback / six methods | General-text NLL |
| FLUX-4B | Cat / six methods | Image drift / CLIP |
| FLUX-4B | 28 objects | Flow loss / drift / CLIP |
| Adapter | Params (M) | LR | Dev | Test | NLL |
|---|---|---|---|---|---|
| LoRA r7 | 17.66 | ||||
| LoRA r14 | 35.32 | ||||
| LoRA r28 | 70.65 | ||||
| OFT b32 | 17.55 | ||||
| OFT b64 | 35.68 | ||||
| OFT b128 | 71.92 |
| Adapter | Params (M) | LR | Dev | Test | NLL |
|---|---|---|---|---|---|
| LoRA r7 | 18.35 | ||||
| LoRA r15 | 39.32 | ||||
| LoRA r30 | 78.64 | ||||
| OFT b32 | 19.30 | ||||
| OFT b64 | 39.22 | ||||
| OFT b128 | 79.07 |
| Adapter | LR | Dev (%) | Test (%) | NLL | |
|---|---|---|---|---|---|
| HRA r16 | 3 | ||||
| HRA r32 | 3 | ||||
| HRA r64 | 3 |
| Model | Method | Selected LR | HumanEval | HE+ | MBPP | MBPP+ |
|---|---|---|---|---|---|---|
| Qwen7B | Base | — | ||||
| LoRA | ||||||
| OFT | ||||||
| DoRA | ||||||
| PiSSA | ||||||
| MiLoRA |
| Method | Capacity | Params | LR | Validation | Test DINO | Gen. drift | Gen. CLIP |
|---|---|---|---|---|---|---|---|
| LoRA | r4 | 4.79 | |||||
| OFT | b32 | 4.76 | |||||
| DoRA | r4 | 5.68 | |||||
| PiSSA | r4 | 4.79 | |||||
| MiLoRA | r4 | 4.79 | |||||
| HRA | h16 | 4.92 |
| Original selection | LR chosen on other seeds | |||
|---|---|---|---|---|
| Method | Lowest forgetting | Average rank | Lowest forgetting | Average rank |
| LoRA | 9/12 | 1.58 | 8/12 | 2.08 |
| OFT | 1/12 | 4.58 | 1/12 | 4.33 |
| DoRA | 0/12 | 2.75 | 1/12 | 2.50 |
| PiSSA | 0/12 | 5.25 | 0/12 | 5.42 |
| MiLoRA | 1/12 | 3.42 | 1/12 | 3.25 |
| Original selection | LR chosen on other seeds | |||
|---|---|---|---|---|
| Check | Frequency | Average rank | Frequency | Average rank |
| Equal model/task weights | 1/1 | 1/1 | 1/1 | 1/1 |
| Remove one backbone | 3/3 | 3/3 | 3/3 | 2/3 |
| Remove one setting | 12/12 | 12/12 | 12/12 | 12/12 |
| Use individual seeds | 9/9 | 9/9 | 9/9 | 6/9 |
| Method | Capacity | LR | |||
|---|---|---|---|---|---|
| LoRA | 7 | ||||
| LoRA | 14 | ||||
| LoRA | 28 | ||||
| OFT | 32 | ||||
| OFT | 64 | ||||
| OFT | 128 |
| Method | Capacity | LR | Trained acc. | Rebuilt acc. | Restored acc. | NLL vs. rebuilt |
|---|---|---|---|---|---|---|
| LoRA | 7 | |||||
| LoRA | 14 | |||||
| LoRA | 28 | |||||
| OFT | 32 | |||||
| OFT | 64 | |||||
| OFT | 128 |
| Method | Capacity | LR | Top acc. | Middle acc. | Tail acc. | Top NLL |
|---|---|---|---|---|---|---|
| LoRA | 7 | |||||
| LoRA | 14 | |||||
| LoRA | 28 | |||||
| OFT | 32 | |||||
| OFT | 64 | |||||
| OFT | 128 |
| Method | Capacity | LR | |||
|---|---|---|---|---|---|
| LoRA | 7 | ||||
| LoRA | 15 | ||||
| LoRA | 30 | ||||
| OFT | 32 | ||||
| OFT | 64 | ||||
| OFT | 128 |
| Method | Capacity | LR | Trained acc. | Rebuilt acc. | Restored acc. | NLL vs. rebuilt |
|---|---|---|---|---|---|---|
| LoRA | 7 | |||||
| LoRA | 15 | |||||
| LoRA | 30 | |||||
| OFT | 32 | |||||
| OFT | 64 | |||||
| OFT | 128 |
| Method | Capacity | LR | Top acc. | Middle acc. | Tail acc. | Top NLL |
|---|---|---|---|---|---|---|
| LoRA | 7 | |||||
| LoRA | 15 | |||||
| LoRA | 30 | |||||
| OFT | 32 | |||||
| OFT | 64 | |||||
| OFT | 128 |
| Method | Qwen2.5-7B | Llama-3.1-8B | FLUX-4B cat |
|---|---|---|---|
| (a) Absolute relative HE change ( ) | |||
| LoRA | |||
| OFT | |||
| DoRA | |||
| PiSSA | |||
| MiLoRA | |||
| Model | Method | Small | Medium | Large | |||
|---|---|---|---|---|---|---|---|
| Top | Complement | Top | Complement | Top | Complement | ||
| Qwen2.5-7B | LoRA | 3/3 | 0/3 | 0/3 | 0/3 | 3/3 | 3/3 |
| OFT | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | |
| DoRA | 3/3 | 3/3 | 3/3 | 1/3 | 3/3 | 3/3 | |
| PiSSA | 3/3 | 0/3 | 3/3 | 0/3 | 3/3 | 0/3 | |
| MiLoRA | 0/3 | 0/3 | 0/3 | 3/3 | 0/3 | 3/3 | |
| Model | Checkpoints | Max. | Max. entry | Max. polar residual |
|---|---|---|---|---|
| FLUX-4B | 33 | |||
| Llama-8B | 21 | |||
| Qwen-7B | 21 |
| Method | Capacity | LR | Params. (M) | |||
|---|---|---|---|---|---|---|
| LoRA | 32 | |||||
| OFT | 256 | |||||
| DoRA | 32 | |||||
| PiSSA | 32 | |||||
| MiLoRA | 32 | |||||
| HRA | 124 |
| Method | Trained | Reconstructed | Restored | Values only | vs. recon. |
|---|---|---|---|---|---|
| LoRA | |||||
| OFT | |||||
| DoRA | |||||
| PiSSA | |||||
| MiLoRA | |||||
| HRA |
| Control | Required evidence |
|---|---|
| Training runs and uncertainty | At least two independent training runs per method with a defined SD, SE or CI attributable to training variation. |
| LR selection | At least three learning rates evaluated per method, with separate selection using a stated validation criterion. |
| LR candidates | Numerical candidate values or a search range stated for both methods. |
| Comparable conditions | Shared checkpoint and data, comparable exposure, and adapted modules and trainable budgets matched or addressed through an appropriate controlled comparison. |
| Control | Documented | Explicit departure | Unclear | All main comparisons |
|---|---|---|---|---|
| Training runs and uncertainty | 9 | 5 | 50 | 2 |
| LR selection | 7 | 4 | 53 | 1 |
| LR candidates | 13 | 0 | 51 | 4 |
| Comparable conditions | 25 | 10 | 29 | 1 |
| Domain | Training uncertainty | LR selection | LR candidates | Conditions |
|---|---|---|---|---|
| Decoder language | 4/51 | 4/51 | 7/51 | 18/51 |
| Other language | 5/40 | 3/40 | 6/40 | 7/40 |
| Vision and multimodal | 3/20 | 1/20 | 3/20 | 2/20 |
| Diffusion | 0/11 | 0/11 | 1/11 | 3/11 |
| Other | 0/1 | 0/1 | 0/1 | 0/1 |
| Decoder subset | 3/40 | 2/40 | 4/40 | 12/40 |