Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability Retention
Organizations: School of Mathematics and Statistics, Xi’an Jiaotong University, Xi’an, Shaanxi, China
Abstract
Parameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining. Existing approaches primarily mitigate this trade-off through data replay or regularization, relying on additional data or explicit optimization constraints. We instead focus on a different question: where should adaptation be applied? We find that fine-tuning different Transformer layers produces different target-task gains and degrees of capability degradation, suggesting that not all layers are equally suitable for adaptation. To characterize this difference, we use layer-wise empirical Fisher information to measure target-task sensitivity. However, computing Fisher scores requires backward computation and becomes increasingly expensive for large models. We therefore introduce input--output cosine similarity as a lightweight, forward-only proxy for ranking layer sensitivity. Across models and tasks, layers with lower input--output similarity consistently exhibit higher empirical Fisher scores. Building on this observation, we propose Layer-Selective LoRA (LS-LoRA), which places trainable LoRA adapters only in layers with low input--output similarity. Experiments on mathematical reasoning and code generation show that LS-LoRA improves average target-task performance while retaining substantially more commonsense reasoning capability than standard all-layer LoRA, demonstrating that carefully choosing where to adapt can provide a simple and effective way to balance target-task adaptation and general capability retention.
Figures & tables
| Task | Llama-3.2-1B | Llama-3.2-3B | Llama-3-8B | Mistral-7B |
|---|---|---|---|---|
| Mathematical Reasoning | -0.82 (1e-8) | -0.84 (2e-9) | -0.89 (6e-12) | -0.85 (7e-10) |
| Code Generation | -0.76 (5e-7) | -0.82 (8e-9) | -0.86 (3e-10) | -0.91 (9e-13) |
| Model | Method | Target-Task Performance (Math) | General Capability Retention | ||||||
|---|---|---|---|---|---|---|---|---|---|
| GSM8K | AQuA | SVAMP | Avg. | ARC-C | OB-QA | SI-QA | Avg. | ||
| Llama-3.2-1B | Base | 31.0 | 6.7 | 60.1 | 32.6 | 28.8 | 30.6 | 40.8 | 33.4 |
| LoRA | 49.5 | 38.2 | 67.6 | 51.8 | 22.8 | 24.6 | 29.9 | 25.8 | |
| L2-SP-LoRA | 42.8 | 25.6 | 70.3 | 46.2 | 26.1 | 29.2 | 34.0 | 29.8 | |
| EWC-LoRA | 50.3 | 36.6 | 68.6 | 51.9 | 22.8 | 27.4 | 32.3 | 27.5 | |
| LS-LoRA | 50.3 | 37.8 | 68.4 | 52.2 | 30.0 | 32.2 | 33.1 | 31.8 | |
| Model | Method | Target-Task Performance (Code) | General Capability Retention | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Pass@1 | Pass@5 | Pass@10 | Avg. | ARC-C | OB-QA | SI-QA | Avg. | ||
| Llama-3.2-1B | Base | 33.0 | 43.3 | 47.2 | 41.2 | 28.8 | 30.6 | 40.8 | 33.4 |
| LoRA | 35.6 | 44.7 | 47.8 | 42.7 | 26.1 | 28.6 | 39.1 | 31.3 | |
| L2-SP-LoRA | 32.9 | 42.7 | 45.7 | 40.4 | 26.4 | 30.4 | 39.9 | 32.2 | |
| EWC-LoRA | 34.3 | 42.8 | 46.6 | 41.3 | 26.5 | 31.0 | 39.2 | 32.2 | |
| LS-LoRA | 35.8 | 45.6 | 49.2 | 43.5 | 26.9 | 30.6 | 39.9 | 32.5 | |
| Model | Method | Target-Task Performance (Math) | General Capability Retention | ||||||
|---|---|---|---|---|---|---|---|---|---|
| GSM8K | AQuA | SVAMP | Avg. | ARC-C | OB-QA | SI-QA | Avg. | ||
| Llama-3.2-1B | Base | 31.0 | 6.7 | 60.1 | 32.6 | 28.8 | 30.6 | 40.8 | 33.4 |
| LoRA | 49.5 | 38.2 | 67.6 | 51.8 | 22.8 | 24.6 | 29.9 | 25.8 | |
| Rand-LoRA | 50.8 | 36.6 | 66.8 | 51.4 | 25.3 | 28.0 | 28.1 | 27.1 | |
| LS-LoRA | 50.3 | 37.8 | 68.4 | 52.2 | 30.0 | 32.2 | 33.1 | 31.8 | |
| Metric | Llama-3.2-1B | Llama-3.2-3B | Llama-3-8B | Mistral-7B | ||||
|---|---|---|---|---|---|---|---|---|
| LoRA | LS-LoRA | LoRA | LS-LoRA | LoRA | LS-LoRA | LoRA | LS-LoRA | |
| GPU Memory (GB) | 7.2 | 6.9 | 13.1 | 11.8 | 24.8 | 22.5 | 21.5 | 18.4 |
| Training Time (h) | 1.0 | 0.9 | 1.7 | 1.2 | 2.1 | 1.8 | 2.7 | 1.9 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Method | Target-Task Performance (Math) | General Capability Retention | ||||||
|---|---|---|---|---|---|---|---|---|---|
| GSM8K | AQuA | SVAMP | Avg. | ARC-C | OB-QA | SI-QA | Avg. | ||
| Llama-3.2-3B | Base | 62.0 | 19.3 | 83.7 | 55.0 | 62.2 | 56.4 | 58.8 | 59.1 |
| LoRA | 69.6 | 47.6 | 80.3 | 65.8 | 44.1 | 52.4 | 34.9 | 43.8 | |
| Rand-LoRA | 69.6 | 46.1 | 80.2 | 65.3 | 57.3 | 59.0 | 53.8 | 56.7 | |
| LS-LoRA | 69.5 | 48.8 | 82.3 | 66.9 | 59.3 | 62.4 | 52.6 | 58.1 | |
| Llama-3-8B | Base | 59.6 | 49.6 | 53.8 | 54.3 | 33.1 | 25.6 | 37.4 | 32.0 |
| Method | Target-Task Performance (Math) | General Capability Retention | ||||||
|---|---|---|---|---|---|---|---|---|
| GSM8K | AQuA | SVAMP | Avg. | ARC-C | OB-QA | SI-QA | Avg. | |
| DoRA | 68.1 | 48.8 | 81.2 | 66.0 | 36.4 | 43.8 | 33.4 | 37.9 |
| LS-DoRA | 69.5 | 49.2 | 82.2 | 67.0 | 60.3 | 62.8 | 53.7 | 58.9 |
| PiSSA | 70.6 | 48.0 | 79.9 | 66.2 | 59.2 | 61.8 | 55.9 | 59.0 |
| LS-PiSSA | 71.3 | 50.0 | 81.4 | 67.6 | 68.6 | 66.0 | 59.5 | 64.7 |