SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing
Organizations: North China Electric Power University · Southeast University · Soochow University · Engineering Research Center of Intelligent Computing for Complex Energy Systems, Ministry of Education
Abstract
Fine-tuning-as-a-service enables users to adapt aligned large language models (LLMs) to specialized tasks, but malicious fine-tuning can erode refusal behavior while preserving task performance on legitimate inputs. We revisit recent layer-wise safety diagnostics and find that safety sensitivity is signed: scaling different layers can strengthen refusal, weaken it, or have little effect. Motivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing. SLDR trains a LoRA recovery adapter only on the layers with the maximum and minimum sensitivity scores in the signed spectrum, and uses representation-based dynamic routing inference to activate the adapter only for malicious queries. Across four model architectures, five downstream tasks, and four harmful benchmarks, SLDR substantially reduces harmful outputs while preserving downstream utility. On Llama3.1/SST2, SLDR reduces the average harmful score from 11.54 to 0.08 while maintaining downstream accuracy, and the harmful score remains near zero under poisoning ratios up to 0.9. The code is available at https://github.com/Stardust457/SLDR.
Figures & tables
| Harmful Score | Finetune Accuracy | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | clean | Average | clean | Average | ||||||||
| SFT | 1.60 | 5.50 | 11.40 | 18.30 | 20.90 | 11.54 | 92.43 | 92.20 | 92.43 | 92.32 | 92.55 | 92.39 |
| Lisa | 1.50 | 1.60 | 1.90 | 1.80 | 1.70 | 1.70 | 73.74 | 74.08 | 74.54 | 74.31 | 74.20 | 74.17 |
| Antidote | 1.50 | 1.40 | 1.60 | 1.50 | 1.50 | 1.50 | 84.86 | 84.75 | 84.75 | 85.09 | 84.86 | 84.86 |
| Panacea | 1.20 | 3.60 | 6.40 | 7.00 | 7.80 | 5.20 | 87.84 | 90.02 | 89.79 | 89.56 | 89.22 | 89.29 |
| STAR-DSS | 1.60 | 1.70 | 1.20 | 1.40 | 1.70 | 1.52 | 84.75 | 85.67 | 85.78 | 85.55 | 85.67 | 85.48 |
| Harmful Score | ||||||
|---|---|---|---|---|---|---|
| Method | clean | Average | ||||
| BDS | 1.40 | 1.46 | 1.47 | 1.51 | 1.49 | 1.47 |
| One-shot FT | 1.77 | 1.82 | 1.78 | 1.84 | 1.86 | 1.81 |
| Ours | 1.08 | 1.08 | 1.06 | 1.09 | 1.14 | 1.09 |
| Harmful Score | Finetune Accuracy | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | ||||||||||||||||
| SFT | 29.30 | 30.40 | 34.30 | 34.20 | 36.00 | 36.70 | 35.70 | 36.70 | 92.32 | 92.09 | 91.97 | 91.63 | 91.74 | 91.17 | 90.37 | – |
| BDS | 2.90 | 2.70 | 2.60 | 2.90 | 3.30 | 3.10 | 3.50 | 3.60 | 91.74 | 91.40 | 91.97 | 91.74 | 91.74 | 91.40 | 91.28 | – |
| One-shot FT | 5.00 | 5.90 | 6.50 | 6.30 | 7.30 | 7.10 | 7.40 | 8.00 | 89.79 | 89.56 | 89.68 | 89.22 | 87.96 | 87.96 | 87.61 | – |
| Ours | 0.20 | 0.30 | 0.20 | 0.30 | 0.20 | 0.40 | 0.20 | 0.20 | 92.32 | 92.09 | 91.97 | 91.51 | 91.51 | 90.94 | 90.14 | – |
| Statistic | Top-5 layers (ranked; score) |
|---|---|
| SLDR | 13 (+200), 12 ( 185), 10 (+180), 8 ( 170), 15 (+135) |
| Panacea | 11 (58.13), 9 (49.03), 8 (46.75), 12 (45.06), 17 (44.15) |
| Targeted Vaccine | 0 (0.992), 18 (0.490), 19 (0.484), 1 (0.482), 16 (0.482) |
| Surgery | 0 (+0.154), 2 ( 0.178), 1 ( 0.182), 3 ( 0.371), 4 ( 0.623) |
| Statistic | Top-5 overlap | Top-10 overlap |
|---|---|---|
| Panacea | 2/5 | 4/10 |
| Targeted Vaccine | 0/5 | 4/10 |
| Surgery | 0/5 | 2/10 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Harmful Score | Finetune Accuracy | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Method | Seed 0 | Seed 1 | Seed 2 | Seed 42 | Mean Std | Seed 0 | Seed 1 | Seed 2 | Seed 42 | Mean Std |
| BDS | 1.80 | 2.40 | 1.70 | 1.60 | 90.71 | 91.86 | 91.28 | 91.63 | ||
| One-shot FT | 4.50 | 3.90 | 4.40 | 3.90 | 90.83 | 90.60 | 90.71 | 90.48 | ||
| Ours | 0.10 | 0.00 | 0.20 | 0.00 | 93.00 | 93.12 | 93.12 | 92.43 | ||
| Method | Harmful Score | Finetune Accuracy |
|---|---|---|
| BDS | 47.70 | 88.42 |
| One-shot FT | 51.10 | 93.46 |
| Ours | 1.10 | 92.55 |
| Component | Cost |
|---|---|
| Layer probing across 32 layers | 0.3663 hours |
| Standard SFT inference | 0.3121 seconds per prompt |
| SLDR inference with routing | 0.3293 seconds per prompt |
| Additional online overhead | 0.0172 seconds per prompt |
| Method | Zulu HS | IJP HS |
|---|---|---|
| Original aligned model | 66.82 | 12.94 |
| SFT (no defense) | 66.82 | 18.24 |
| EnchTable | 72.71 | 13.53 |
| Ours | 61.53 | 15.53 |
| Method | BeaverTails HS | SST2 FA |
|---|---|---|
| Original aligned model | 1.50 | 70.64 |
| SFT (no defense) | 11.40 | 92.43 |
| EnchTable | 1.50 | 88.65 |
| Ours | 0.00 | 92.43 |
| Recovery Epochs | HS | FA |
|---|---|---|
| 5 | 0.10 | 92.20 |
| 10 | 0.00 | 92.20 |
| 20 (default) | 0.00 | 92.43 |
| 30 | 0.30 | 92.20 |
| 40 | 0.60 | 92.20 |
| 50 | 0.60 | 92.20 |
| Metric / Threshold | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| TPR | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 99.25 | 94.75 | 89.25 | 76.25 | 59.25 |
| FPR | 100.00 | 63.00 | 15.25 | 2.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| Benign over refusal rate | 14.50 | 14.50 | 7.50 | 1.75 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |