Organizations: North China Electric Power University · Southeast University · Soochow University · Engineering Research Center of Intelligent Computing for Complex Energy Systems, Ministry of Education
Fine-tuning-as-a-service enables users to adapt aligned large language models (LLMs) to specialized tasks, but malicious fine-tuning can erode refusal behavior while preserving task performance on legitimate inputs. We revisit recent layer-wise safety diagnostics and find that safety sensitivity is signed: scaling different layers can strengthen refusal, weaken it, or have little effect. Motivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing. SLDR trains a LoRA recovery adapter only on the layers with the maximum and minimum sensitivity scores in the signed spectrum, and uses representation-based dynamic routing inference to activate the adapter only for malicious queries. Across four model architectures, five downstream tasks, and four harmful benchmarks, SLDR substantially reduces harmful outputs while preserving downstream utility. On Llama3.1/SST2, SLDR reduces the average harmful score from 11.54 to 0.08 while maintaining downstream accuracy, and the harmful score remains near zero under poisoning ratios up to 0.9. The code is available at https://github.com/Stardust457/SLDR.
Figures & tables
Figure 1: Signed layer-sensitivity scores for Llama3.1-8B-Instruct. Positive scores (green) indicate layers whose scaling increases refusal, while negative scores (red) indicate layers whose scaling decreases refusal.
Figure 2: Overview of SLDR. (a) Signed layer selection; (b) two-stage fine-tuning with targeted safety recovery; (c) representation-based dynamic routing for inference-time safety-utility balance.
Harmful Score ↓
Finetune Accuracy ↑
Method
clean
ρ=0.05
ρ=0.1
ρ=0.15
ρ=0.2
Average
clean
ρ=0.05
ρ=0.1
ρ=0.15
ρ=0.2
Average
SFT
1.60
5.50
11.40
18.30
20.90
11.54
92.43
92.20
92.43
92.32
92.55
92.39
Lisa
1.50
1.60
1.90
1.80
1.70
1.70
73.74
74.08
74.54
74.31
74.20
74.17
Antidote
1.50
1.40
1.60
1.50
1.50
1.50
84.86
84.75
84.75
85.09
84.86
84.86
Panacea
1.20
3.60
6.40
7.00
7.80
5.20
87.84
90.02
89.79
89.56
89.22
89.29
STAR-DSS
1.60
1.70
1.20
1.40
1.70
1.52
84.75
85.67
85.78
85.55
85.67
85.48
Table 1: Comparison with SOTA baselines. Using the Llama3.1 and the SST2 dataset.
Table 4
Harmful Score ↓
Method
clean
ρ=0.05
ρ=0.1
ρ=0.15
ρ=0.2
Average
BDS
1.40
1.46
1.47
1.51
1.49
1.47
One-shot FT
1.77
1.82
1.78
1.84
1.86
1.81
Ours
1.08
1.08
1.06
1.09
1.14
1.09
Table 4: Robustness Across Different LLM Judge. Using the Llama3.1 and the SST2 dataset.
Harmful Score ↓
Finetune Accuracy ↑
Method
ρ=0.3
ρ=0.4
ρ=0.5
ρ=0.6
ρ=0.7
ρ=0.8
ρ=0.9
ρ=1.0
ρ=0.3
ρ=0.4
ρ=0.5
ρ=0.6
ρ=0.7
ρ=0.8
ρ=0.9
ρ=1.0
SFT
29.30
30.40
34.30
34.20
36.00
36.70
35.70
36.70
92.32
92.09
91.97
91.63
91.74
91.17
90.37
–
BDS
2.90
2.70
2.60
2.90
3.30
3.10
3.50
3.60
91.74
91.40
91.97
91.74
91.74
91.40
91.28
–
One-shot FT
5.00
5.90
6.50
6.30
7.30
7.10
7.40
8.00
89.79
89.56
89.68
89.22
87.96
87.96
87.61
–
Ours
0.20
0.30
0.20
0.30
0.20
0.40
0.20
0.20
92.32
92.09
91.97
91.51
91.51
90.94
90.14
–
Table 5: Robustness across different poisoning ratios. Using the Llama3.1 and the SST2 dataset.
Table 12: Comparison of layer-wise statistics on Llama3.1-8B-Instruct. Scores are shown for the Top-5 layers. Since the statistics have different scales, scores should not be compared across methods.
Statistic
Top-5 overlap
Top-10 overlap
Panacea
2/5
4/10
Targeted Vaccine
0/5
4/10
Surgery
0/5
2/10
Table 13: Overlap between the Top- K layers identified by SLDR and other layer-wise statistics.
Figure 5: Refusal count results for layer-wise scaling of Llama3.1-8B-Instruct. (a) shows the results for the first 16 layers, and (b) shows the results for the last 16 layers. The x-axis denotes the scaling parameter (1±α) applied to all weight matrices in layer l , where α∈{0.1,0.2} . The value 1.0 corresponds to the unscaled baseline.
Figure 6: Refusal count results for layer-wise scaling of Llama3-8B-Instruct. (a) shows the results for the first 16 layers, and (b) shows the results for the last 16 layers. The x-axis denotes the scaling parameter (1±α) applied to all weight matrices in layer l , where α∈{0.1,0.2} . The value 1.0 corresponds to the unscaled baseline.
Figure 7: Refusal count results for layer-wise scaling of Qwen2.5-7B-Instruct. (a) shows the results for the first 14 layers, and (b) shows the results for the last 14 layers. The x-axis denotes the scaling parameter (1±α) applied to all weight matrices in layer l , where α∈{0.1,0.2} . The value 1.0 corresponds to the unscaled baseline.
Figure 8: Refusal count results for layer-wise scaling of Mistral-7B-Instruct-v0.2. (a) shows the results for the first 16 layers, and (b) shows the results for the last 16 layers. The x-axis denotes the scaling parameter (1±α) applied to all weight matrices in layer l , where α∈{0.1,0.2} . The value 1.0 corresponds to the unscaled baseline.
Figure 9: Layer-wise sensitivity scores of Llama3-8B-Instruct. Positive scores (green) indicate layers that enhance safety refusal, while negative scores (red) indicate layers that undermine it.
Figure 10: Layer-wise sensitivity scores of Qwen2.5-7B-Instruct. Positive scores (green) indicate layers that enhance safety refusal, while negative scores (red) indicate layers that undermine it.
Figure 11: Layer-wise sensitivity scores of Mistral-7B-Instruct-v0.2. Positive scores (green) indicate layers that enhance safety refusal, while negative scores (red) indicate layers that undermine it.
Harmful Score ↓
Finetune Accuracy ↑
Method
Seed 0
Seed 1
Seed 2
Seed 42
Mean ± Std
Seed 0
Seed 1
Seed 2
Seed 42
Mean ± Std
BDS
1.80
2.40
1.70
1.60
1.88±0.36
90.71
91.86
91.28
91.63
91.37±0.50
One-shot FT
4.50
3.90
4.40
3.90
4.18±0.32
90.83
90.60
90.71
90.48
90.66±0.15
Ours
0.10
0.00
0.20
0.00
0.08±0.10
93.00
93.12
93.12
92.43
92.92±0.33
Appendix
Table 14: Multi-seed stability analysis on Llama3.1 with the SST2 dataset.
Method
Harmful Score ↓
Finetune Accuracy ↑
BDS
47.70
88.42
One-shot FT
51.10
93.46
Ours
1.10
92.55
Appendix
Table 15: Performance under full-parameter downstream fine-tuning. Using the Llama3.1 and the SST2 dataset.
Component
Cost
Layer probing across 32 layers
0.3663 hours
Standard SFT inference
0.3121 seconds per prompt
SLDR inference with routing
0.3293 seconds per prompt
Additional online overhead
0.0172 seconds per prompt
Appendix
Table 16: Computational overhead of SLDR on Llama3.1. Inference time is averaged over 1,000 BeaverTails prompts with a generation length of 64 tokens.
Method
Zulu HS ↓
IJP HS ↓
Original aligned model
66.82
12.94
SFT (no defense)
66.82
18.24
EnchTable
72.71
13.53
Ours
61.53
15.53
Appendix
Table 17: Comparison under jailbreak distribution shifts on Llama3.1 with the SST2 dataset. Lower Harmful Score (HS) indicates better safety.
Method
BeaverTails HS ↓
SST2 FA ↑
Original aligned model
1.50
70.64
SFT (no defense)
11.40
92.43
EnchTable
1.50
88.65
Ours
0.00
92.43
Appendix
Table 18: Safety and downstream utility comparison on Llama3.1 with the SST2 dataset. HS is evaluated on BeaverTails.
Recovery Epochs
HS ↓
FA ↑
5
0.10
92.20
10
0.00
92.20
20 (default)
0.00
92.43
30
0.30
92.20
40
0.60
92.20
50
0.60
92.20
Appendix
Table 19: Impact of the number of recovery epochs on SLDR. Experiments are conducted on Llama3.1 with the SST2 dataset under ρ=0.1 and ∣Dtrain∣=1000 .
Metric / Threshold
−0.10
−0.08
−0.06
−0.04
−0.02
0.00
0.02
0.04
0.06
0.08
0.10
TPR ↑
100.00
100.00
100.00
100.00
100.00
100.00
99.25
94.75
89.25
76.25
59.25
FPR ↓
100.00
63.00
15.25
2.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
Benign over refusal rate ↓
14.50
14.50
7.50
1.75
0.00
0.00
0.00
0.00
0.00
0.00
0.00
Appendix
Table 20: Router performance under different routing thresholds τ .
Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification, which can be costly and may hurt task performance, or they use model-agnostic safety classifiers, which may miss failures specific to a given fine-tuned checkpoint. These limitations motivate a post hoc, model-specific, and non-invasive approach to safety restoration. To meet these requirements, we propose HyperSafe, a framework that restores safety behavior by generating a model-specific Safe Side Network (SSN) for each fine-tuned checkpoint. HyperSafe uses layer-wise activation fingerprints to capture how fine-tuning changes the model's inner representations. With a small set of given calibration prompts, the hypernetwork maps these fingerprints to the parameters of the \ssn{} in a single forward pass. The generated \ssn{} runs alongside the frozen fine-tuned model and performs prompt-level safety classification: harmful prompts are routed to refusal, while safe prompts are answered by the original fine-tuned model. Thus, HyperSafe requires no gradient updates, no safety data at deployment time, and no modification to the deployed model weights. We evaluate HyperSafe on two model families, Qwen2-7B and LLaMA-3-8B, across multiple safety benchmarks. HyperSafe reduces harmful response rates from 19-31% to below 1% on every held-out checkpoint, while keeping downstream task accuracy within 1% of the fine-tuned baseline on average. Code is available at https://github.com/nokronim/project-safety-remedy.
Aznaur Aliev, Carlos Hinojosa, Abdelrahman Eldesokey +3
King Abdullah University of Science and Technology, Saudi Arabia
Fine-tuning-as-a-Service (FaaS) enables personalization of large language models (LLMs), but it can weaken safety-alignment under harmful fine-tuning attacks. Recent work has shown that activating harmful-behavior modules during fine-tuning can prevent models from learning undesired behaviors, but its mechanism remains unclear. In this paper, we revisit temporary jailbreaking as a defense against harmful fine-tuning and provide a gradient-level analysis showing that it saturates safety-degrading gradients while preserving benign task-relevant gradients. Based on this insight, we propose a Buffer-and-Reinforce fine-tuning framework that buffers harmful updates during user fine-tuning and reinforces safety after adaptation. Specifically, BufferLoRA induces temporary jailbreaking as a removable adapter to reduce harmful updates during user fine-tuning. After adaptation, ReinforceLoRA, trained to recover refusal behavior under the temporarily jailbroken state, is integrated with UserLoRA via QR decomposition-based merging to reinforce safety while preserving user-task performance. Extensive experiments show that our framework achieves superior safety and utility with no additional safety data during user fine-tuning and minimal computational cost.
Seokil Ham, Jaehyuk Jang, Wonjun Lee +1
School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Republic of Korea.
Benign fine-tuning severely weakens the safety alignment of large language models (LLMs), so we study why refusal behavior is so fragile. While prior work often attributes this failure to gradient conflict, we propose a fundamentally different Fisher-geometric explanation: safety Fisher is low-rank, and alignment makes the safety geometry flatter while preserving an output-routing pathway. After 100 benign fine-tuning examples, this pathway is selectively re-sharpened in output-side MLP modules, explaining the asymmetric fragility: safety can collapse to high attack success rates, while general utility degrades mildly. The routing view also explains why few safety examples can restore refusal behavior, indicating that internal safety-relevant representations are preserved. Finally, we show that LoRA and ASAM mitigate early collapse by suppressing output-side sharpness, but their protection weakens at larger fine-tuning scales. Overall, safety failure is best understood as a disruption of a low-rank output-routing mechanism
Yitong Guo, Xiaoyi Chen, Siyuan Zhang +2
Indiana University Bloomington · Tsinghua University · Nanyang Technological University