Fine-Tuning-as-a-Service (FTaaS) platforms let users perform supervised fine-tuning (SFT) on customized data, but this pipeline can erode model safety alignment. To recover safety without re-running full alignment, existing realignment methods focus on calibrating the integration of safety patches into fine-tuned models. These methods exhibit a persistent safety-utility trade-off: weak repair leaves harmful behavior intact, while stronger repair increasingly damages the benign task. This paper shifts the focus from online calibration to offline patch learning and aims to learn a safety patch that restores safety while preserving task-specific capabilities. To this end, we propose TRACE, which simulates harmful SFT trajectories to produce progressively corrupted model states, and optimizes a safety patch simultaneously across these states. TRACE trains the safety patch during the offline stage, and reuses it across all user checkpoints without per-user calibration. We evaluate two representative models using three harmful SFT datasets, together with three utility benchmarks. Across six benchmarks and two models, TRACE consistently dominates the safety-utility frontier. TRACE improves the safety rate by up to 77 percentage points over the best baselines, while maintaining comparable utility to the undefended fine-tuned model.
Figures & tables
Paradigm
Methods
Safety
Utility
Robustness
Efficiency
Extensibility
Stage
Mechanism
Pre-FT
Alignment Hardening
Vaccine ( Huang et al., 2024 ) , Booster ( Huang et al., 2025 )
0
In-FT
Safety-preserving FT
SaLoRA ( Li et al., 2025 ) , SPF ( Zhang et al., 2026b )
∼103
Post-FT
Safety Re-training
OneShot ( Zhang et al., 2026a )
∼101
Online Merge Calibration
RESTA ( Bhardwaj et al., 2024 ) , EnchTable ( Wu et al., 2025 ) SafeLoRA ( Hsu et al., 2024 ) , SafeDelta ( Lu et al., 2025 )
101∼102
Offline Patch Learning
TRACE (Ours)
∼0.5
Table 1: Comparison of intervention paradigms for post-training safety recovery.
Figure 1: Overview of TRACE. Offline (left), harmful SFT trajectory simulation alternates with safety patch optimization. Online (right), the provider applies the learned patch directly to a fine-tuned model built from the same backbone, without per-user calibration.
Safety Rate (%) ↑
Pure Profile
Mixed Profile
Llama
Qwen
Llama
Qwen
Dataset
Shadow
PureBad
SafeRLHF
Shadow
PureBad
SafeRLHF
MixedSamSum
MixedSQL
MixedGSM8K
MixedSamSum
MixedSQL
MixedGSM8K
No Defense
18
10
11
3
5
2
11
20
12
2
4
1
RESTA
99
15
8
10
3
3
13
21
21
30
4
50
EnchTable
100
16
10
10
21
4
14
21
22
24
5
44
Table 2: Safety rate and task accuracy under pure and mixed profiles. In each column, bold and underline indicate the best and second-best safety rate, respectively. For TRACE, ↑ denotes the gain over the second-best value, while Δ denotes the change in task accuracy relative to No Defense.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
SPF
OneShot
RESTA
EnchTable
SafeLoRA
SafeDelta
TRACE
Offline preparation (s)
/
/
140.25
140.25
217.66
216.61
1244.86
Per-user defense (s)
2343.10
14.99
6.71
430.29
123.08
26.86
0.41 ↓ 93.9%
Appendix
Table 3: Method-specific computation for one-time offline preparation and per-user defense (seconds). Common checkpoint loading, serialization, and API serving costs are excluded.
Shadow
PureBad
SafeRLHF
Method
Harmfulness ↓
Refusal ↑
Harmfulness ↓
Refusal ↑
Harmfulness ↓
Refusal ↑
No Defense
0.33
0.23
0.17
0.58
0.26
0.43
TRACE
0.00
1.00
0.02
0.96
0.03
0.94
Appendix
Table 4: Safety evaluation using the original StrongReject rubric.
Metric(%)
Base
TRACE
Δ (pp)
Overall
68.30±0.37
67.97
−0.33
Humanities
64.21
64.23
+0.02
Other
72.96
72.71
−0.25
Social Sciences
77.74
77.51
−0.23
STEM
60.61
59.56
−1.05
Appendix
Table 5: MMLU utility evaluation before and after applying the TRACE safety patch.
Safety (↑)
Utility (↑)
Ablation
Score (%)
Δ (pp)
Score(%)
Δ (pp)
No Defense
9.58
/
68.30
/
Base-only
15.97
+6.39
65.93
-2.37
Endpoint-only
9.00
-0.58
66.35
-1.95
No task preservation
5.11
-4.47
24.73
-43.57
Harmful-only trajectory
2.88
-6.70
24.73
-43.57
Appendix
Table 6: Ablation of the TRACE training design. Δ denotes the deviation relative to No Defense.
Task Accuracy (%) ↑
Pure Profile
Mixed Profile
Llama
Qwen
Llama
Qwen
Method
SamSum
SQL
GSM8K
SamSum
SQL
GSM8K
MixedSamSum
MixedSQL
MixedGSM8K
MixedSamSum
MixedSQL
MixedGSM8K
No Defense
51
81
55
46
66
42
47
76
40
50
77
59
RESTA
49
82
58
46
65
44
48
76
43
49
78
63
EnchTable
50
81
58
46
65
44
48
77
43
50
78
64
Appendix
Table 7: Task accuracy (%) for all methods under pure and mixed profiles. For TRACE, Δ denotes its performance deviation (pp) relative to No Defense.
Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification, which can be costly and may hurt task performance, or they use model-agnostic safety classifiers, which may miss failures specific to a given fine-tuned checkpoint. These limitations motivate a post hoc, model-specific, and non-invasive approach to safety restoration. To meet these requirements, we propose HyperSafe, a framework that restores safety behavior by generating a model-specific Safe Side Network (SSN) for each fine-tuned checkpoint. HyperSafe uses layer-wise activation fingerprints to capture how fine-tuning changes the model's inner representations. With a small set of given calibration prompts, the hypernetwork maps these fingerprints to the parameters of the \ssn{} in a single forward pass. The generated \ssn{} runs alongside the frozen fine-tuned model and performs prompt-level safety classification: harmful prompts are routed to refusal, while safe prompts are answered by the original fine-tuned model. Thus, HyperSafe requires no gradient updates, no safety data at deployment time, and no modification to the deployed model weights. We evaluate HyperSafe on two model families, Qwen2-7B and LLaMA-3-8B, across multiple safety benchmarks. HyperSafe reduces harmful response rates from 19-31% to below 1% on every held-out checkpoint, while keeping downstream task accuracy within 1% of the fine-tuned baseline on average. Code is available at https://github.com/nokronim/project-safety-remedy.
Aznaur Aliev, Carlos Hinojosa, Abdelrahman Eldesokey +3
King Abdullah University of Science and Technology, Saudi Arabia
Fine-tuning-as-a-Service (FaaS) enables personalization of large language models (LLMs), but it can weaken safety-alignment under harmful fine-tuning attacks. Recent work has shown that activating harmful-behavior modules during fine-tuning can prevent models from learning undesired behaviors, but its mechanism remains unclear. In this paper, we revisit temporary jailbreaking as a defense against harmful fine-tuning and provide a gradient-level analysis showing that it saturates safety-degrading gradients while preserving benign task-relevant gradients. Based on this insight, we propose a Buffer-and-Reinforce fine-tuning framework that buffers harmful updates during user fine-tuning and reinforces safety after adaptation. Specifically, BufferLoRA induces temporary jailbreaking as a removable adapter to reduce harmful updates during user fine-tuning. After adaptation, ReinforceLoRA, trained to recover refusal behavior under the temporarily jailbroken state, is integrated with UserLoRA via QR decomposition-based merging to reinforce safety while preserving user-task performance. Extensive experiments show that our framework achieves superior safety and utility with no additional safety data during user fine-tuning and minimal computational cost.
Seokil Ham, Jaehyuk Jang, Wonjun Lee +1
School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Republic of Korea.
Fine-tuning well-aligned large language models (LLMs) on new domains often degrades their safety alignment, even when using benign datasets. Existing safety alignment techniques primarily focus on pretraining, leaving fine-tuned models vulnerable to behavioral shifts. In this work, we introduce safety token regularization (STR), a lightweight method designed to preserve safety properties during fine-tuning. Our approach identifies salient tokens from rejection templates of well-aligned models and constrains their associated logits during training, preventing the loss of critical safety behaviors. Unlike reinforcement learning or preference optimization methods, STR requires minimal additional computation and seamlessly integrates with parameter-efficient fine-tuning techniques such as LoRA. Comprehensive experiments demonstrate that our approach achieves safety performance on par with state-of-the-art methods, while preserving task-specific utility and requiring minimal implementation overhead. Furthermore, we show that safety token regularization enhances training stability and overall performance beyond safety considerations alone. This work offers a practical and readily deployable strategy for continual safety alignment in fine-tuned LLMs.
Thong Bach, Truyen Tran
Applied Artificial Intelligence Initiative (A2I2), Deakin University