SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models
Authors: Miao Yu, Hao Huang, Lu Yuan, Yunpeng Li, Kun Wang, Zuming Jiang
Organizations: The University of Hong Kong (HKU) · Chinese Academy of Sciences (CAS) · Information Engineering University (IEU) · Nanyang Technological University (NTU)
Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints. To address this, we propose SafeEvo, an interpretability framework from the circuit (sparse subgraphs of an LLM) perspective. SafeEvo first applies an optimization-based extraction algorithm to identify weak refusal circuits in pretrained base LLMs that can independently express refusal behavior. Causally ablating these circuits completely eliminates the base model's refusal of harmful inputs. SafeEvo then traces the evolution of refusal circuits across successive alignment checkpoints and finds that their structures change progressively, suggesting that the alignment tax may result from refusal-circuit updates affecting utility-related parameters. To validate this, SafeEvo introduces Safety Circuit Alignment (SCA), which confines safety updates to the refusal circuits. Experiments across three LLMs and two alignment algorithms show that, on average, SCA outperforms vanilla alignment in three aspects: \textbf{(1) stronger alignment}, lowering harmfulness score by 63.21%; \textbf{(2) less over-refusal}, yielding a 58.44% decrease in refusal rates for benign queries; and \textbf{(3) better utility}, retaining 99.58% of the original model capabilities.
Figures & tables
Figure 1: From circuit discovery to circuit-targeted alignment. We move beyond endpoint interpretability by tracing refusal circuits from pretrained-only models and their subsequent evolution ( Phase 1 ). We then develop alignment methods that directly target the identified circuits for better safety-utility trade-off ( Phase 2 ).
Figure 2: Overview of our Safety Circuit Alignment algorithm. It first extracts a refusal circuit before LLM alignment ( Step A ) and only fine-tunes the identified parameters in that circuit for precise alignment ( Step B ).
Figure 3: Refusal rate and ASR on HarmBench for base LLMs and the same models with the identified refusal circuits C1 ablated, compared to the density-matched random ablation models.
Figure 4: Evolution of refusal circuits across alignment, traced and quantified by Jaccard similarity.
Method
Objective
Safety
Utility
Over-refusal
HarmB. ( ↓ )
AdvB. (↓ )
GSM8K ( ↑ )
MMLU ( (↑ )
HSwag ( ↑ )
XSTest ( ↓ )
LLAMA-3-8B
Base Model
–
50.31
52.88
50.34
65.36
53.05
0.8
LoRA
SFT
16.98 ↓33.33
12.50 ↓40.38
45.79 ↓4.55
63.75 ↓1.61
55.25 ↑2.20
18.8 ↑18.0
Random
SFT
18.87 ↓31.44
3.27 ↓49.61
49.05 ↓1.29
64.49 ↓0.87
53.35 ↑0.30
2.8 ↑2.0
SN-Tune
SFT
0.00 ↓50.31
0.00 ↓52.88
31.54 ↓18.80
57.78 ↓7.58
51.85 ↓1.20
86.4 ↑85.6
Table 2: Safety, utility, and over-refusal for different baselines under SFT and DPO. Colored marker ↓ and ↑ indicate changes compared to base models, while those following dataset names denote whether lower or higher values are better. Bold and underline mark the best and second-best within each model–objective block.
Method
Objective
Llama-3-8B
Qwen-2.5-7B
Mistral-7B-v0.1
ASR ( ↓ )
Over Refusal ( ↓)
ASR ( ↓ )
Over Refusal ( ↓)
ASR ( ↓ )
Over Refusal ( ↓)
Base Model
–
64.5
0.8
61.5
5.2
66.0
0.8
LoRA
SFT
50.5 ↓14.0
18.8 ↑18.0
30.5 ↓31.0
28.4 ↑23.2
40.0 ↓26.0
7.6 ↑6.8
SCA (Ours)
SFT
62.0 ↓2.5
1.6 ↑0.8
47.5 ↓14.0
6.8 ↑1.6
51.0 ↓15.0
5.6 ↑4.8
LoRA
DPO
53.0 ↓11.5
2.8 ↑2.0
1.0 ↓60.5
34.8 ↑29.6
26.5 ↓39.5
4.8 ↑4.0
SCA (Ours)
DPO
48.5 ↓16.0
3.2 ↑2.4
36.0 ↓25.5
17.6 ↑12.4
21.5 ↓44.5
5.6 ↑4.8
Table 3: Over-refusal and ASR trade-off on XSTest datasets. Marker meanings are the same as above.
Models
SCA ASR (Ours) ↓
Random Baseline ASR ↓
Δ (95% CI)
Llama-3-8B
5.48±3.27
20.28±6.00
14.80 [ 11.48,18.55 ]
Qwen-2.5-7B
11.32±5.51
19.81±3.43
8.49 [ 4.72,12.42 ]
Mistral-7B-v0.1
5.35±2.44
12.58±7.00
7.23 [ 3.77,10.85 ]
Table 4: HarmBench ASR of multiple runs of SCA and the random baseline. ASR is reported as mean ± std over four training seeds. The Δ column denotes the decrease of SCA ASR (relative to the random baseline), with brackets reporting the paired prompt-level 95% bootstrap confidence interval (CI).
Figure 5: Safety–utility trade-offs across different LLM families and alignment methods.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Family
active circuit neurons
patched Linears
trainable params
trainable %
Llama-3-8B
10,376
96
28.31M
0.3513
Qwen-2.5-7B
3,594
84
30.28M
0.3960
Mistral-Inst-v0.1
5,724
96
28.31M
0.3894
Appendix
Table 5: Per-model circuit budget (seed-1 representative; other seeds identical recipe). Trainable params are fixed by LoRA rank × patched-Linear dimensions, independent of mask size.
Family
q
k
v
MLP
tuned slices
Llama-3-8B
3,200
3,200
1,461
1,679
11,219
Qwen-2.5-7B
2,781
2,771
1,270
1,650
10,122
Mistral-Inst-v0.1
3,200
3,200
1,107
1,548
10,603
Appendix
Table 6: SN-Tune selector sizes used in the external baseline. MLP coordinates gate both up and down , so tuned projection slices equal q+k+v+2MLP .
Experiment
Mask object
Neurons (Llama/Qwen/Mistral)
Extracted on
Ablation (Section 3.2 )
C1
10,376 / 9,972 / 7,677
base weights
SCA tuning (Section 4.3 , Table 2 )
C
10,376 / 3,594 / 5,724
base weights
Appendix
Table 7: Mask provenance for ablation and SCA tuning. Each row reports the circuit object, neuron count, and extraction checkpoint used by that experiment.
Setting
Recorded value
Data and split
LLM-LAT harmful-prompt records, each paired with a refusal and harmful-compliance completion (4,948 available); first 100 for training and the next 50 for validation; model-native chat template; maximum length 256.
Search space
One logit per output row of MLP gate / up / down ; model weights frozen; mask logits only are optimized in bf16.
Loss coefficients
α=1 , β=1 , and λmlp=0.05 , applied to the sum of per-module mean soft masks in Equation 6 .
Mask relaxation
Logit initialization q0=0.2 ; temperature 1 ; hard forward gate 1[σ(q)>0.5] with a straight-through estimator.
Optimization
Fused AdamW, learning rate 10−2 , batch size 4, no gradient accumulation, linear schedule, zero warmup, zero weight decay.
Horizon and selection
100-epoch schedule horizon; validation after each epoch; checkpoint comparison uses the final validation minibatch’s two completion losses (sparsity excluded); stop after epoch 2 (50 optimizer updates), which was the selected mask for all three families.
Appendix
Table 8: Recorded circuit-extraction configuration. Circuit-extraction configuration for the ablation experiments in Section 3.2.