SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models
Authors: Miao Yu, Hao Huang, Lu Yuan, Yunpeng Li, Kun Wang, Zuming Jiang
Organizations: The University of Hong Kong (HKU) · Chinese Academy of Sciences (CAS) · Information Engineering University (IEU) · Nanyang Technological University (NTU)
Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints. To address this, we propose SafeEvo, an interpretability framework from the circuit (sparse subgraphs of an LLM) perspective. SafeEvo first applies an optimization-based extraction algorithm to identify weak refusal circuits in pretrained base LLMs that can independently express refusal behavior. Causally ablating these circuits completely eliminates the base model's refusal of harmful inputs. SafeEvo then traces the evolution of refusal circuits across successive alignment checkpoints and finds that their structures change progressively, suggesting that the alignment tax may result from refusal-circuit updates affecting utility-related parameters. To validate this, SafeEvo introduces Safety Circuit Alignment (SCA), which confines safety updates to the refusal circuits. Experiments across three LLMs and two alignment algorithms show that, on average, SCA outperforms vanilla alignment in three aspects: \textbf{(1) stronger alignment}, lowering harmfulness score by 63.21%; \textbf{(2) less over-refusal}, yielding a 58.44% decrease in refusal rates for benign queries; and \textbf{(3) better utility}, retaining 99.58% of the original model capabilities.
Figures & tables
Figure 1: From circuit discovery to circuit-targeted alignment. We move beyond endpoint interpretability by tracing refusal circuits from pretrained-only models and their subsequent evolution ( Phase 1 ). We then develop alignment methods that directly target the identified circuits for better safety-utility trade-off ( Phase 2 ).
Figure 2: Overview of our Safety Circuit Alignment algorithm. It first extracts a refusal circuit before LLM alignment ( Step A ) and only fine-tunes the identified parameters in that circuit for precise alignment ( Step B ).
Figure 3: Refusal rate and ASR on HarmBench for base LLMs and the same models with the identified refusal circuits C1 ablated, compared to the density-matched random ablation models.
Figure 4: Evolution of refusal circuits across alignment, traced and quantified by Jaccard similarity.
Method
Objective
Safety
Utility
Over-refusal
HarmB. ( ↓ )
AdvB. (↓ )
GSM8K ( ↑ )
MMLU ( (↑ )
HSwag ( ↑ )
XSTest ( ↓ )
LLAMA-3-8B
Base Model
–
50.31
52.88
50.34
65.36
53.05
0.8
LoRA
SFT
16.98 ↓33.33
12.50 ↓40.38
45.79 ↓4.55
63.75 ↓1.61
55.25 ↑2.20
18.8 ↑18.0
Random
SFT
18.87 ↓31.44
3.27 ↓49.61
49.05 ↓1.29
64.49 ↓0.87
53.35 ↑0.30
2.8 ↑2.0
SN-Tune
SFT
0.00 ↓50.31
0.00 ↓52.88
31.54 ↓18.80
57.78 ↓7.58
51.85 ↓1.20
86.4 ↑85.6
Table 2: Safety, utility, and over-refusal for different baselines under SFT and DPO. Colored marker ↓ and ↑ indicate changes compared to base models, while those following dataset names denote whether lower or higher values are better. Bold and underline mark the best and second-best within each model–objective block.
Method
Objective
Llama-3-8B
Qwen-2.5-7B
Mistral-7B-v0.1
ASR ( ↓ )
Over Refusal ( ↓)
ASR ( ↓ )
Over Refusal ( ↓)
ASR ( ↓ )
Over Refusal ( ↓)
Base Model
–
64.5
0.8
61.5
5.2
66.0
0.8
LoRA
SFT
50.5 ↓14.0
18.8 ↑18.0
30.5 ↓31.0
28.4 ↑23.2
40.0 ↓26.0
7.6 ↑6.8
SCA (Ours)
SFT
62.0 ↓2.5
1.6 ↑0.8
47.5 ↓14.0
6.8 ↑1.6
51.0 ↓15.0
5.6 ↑4.8
LoRA
DPO
53.0 ↓11.5
2.8 ↑2.0
1.0 ↓60.5
34.8 ↑29.6
26.5 ↓39.5
4.8 ↑4.0
SCA (Ours)
DPO
48.5 ↓16.0
3.2 ↑2.4
36.0 ↓25.5
17.6 ↑12.4
21.5 ↓44.5
5.6 ↑4.8
Table 3: Over-refusal and ASR trade-off on XSTest datasets. Marker meanings are the same as above.
Models
SCA ASR (Ours) ↓
Random Baseline ASR ↓
Δ (95% CI)
Llama-3-8B
5.48±3.27
20.28±6.00
14.80 [ 11.48,18.55 ]
Qwen-2.5-7B
11.32±5.51
19.81±3.43
8.49 [ 4.72,12.42 ]
Mistral-7B-v0.1
5.35±2.44
12.58±7.00
7.23 [ 3.77,10.85 ]
Table 4: HarmBench ASR of multiple runs of SCA and the random baseline. ASR is reported as mean ± std over four training seeds. The Δ column denotes the decrease of SCA ASR (relative to the random baseline), with brackets reporting the paired prompt-level 95% bootstrap confidence interval (CI).
Figure 5: Safety–utility trade-offs across different LLM families and alignment methods.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Family
active circuit neurons
patched Linears
trainable params
trainable %
Llama-3-8B
10,376
96
28.31M
0.3513
Qwen-2.5-7B
3,594
84
30.28M
0.3960
Mistral-Inst-v0.1
5,724
96
28.31M
0.3894
Appendix
Table 5: Per-model circuit budget (seed-1 representative; other seeds identical recipe). Trainable params are fixed by LoRA rank × patched-Linear dimensions, independent of mask size.
Family
q
k
v
MLP
tuned slices
Llama-3-8B
3,200
3,200
1,461
1,679
11,219
Qwen-2.5-7B
2,781
2,771
1,270
1,650
10,122
Mistral-Inst-v0.1
3,200
3,200
1,107
1,548
10,603
Appendix
Table 6: SN-Tune selector sizes used in the external baseline. MLP coordinates gate both up and down , so tuned projection slices equal q+k+v+2MLP .
Experiment
Mask object
Neurons (Llama/Qwen/Mistral)
Extracted on
Ablation (Section 3.2 )
C1
10,376 / 9,972 / 7,677
base weights
SCA tuning (Section 4.3 , Table 2 )
C
10,376 / 3,594 / 5,724
base weights
Appendix
Table 7: Mask provenance for ablation and SCA tuning. Each row reports the circuit object, neuron count, and extraction checkpoint used by that experiment.
Setting
Recorded value
Data and split
LLM-LAT harmful-prompt records, each paired with a refusal and harmful-compliance completion (4,948 available); first 100 for training and the next 50 for validation; model-native chat template; maximum length 256.
Search space
One logit per output row of MLP gate / up / down ; model weights frozen; mask logits only are optimized in bf16.
Loss coefficients
α=1 , β=1 , and λmlp=0.05 , applied to the sum of per-module mean soft masks in Equation 6 .
Mask relaxation
Logit initialization q0=0.2 ; temperature 1 ; hard forward gate 1[σ(q)>0.5] with a straight-through estimator.
Optimization
Fused AdamW, learning rate 10−2 , batch size 4, no gradient accumulation, linear schedule, zero warmup, zero weight decay.
Horizon and selection
100-epoch schedule horizon; validation after each epoch; checkpoint comparison uses the final validation minibatch’s two completion losses (sparsity excluded); stop after epoch 2 (50 optimizer updates), which was the selected mask for all three families.
Appendix
Table 8: Recorded circuit-extraction configuration. Circuit-extraction configuration for the ablation experiments in Section 3.2.
Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their outputs violate a safety policy. Although useful, output-level evaluation is expensive, sensitive to judge choice, and easily tied to fixed question banks. We propose SafeVec, a white-box evaluation procedure that measures safety from internal representations rather than generated answers. SafeVec first extracts layer-wise refusal directions from a safety-aligned reference model, then selects stable layer windows where safe and unsafe behaviors are separable, and finally scores a target model by measuring whether its hidden states align with these refusal directions under unsafe and jailbreak prompts. The resulting metric, RAS (Refusal Alignment Score), maps representation-level refusal alignment to a calibrated 0-100 safety score. Across Llama, Gemma, and Qwen model families, RAS separates aligned models from uncensored and abliterated variants, tracks output-level attack success rate, and is substantially faster than judge-based evaluation. These results suggest that refusal alignment provides a compact and efficient signal for white-box LLM safety evaluation.
Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu +1
National Yang Ming Chiao Tung University · Hon Hai Research Institute
While explicit Chain-of-Thought (CoT) empowers large reasoning models (LRMs), it enables the generation of riskier final answers. Current alignment paradigms primarily rely on externally enforced compliance, optimizing models to detect malicious prompts rather than evaluating the safety of their own outputs. We argue that this approach remains largely behavioral: our empirical analysis reveals that ostensibly aligned models lack intrinsic safety understanding, often failing to verify their own response safety and remaining vulnerable to adversarial jailbreaks. To address this fundamental limitation, we propose Safety Internal (SInternal), a framework that internalizes safety specifications by training LRMs exclusively on safety verification tasks to critique their own generated answers using expert reasoning trajectories. We demonstrate that learning to verify induces a strong generalization for response safety, significantly enhancing robustness against out-of-domain jailbreaks. Furthermore, when combined with reinforcement learning, SInternal serves as a superior initialization compared to standard supervised fine-tuning, suggesting that internalizing safety understanding creates a more robust foundation for alignment than merely mimicking safe behaviors. Our codes are available at https://github.com/AlphaLab-USTC/SInternal
Yi Zhang, Yuxin Chen, Leheng Sheng +4
University of Science and Technology of China · National University of Singapore · Shanghai Artificial Intelligence Laboratory
How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a safety decision), and preference optimization (ORPO) - across three architecturally distinct models (Llama-3.1-8B, Gemma-2-9B, Qwen3-8B). We find that training method, not just data, reshapes how refusal is computed internally: reasoning-augmented training consistently produces a distinct kind of refusal computation, visible across all three models, while architecture independently shapes internal structure and how reliably refusal can be steered. Most importantly, no method we study achieves all three properties we would want from safe alignment at once: refusal that isn't concentrated in a few fragile components, safety gains that don't cost general capability, and safety behavior correctable through small, targeted edits. We caution against treating current post-training methods as a solved, reliable defense, especially for security-critical use. Code and models are available in https://github.com/hoangcuongnguyen2001/Beyond-Shallow-Alignment.