cs.AIOct 7, 2026

SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models

Authors: Miao Yu, Hao Huang, Lu Yuan, Yunpeng Li, Kun Wang, Zuming Jiang

Organizations: The University of Hong Kong (HKU) · Chinese Academy of Sciences (CAS) · Information Engineering University (IEU) · Nanyang Technological University (NTU)

Abstract

Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints. To address this, we propose SafeEvo, an interpretability framework from the circuit (sparse subgraphs of an LLM) perspective. SafeEvo first applies an optimization-based extraction algorithm to identify weak refusal circuits in pretrained base LLMs that can independently express refusal behavior. Causally ablating these circuits completely eliminates the base model's refusal of harmful inputs. SafeEvo then traces the evolution of refusal circuits across successive alignment checkpoints and finds that their structures change progressively, suggesting that the alignment tax may result from refusal-circuit updates affecting utility-related parameters. To validate this, SafeEvo introduces Safety Circuit Alignment (SCA), which confines safety updates to the refusal circuits. Experiments across three LLMs and two alignment algorithms show that, on average, SCA outperforms vanilla alignment in three aspects: \textbf{(1) stronger alignment}, lowering harmfulness score by 63.21%; \textbf{(2) less over-refusal}, yielding a 58.44% decrease in refusal rates for benign queries; and \textbf{(3) better utility}, retaining 99.58% of the original model capabilities.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RAS: Measuring LLM Safety Through Refusal Alignment

    Jun 24, 2026Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu +1Large Language Model SafetySafety Alignment

  2. Internalizing Safety Understanding in Large Reasoning Models via Verification

    May 9, 2026Yi Zhang, Yuxin Chen, Leheng Sheng +4Safety ClaimsLarge Reasoning Models

  3. Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

    Sep 3, 2026Hoang Cuong Nguyen, Mark Dras, Usman NaseemRefusalsLarge Language Model Alignment