cs.CLAug 30, 2026

From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling

Authors: Kuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng

Organizations: Computer Science and Engineering, University of California San Diego · Halıcıoğlu Data Science Institute, University of California San Diego

Abstract

Despite extensive alignment efforts, Large Language Models (LLMs) remain vulnerable to generating unsafe content under adversarial prompting, yet the internal mechanisms by which safety behaviors are implemented remain poorly understood. We study LLM safety from a mechanistic interpretability perspective and characterize a multi-stage safety circuit that organizes refusal behavior, consisting of (i) Harmful Detection Heads\textbf{Harmful Detection Heads} that respond to harmful inputs, (ii) Safety Neurons\textbf{Safety Neurons} that mediate and stabilize safety signals in the residual stream, and (iii) Refusal Heads\textbf{Refusal Heads} that translate these signals into safe response generation. Using targeted attention-head and neuron-level interventions, we provide causal evidence consistent with this circuit organization, showing that suppressing upstream Harmful Detection Heads disrupts downstream refusal behavior and that safety neurons mediate this interaction. We validate that this decomposition recurs across multiple LLM architectures and adversarial attack settings, and use simple, architecture-preserving weight scaling as a mechanistic probe to test its functional relevance. Across six LLMs, circuit-guided scaling improves safety rates under attacks by 26.5%, while incurring only a 1.7% accuracy drop across four standard benchmarks. Overall, our results support a circuit-level interpretation of LLM safety and suggest that mechanistic abstractions can reveal stable and transferable patterns underlying aligned behavior.

Explore similar work

CardsList
  1. Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

    Sep 16, 2026Alizishaan Khatri, Chiquita Prabhu, Omkar NeogiSafety FilteringLLM Safety

  2. LLMs Encode Harmfulness and Refusal Separately

    Jul 16, 2025Jiachen Zhao, Jing Huang, Zhengxuan Wu +2LLM InterpretabilityLLM Refusal Behavior

  3. Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs

    Apr 22, 2026Krishiv Agarwal, Ramneet Kaur, Colin Samplawski +6Language Model Safety EvaluationLLM Auditing