cs.LGSep 28, 2026

See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs

Authors: Weiqiao Que, Ruizhe Li, Chengyu Wang, Dakan Wang, Emine Yilmaz, Xiaofeng He

Organizations: School of Computer Science and Technology, East China Normal University, China · School of Computer Science, University of Birmingham, UK · Alibaba Group, China · Exacity Inc. · Centre for Artificial Intelligence, University College London, UK

Abstract

Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation unexpectedly triggers catastrophic safety failures across unrelated domains. Prior static analyses leave training dynamics unmapped, while existing defenses rely on heuristics that degrade utility. We present a dynamic, second-order geometric study of EM. Tracking training trajectories reveals that directional Hessian curvature concentrates sharply on semantic pivot tokens. Grassmannian projections show that, in most settings, harmful-safe gap widens mainly because safe-gradient overlap declines. Leveraging these insights, we introduce a parameter-level Geometric Mitigation Framework that orthogonally projects empirical harmful gradient subspace out of parameter updates. On Qwen2.5-14B-IT, our defense suppresses free-generation EM by up to 80.0%; across the other three of four open-weight instruction-based model families (3B--20B), where single-layer behavioral EM is already near zero, teacher-forced evaluation shows same harmful subspace controls the conditional support of frozen EM responses. Crucially, these diagnostics unmask the illusion of behavioral safety: the same subspace remains measurable and steerable in models where behavioral EM is near zero. Code: https://github.com/WeiqiaoQUE/mechanistic-emergent-misalignment.

Figures & tables

Appendix figures & tables35 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs

    May 11, 2026Krishak Aneja, Manas Mittal, Anmol Goel +2Emergent MisalignmentLarge Language Model Alignment

  2. Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

    Apr 10, 2026Hadas Orgad, Boyi Wei, Kaden Zheng +4Large Language Model SafetyHarmful Content

  3. TAME: Token Attribution and Masking for Emergent misalignment

    Sep 15, 2026Md Rayhanul Masud, Md Rizwan ParvezEmergent MisalignmentLarge Language Model Alignment