cs.AISep 28, 2026

When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety

Authors: Tianyi Guan, Jianhui Chen, Liangming Pan

Organizations: State Key Laboratory of Multimedia Information Processing, Peking University · School of Computer Science, Peking University

Abstract

Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactions. Established behavioral safeguards include alignment methods that optimize model outputs and text monitors that assess interaction text. Representation engineering instead reads or modifies internal model states, but the relative strengths of these approaches remain unclear because they are often evaluated under different settings. We present a matched evaluation across two tracks. For safety control, we compare DPO, a behavioral alignment method, with three representation steering methods across robustness, practicality, and granularity. DPO provides the strongest overall control and generally improves with increasing training data, although its safety can degrade after subsequent benign fine-tuning. Representation steering remains competitive primarily in low-data settings, particularly with high-quality contrastive data. For safety monitoring, we compare representation probes with fine-tuned and open-weight text monitors across full-response detection, early detection, and computational cost. Specialized text monitors achieve the strongest overall detection accuracy, while representation probes remain competitive at substantially lower marginal cost. Finally, monitor-guided interventions recover much of the safety lost by DPO after benign fine-tuning, with little additional over-refusal. Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

    Jun 6, 2026Enyi Jiang, Anders Gjølbye, Yibo Jacky Zhang +1Large Language Model SafetyRepresentation-Level Misalignment

  2. Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

    Sep 16, 2026Alizishaan Khatri, Chiquita Prabhu, Omkar NeogiLarge Language Model SafetyGuardrail

  3. LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards

    Aug 4, 2026Zhinan Liu, Jie Li, Mingyu Kang +1Large Language Model SafetyEfficient Latent Reasoning