cs.AISep 29, 2026

The Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization

Authors: Benoit Dherin, Michael Munn, Xavier Gonzalvo, Adrian Goldwaser, Blaz Bratanic, Ananth Balashankar, Andrey Vlasov, Pinzhi Huang, +9 more

Organizations: Google Research · Cambridge University · Google DeepMind · Google · Tel Aviv University

Abstract

Context tokens in a transformer-based language model can be absorbed into the model's weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly the instruction shapes generation. We derive a Contrastive Safety Loss with a suppression weight that controls the tradeoff between emphasizing the safety instruction on harmful queries while suppressing it on harmless queries. Varying the suppression weight maps a relationship between the attack success and the over-refusal rates, supporting the hypothesis that the operator's eigenvalue acts as a continuous dial for the instruction's influence. Moreover, this relationship holds relatively independently of how the Safety Loss is parameterized, yielding Pareto-improved safety instructions for appropriate values of suppression weight.

Figures & tables

Appendix figures & tables32 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails

    Sep 27, 2026Gert Lek, Abele Malan, Chaoyi Zhu +3Large Language Model SafetyMasked Diffusion Language Models

  2. Test-Time Safety Alignment

    Apr 28, 2026Baturay Saglam, Dionysis KalogeriasSafety AlignmentHarmful Content

  3. The Attentional White Bear Effect in Transformer Language Models

    May 27, 2026Rebecca Ramnauth, Brian ScassellatiSuppressionTransformer Architectures