cs.CLOct 6, 2026

Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models

Authors: Ziyuan Yang, Wenxuan Ding, Shangbin Feng, Yulia Tsvetkov

Organizations: University of Washington · New York University

Abstract

Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering. However, these methods show inconsistent improvements across models and evaluation distributions, and the intervention into model behavior or internal states introduce safety-utility tradeoffs by over-refusal. Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety. We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not. Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing. Across Qwen3-VL and Kimi-VL, the proposed detector substantially reduces safety errors on the HoliSafe benchmark and resoundingly generalizes to out-of-distribution safety benchmarks featuring different safety patterns, including MISHard and MM-SafetyBench. The success of the proposed router-logit detector also suggests a broader perspective on model internals: rather than focusing only on manipulating internal components to steer behavior, simply reading naturally emerging signals and linking them to an external safety mechanism can provide a simple, effective, and non-intrusive complement to existing safety interventions.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SafeRI: Recognition and Intervention for Token-Level Safety Intervention in Large Vision Language Models

    Sep 3, 2026Caoyuan Ma, Tian Gu, Wenpu Liu +11Safety AlignmentLarge Language Model Alignment

  2. Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment

    Sep 29, 2026Haotian Deng, Wenbin Xing, Gang Xu +5Process-Level Supervision

  3. Do VLMs Share Safety Neurons Across Modalities?

    Aug 31, 2026Jiaxuan Li, Jiahao Zhang, Duc Minh Vo +3Multimodal BenchmarksNeurons