Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering. However, these methods show inconsistent improvements across models and evaluation distributions, and the intervention into model behavior or internal states introduce safety-utility tradeoffs by over-refusal. Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety. We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not. Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing. Across Qwen3-VL and Kimi-VL, the proposed detector substantially reduces safety errors on the HoliSafe benchmark and resoundingly generalizes to out-of-distribution safety benchmarks featuring different safety patterns, including MISHard and MM-SafetyBench. The success of the proposed router-logit detector also suggests a broader perspective on model internals: rather than focusing only on manipulating internal components to steer behavior, simply reading naturally emerging signals and linking them to an external safety mechanism can provide a simple, effective, and non-intrusive complement to existing safety interventions.
Figures & tables
Figure 1: Safety–utility trade-off on HoliSafe, reporting the false-refusal rate (RFR) on safe SSS requests against the mean attack success rate (ASR) across the four unsafe compositions. Lower values on both axes indicate better performance. Points for SFT, prompting, and steering average within each intervention family. Our router-logit detector achieves low ASR while maintaining a moderate false-refusal rate across both models.
Figure 2: Overview of our approach. Existing safety interventions directly modify model inputs, parameters, or routing behavior. In contrast, we non-intrusively read out router logits during prompt prefill and use a lightweight linear detector to identify unsafe multimodal requests before generation.
In-distribution Safety ↓
Out-of-distribution
Method
SSS
SSU
SUU
USU
UUU
Overall
MISHard
MMSafety
MMMU
RFR
ASR
ASR
ASR
ASR
Error
ASR ↓
ASR ↓
Accuracy ↑
Qwen3-VL-30B-A3B-Instruct
Base Model
3.00
50.50
15.00
84.50
15.00
33.60
57.00
23.00
43.00
SFT
Text SFT
4.00
49.00
15.00
89.00
15.00
34.40
58.50
24.50
42.00
Table 1: Performance comparison across Qwen3-VL-30B-A3B-Instruct and Kimi-VL-A3B-Instruct. We compare the base model, SFT, system prompting, expert steering, and our router-logit detector. We report per-category results and overall error on in-distribution HoliSafe, together with out-of-distribution safety and multimodal capability results. Lower is better for all safety metrics. The best results within each model group are highlighted in bold .
Figure 3: Layer-wise router-logit shifts on benign SSS inputs under different safety prompts relative to the base prompt. Image-only prompting induces the largest shifts, particularly in middle and late layers, and is associated with substantially higher false-refusal rates.
Figure 4: Alignment between expert-steering scores and router-detector weights. Steering scores are only weakly correlated with detector coefficient magnitudes across both models, suggesting that features useful for safety readout differ from those identified for direct steering.
Layers
SSS ↓
SSU ↓
SUU ↓
USU ↓
Early
14.50
2.00
1.50
14.00
Middle
9.50
2.00
0.00
13.00
Late
5.50
2.50
0.00
15.00
Early+Middle
11.50
2.50
0.00
13.00
Early+Late
7.00
2.00
0.00
13.00
Middle+Late
6.50
2.50
0.00
13.50
Table 2: Layer-wise ablation of the router-logit detector on Qwen3-VL using different subsets of MoE layers.
k
Selection
SSS ↓
SSU ↓
SUU ↓
USU ↓
8
Selected
8.5
2.5
0.0
12.5
8
Random
10.7±1.7
2.6±0.4
0.1±0.2
17.0±1.2
16
Selected
6.5
2.0
0.0
13.5
16
Random
8.6±1.4
2.4±0.2
0.0±0.0
15.0±1.3
32
Selected
6.0
2.0
0.0
15.0
32
Random
7.8±1.0
1.9±0.2
0.1±0.2
14.0±0.4
Table 3: Experts ablation of router-logit detector on Qwen3-VL. Selected retains the k experts with the largest absolute classifier weights per layer, while Random samples k experts per layer and reports mean/std over five seeds.
Figure 5: Cross-composition generalization of the router-logit detector. Each row excludes one unsafe composition from training and evaluates on all four unsafe test compositions. Cells report absolute ASR, with changes relative to the detector trained on all compositions shown in parentheses. Red boxes indicate the held-out composition.
Signal
SSS
SSU
SUU
USU
UUU
Overall
Base
3.00
50.50
15.00
84.50
15.00
33.60
Router Logit
5.50
2.00
0.00
14.50
0.00
4.40
Hidden State
9.50
1.50
0.00
15.50
0.00
5.30
Table 4: Router-logit and hidden-state safety detectors on Qwen3-VL. SSS reports RFR, while unsafe categories report ASR. Lower is better.
Figure 6: Category-wise performance of the router-logit detector across detection thresholds. Higher thresholds reduce false refusals while increasing ASR.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Layer–expert maps of routing signals used in the steering analysis for Qwen3-VL (top) and Kimi-VL (bottom). Safety-associated routing patterns vary across layers and experts, providing the basis for constructing sparse expert-steering priors.
Figure 8: SSS examples showing both correct acceptance and false refusal. The detector correctly accepts the benign gardening question on the left, but unnecessarily refuses the house-identification question on the right.
Figure 9: SSU and SUU examples where the detector refuses. The SSU prompt is a neutral caption request; some baseline captions introduce disability-related stereotypes. The SUU prompt explicitly requests disruption.
Figure 10: USU examples: a refusal consistent with the benchmark label and a missed unsafe image context. Label agreement should not be equated with harmfulness of every baseline response.
Figure 11: An UUU example in which the detector refuses a workplace-harassment request.
Existing safety alignment methods for vision-language models usually modify the model behavior globally: once the safety parameters are trained or loaded, they participate in both unsafe and already-safe generations. This always-on intervention can unnecessarily perturb the model's original reasoning path and degrade general multimodal capabilities. We argue that safety alignment should be an on-demand intervention rather than a permanent modification to every decoding trajectory. To this end, we propose a streaming recognition and gated LoRA framework for intrinsic VLM safety. During autoregressive generation, a lightweight recognizer estimates whether the current pre-token generation state is safe or unsafe. Its output updates the LoRA gate for the following decoding step; otherwise, generation follows the frozen-backbone policy. The LoRA module is trained from unsafe prefixes, transition statements, and safe continuations, so that it learns to redirect unsafe generations back to safe responses after activation. Experiments across multiple safety and general-purpose benchmarks demonstrate the effectiveness of our method in post-alignment settings.
Caoyuan Ma, Tian Gu, Wenpu Liu +11
The University of Tokyo · Wuhan University · Shanghai AI Laboratory +4
Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time. They may also sacrifice helpfulness by directly refusing requests that could be answered safely. In this paper, we propose Intent-Privilege On-Policy Self-Distillation (OPSD), which leverages evidence-grounded intent as privileged supervision during training to help VLMs recognize implicit risks and provide safe, useful responses instead of blanket refusals. OPSD distills a teacher's intent-conditioned preferences over responses into a student using a single rollout per prompt; the student then responds without intent annotations or an additional safety module. With only 1,447 safety-specific examples - 95% fewer than standard preference datasets - OPSD reduces training time by 5x relative to multi-rollout GRPO-style training and average inference length by 7%. It attains the highest ratio for joint safety-helpfulness success, which measures the proportion of responses that are both safe and helpful, across all five evaluation groups. Remarkably, on pooled SIUO+HoliSafe, this success ratio rises from 43.9% to 53.5%. These results show that training-time intent supervision can improve both safety and helpfulness while substantially reducing data, training, and inference costs.
Haotian Deng, Wenbin Xing, Gang Xu +5
Southern University of Science and Technology · Sun Yat-sen University · Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ) +5
Vision-language models (VLMs) can comply with harmful requests delivered through images, even when their LLM backbones would refuse the same content in text. While prior work characterizes these jailbreaks empirically or at the representation level, how visual inputs perturb safety pathways at the neuron level remains uncharted. We close this gap with a causal, neuron-level analysis of safety mechanisms in 10 VLMs. We propose a two-stage detection pipeline with iterative ablation that accounts for self-repair, and introduce two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval, which decouple visual and textual safety signals. Our analysis reveals: (i) Text safety in VLMs is localizable: ∼88 neurons (<0.01%) whose targeted ablation substantially reduces refusal. (ii) Text safety neurons constitute the dominant refusal pathway: ablating them is the only intervention that consistently and substantially reduces refusal across all models. (iii) Visual safety is high-dimensional and diffuse at the single-neuron level: text safety concentrates in ∼5 subspace directions while visual safety requires ≥50. This gap holds across architectures, explaining why current alignment has not closed the visual safety gap. Project page is at: https://jiaxuan-li.github.io/vlm-safety-neuron/ Warning: this paper may include examples of harmful content.