Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering. However, these methods show inconsistent improvements across models and evaluation distributions, and the intervention into model behavior or internal states introduce safety-utility tradeoffs by over-refusal. Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety. We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not. Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing. Across Qwen3-VL and Kimi-VL, the proposed detector substantially reduces safety errors on the HoliSafe benchmark and resoundingly generalizes to out-of-distribution safety benchmarks featuring different safety patterns, including MISHard and MM-SafetyBench. The success of the proposed router-logit detector also suggests a broader perspective on model internals: rather than focusing only on manipulating internal components to steer behavior, simply reading naturally emerging signals and linking them to an external safety mechanism can provide a simple, effective, and non-intrusive complement to existing safety interventions.
Figures & tables
Figure 1: Safety–utility trade-off on HoliSafe, reporting the false-refusal rate (RFR) on safe SSS requests against the mean attack success rate (ASR) across the four unsafe compositions. Lower values on both axes indicate better performance. Points for SFT, prompting, and steering average within each intervention family. Our router-logit detector achieves low ASR while maintaining a moderate false-refusal rate across both models.
Figure 2: Overview of our approach. Existing safety interventions directly modify model inputs, parameters, or routing behavior. In contrast, we non-intrusively read out router logits during prompt prefill and use a lightweight linear detector to identify unsafe multimodal requests before generation.
In-distribution Safety ↓
Out-of-distribution
Method
SSS
SSU
SUU
USU
UUU
Overall
MISHard
MMSafety
MMMU
RFR
ASR
ASR
ASR
ASR
Error
ASR ↓
ASR ↓
Accuracy ↑
Qwen3-VL-30B-A3B-Instruct
Base Model
3.00
50.50
15.00
84.50
15.00
33.60
57.00
23.00
43.00
SFT
Text SFT
4.00
49.00
15.00
89.00
15.00
34.40
58.50
24.50
42.00
Table 1: Performance comparison across Qwen3-VL-30B-A3B-Instruct and Kimi-VL-A3B-Instruct. We compare the base model, SFT, system prompting, expert steering, and our router-logit detector. We report per-category results and overall error on in-distribution HoliSafe, together with out-of-distribution safety and multimodal capability results. Lower is better for all safety metrics. The best results within each model group are highlighted in bold .
Figure 3: Layer-wise router-logit shifts on benign SSS inputs under different safety prompts relative to the base prompt. Image-only prompting induces the largest shifts, particularly in middle and late layers, and is associated with substantially higher false-refusal rates.
Figure 4: Alignment between expert-steering scores and router-detector weights. Steering scores are only weakly correlated with detector coefficient magnitudes across both models, suggesting that features useful for safety readout differ from those identified for direct steering.
Layers
SSS ↓
SSU ↓
SUU ↓
USU ↓
Early
14.50
2.00
1.50
14.00
Middle
9.50
2.00
0.00
13.00
Late
5.50
2.50
0.00
15.00
Early+Middle
11.50
2.50
0.00
13.00
Early+Late
7.00
2.00
0.00
13.00
Middle+Late
6.50
2.50
0.00
13.50
Table 2: Layer-wise ablation of the router-logit detector on Qwen3-VL using different subsets of MoE layers.
k
Selection
SSS ↓
SSU ↓
SUU ↓
USU ↓
8
Selected
8.5
2.5
0.0
12.5
8
Random
10.7±1.7
2.6±0.4
0.1±0.2
17.0±1.2
16
Selected
6.5
2.0
0.0
13.5
16
Random
8.6±1.4
2.4±0.2
0.0±0.0
15.0±1.3
32
Selected
6.0
2.0
0.0
15.0
32
Random
7.8±1.0
1.9±0.2
0.1±0.2
14.0±0.4
Table 3: Experts ablation of router-logit detector on Qwen3-VL. Selected retains the k experts with the largest absolute classifier weights per layer, while Random samples k experts per layer and reports mean/std over five seeds.
Figure 5: Cross-composition generalization of the router-logit detector. Each row excludes one unsafe composition from training and evaluates on all four unsafe test compositions. Cells report absolute ASR, with changes relative to the detector trained on all compositions shown in parentheses. Red boxes indicate the held-out composition.
Signal
SSS
SSU
SUU
USU
UUU
Overall
Base
3.00
50.50
15.00
84.50
15.00
33.60
Router Logit
5.50
2.00
0.00
14.50
0.00
4.40
Hidden State
9.50
1.50
0.00
15.50
0.00
5.30
Table 4: Router-logit and hidden-state safety detectors on Qwen3-VL. SSS reports RFR, while unsafe categories report ASR. Lower is better.
Figure 6: Category-wise performance of the router-logit detector across detection thresholds. Higher thresholds reduce false refusals while increasing ASR.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Layer–expert maps of routing signals used in the steering analysis for Qwen3-VL (top) and Kimi-VL (bottom). Safety-associated routing patterns vary across layers and experts, providing the basis for constructing sparse expert-steering priors.
Figure 8: SSS examples showing both correct acceptance and false refusal. The detector correctly accepts the benign gardening question on the left, but unnecessarily refuses the house-identification question on the right.
Figure 9: SSU and SUU examples where the detector refuses. The SSU prompt is a neutral caption request; some baseline captions introduce disability-related stereotypes. The SUU prompt explicitly requests disruption.
Figure 10: USU examples: a refusal consistent with the benchmark label and a missed unsafe image context. Label agreement should not be equated with harmfulness of every baseline response.
Figure 11: An UUU example in which the detector refuses a workplace-harassment request.