VLM safety is commonly evaluated through input- and output-level classification. Such classification is necessary, but it does not reveal whether a safety state is accessible or controllable inside the model. We argue that multimodal safety evaluation should therefore report a \emph{controllability profile} alongside behavioral classification, separating representation-level detectability, cross-modal specificity, intervention sensitivity, and benign-preserving selectivity. Using implicit toxicity as a stress case, we instantiate this profile on LlavaGuard and Qwen3.5 with sparse feature decompositions. LlavaGuard admits localized handles with a narrow benign-preserving intervention range and modest downstream safety gains, whereas Qwen3.5 supports strong representation-level readout but no comparable selective-control regime under the tested operators. These results show that internal readout and controllability can diverge. Future multimodal safety benchmarks should therefore report not only behavioral safety metrics, but also whether safety-relevant internal signals can be intervention-tested and controlled within a validated operating range.
Figures & tables
Figure 1: Two complementary views of multimodal safety evaluation. (A) Output classification measures whether harmful behavior is detected. (B) Controllability evaluation additionally asks whether candidate internal handles can be localized, intervened on, and validated without broad degradation on benign inputs. We argue that these axes should be reported separately rather than treating detection as evidence of control.
Work family
Primary question
Mod.
Behav.
Interv.
X-modal
Output guardrails
Is the case unsafe?
Text/VL
Yes
No
Part.
Cross-modal safety benchmarks
Does context induce harm?
VL
Yes
No
Yes
Mechanistic steering / revision
Can activations alter behavior?
Text/VL
Part.
Yes
Rare
Ours
Is the safety signal controllable?
VL
Down.
Yes
Yes
Table 1: Positioning relative to adjacent safety work. Existing multimodal benchmarks primarily evaluate behavioral safety failures, while mechanistic-control methods test interventions. Our position treats controllability itself as a reporting target and keeps it separate from behavioral classification.
Figure 2: Proximal intervention profiles diverge across models. (a,c) LlavaGuard shows strong proximal sensitivity and a narrow benign-preserving operating range in which top- k clamping suppresses pz(implicit) while preserving pz(benign) around k=4 – 8 . (b,d) Qwen3.5 remains intervention-sensitive but lacks a stable selective range. Small deletions suppress the signal but larger edits rebound, and top- k clamping either preserves benign behavior without suppressing pz(implicit) or enters an unstable collapse regime. The third curve in both clamp panels reports pz(explicit) on benign samples. These results separate intervention sensitivity from selective controllability.
ASR SSU ( by Judge )
α
GPT-5.2
Sonnet 4.5
RR SSS
Δ ASR
0 (baseline)
95.5%
91.0%
0.0%
–
0.5
94.5%
86.5%
0.0%
-1.0% / -4.5%
1.0
92.0%
83.5%
0.0%
-3.5% / -7.5%
Table 3: Output-level safety under the LlavaGuard feature-projection clamp. ASR SSU and RR SSS are evaluated by two independent VLM judges (GPT-5.2 and Claude Sonnet 4.5).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
Value
Notes
Total SAE latents
16,384
SAE sparsity KSAE=128 active latents
Interaction features selected
512
Top features by sf
Top-1 interaction score
42.54
Feature #15926
Top-10 mean score
32.53
Mean over ranks 1–10
Bottom-128 mean score
<0.01
Effectively zero interaction
Appendix
Table 4: LlavaGuard . Summary statistics of SAE interaction features ranked by sf .
Metric
Value
Notes
Total SAE latents
16,384
SAE sparsity KSAE=128 active latents
Interaction features selected
512
Top features by sf
Top-1 interaction score
6.19
Feature #12827
Top-10 mean score
4.14
Mean over ranks 1–10
Bottom-128 mean score
<0.001
Effectively zero interaction
Appendix
Table 5: Qwen3.5-9B . Summary statistics of SAE interaction features ranked by sf .
Rank
Feat. ID
Score
Rank
Feat. ID
Score
1
15926
42.54
11
2613
22.00
2
8180
32.53
12
1652
20.58
3
7410
32.17
13
15937
17.78
4
1462
28.70
14
1519
16.99
5
13199
28.19
15
3628
16.61
6
11594
27.75
16
758
16.39
Appendix
Table 6: LlavaGuard . Top-20 interaction features by interaction score sf .
Rank
Feat. ID
Score
Rank
Feat. ID
Score
1
12827
6.19
11
14744
2.55
2
5204
6.03
12
13061
2.50
3
13718
4.50
13
31
2.34
4
1778
4.47
14
6030
2.22
5
33
4.43
15
2413
1.90
6
5076
3.88
16
157
1.87
Appendix
Table 7: Qwen3.5-9B . Top-20 interaction features by interaction score sf .
Fraction modified
Deletion pz(implicit)
Insertion pz(implicit)
0% (baseline)
87.3%
9.4%
5%
0.0%
10.9%
10%
0.0%
16.7%
20%
0.0%
40.5%
30%
0.0%
83.7%
50%
0.0%
87.3%
Appendix
Table 8: LlavaGuard (layer 24) . Deletion/Insertion interventions measured by the SAE-latent internal classifier pz .
Fraction modified
Deletion pz(implicit)
Insertion pz(implicit)
0% (baseline)
88.7%
9.4%
5%
12.3%
47.5%
10%
37.3%
55.2%
20%
69.4%
79.2%
30%
69.4%
86.2%
50%
69.4%
87.9%
Appendix
Table 9: Qwen3.5-9B (layer 24) . Deletion/Insertion interventions measured by the SAE-latent internal classifier pz .
Clamp k
pz(benign)
pz(explicit)
pz(implicit) on benign
pz(implicit) on implicit
1
0.0%
100.0%
0.0%
0.0%
2
75.3%
24.7%
0.0%
0.0%
4
99.4%
0.6%
0.0%
0.0%
8
100.0%
0.0%
0.0%
0.0%
16
0.0%
100.0%
0.0%
0.0%
32
0.0%
0.0%
100.0%
100.0%
Appendix
Table 10: LlavaGuard . Side effects of clamping top- k SAE interaction features on benign samples, measured by the SAE-latent internal classifier pz .
Clamp k
pz(benign)
pz(explicit)
pz(implicit) on benign
pz(implicit) on implicit
1
82.6%
6.0%
11.5%
88.7%
2
81.5%
5.8%
12.8%
89.9%
4
80.2%
4.0%
15.9%
91.9%
8
12.3%
29.1%
58.6%
97.1%
16
0.1%
70.5%
29.4%
86.8%
32
5.6%
69.8%
24.6%
70.5%
Appendix
Table 11: Qwen3.5-9B . Side effects of clamping top- k SAE interaction features on benign samples, measured by the SAE-latent internal classifier pz .
Direction
Construction
Cosine
Behavioral implication
Raw steering
μimp−μben
0.264
Suppresses pz(implicit) and preserves pz(benign) .
SAE top-10 avg.
Mean of restricted top-10 decoders
0.870
Suppresses pz(implicit) and preserves pz(benign) .
SAE restricted top-1 (#8180)
Decoder of restricted rank-1 feature
0.129
Suppresses pz(implicit) , but collapses benign samples when ablated alone.
Appendix
Table 12: Geometric overlap of representative steering directions with the benign mean hidden vector μben , recomputed on the same SAE and classifier-aware feature set used in the handle-ablation study (LlavaGuard layer 24). The behavioral outcomes in Section A.4 show that cosine overlap alone is not predictive of selectivity.
Feature
Restricted rank
Cosine
Behavioral status
#8180
1
- 0.129
Suppresses pz(implicit) , but collapses benign samples when ablated alone.
#11594
2
- 0.315
Combined with #8180, reduces benign collapse but still causes benign-side errors.
#1519
6
- 0.833
Validated handle. Suppresses pz(implicit) while preserving pz(benign) .
#15926
–
- 0.213
Global top-1 interaction feature, but not read by the pz classifier; ablation has no measurable effect.
Appendix
Table 13: Representative SAE features from the classifier-aware handle-ablation set. Cosine is computed between the decoder vector wfdec (the f -th column of Wdec ) and μben . Behavioral validity is determined by Section A.4 . The validated handle #1519 has high benign overlap but selectively suppresses pz(implicit) while preserving pz(benign) , showing that geometric overlap alone is not a reliable selector.
Handle
Operation
Δpz(implicit)↑
pz(benign)↑
Benign → implicit ↓
Random k=1
zero 1 random latent
0.000
0.780
0.148
Random k=2
zero 2 random latents
0.000
0.780
0.148
SAE top-1 attribution
zero #8180
0.873
0.000
1.000
SAE top-2 attribution
zero #8180, #11594
0.873
0.753
0.238
SAE top-10 mean
project restricted top-10 mean
0.873
1.000
0.000
Raw mean-diff
project dense direction
0.873
1.000
0.000
Appendix
Table 14: LlavaGuard . Behavioral ablation of candidate intervention handles on MDIT group-holdout validation. All probabilities are measured by the SAE-latent classifier pz . We define Δpz(implicit)=pzbase(implicit)−pzedit(implicit) , so higher values indicate stronger suppression of the implicit signal, while higher pz(benign) and lower benign-to-implicit misfire indicate better benign preservation.
Handle
Operation
Δpz(implicit)↑
pz(benign)↑
Benign → implicit ↓
Random k=1
zero 1 random latent
0.000
0.826
0.128
Random k=2
zero 2 random latents
0.000
0.826
0.128
SAE top-1 attribution
zero #12827
−0.001
0.826
0.129
SAE top-2 attribution
zero #12827, #5204
−0.013
0.815
0.142
SAE top-10 mean
project restricted top-10 mean
0.514
0.015
0.985
Raw mean-diff
project dense direction
0.560
0.037
0.963
Appendix
Table 15: Qwen3.5-9B . Behavioral ablation of candidate intervention handles on MDIT validation. All probabilities measured by the SAE-latent classifier pz . In contrast to LlavaGuard, no single-feature or small-set handle simultaneously suppresses pz(implicit) and preserves pz(benign) ; the only strong suppressors cause full benign collapse, consistent with Qwen3.5’s unstable clamp regime ( Table 11 ).
Metric
Definition
Used for
ph(c)
Probability assigned to class c∈C by the pooled-hidden MLP on ho
Text-dominance diagnostic
pr(c)
Probability assigned to class c∈C by the proxy-residual MLP on r
Hidden-space interaction diagnostic
pz(c)
Probability assigned to class c∈C by the SAE-latent classifier
Latent intervention validation
ASRSSU
Fraction of SSU cases judged unsafe after generation
Output-level safety
RRSSS
Failure rate on SSS cases expected to be safe
SSS-side utility / over-refusal proxy
ΔASR
ASR after intervention minus baseline ASR
Mitigation effect
Appendix
Table 16: Metrics used in the case study. Internal metrics evaluate proximal representation-level signals, while output-level metrics evaluate generated behavior.
Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these evaluations target outputs rather than representation-level vulnerability under intervention. We formalize this discrepancy as the audit gap: the difference between behavioral safety and robustness under intervention. To study this gap, we construct dissociated models that preserve safe outward behavior while remaining vulnerable in the latent space. We introduce an intervention-based evaluation framework to test model robustness through soft interventions in parameter and latent spaces, including harmful fine-tuning and layer-wise latent perturbations. To formalize the evaluation, we propose the Latent Vulnerability Score (LVS) to measure how easily harmful behavior can be elicited by bounded latent perturbations. Using this evaluation framework, we show that behavioral safety metrics are insufficient measures of representation-level robustness across multiple safely and unsafely aligned state-of-the-art models. Notably, dissociated models show substantially elevated LVSs despite comparable refusal behavior under harmful intervention, with intermediate representations being the most sensitive to intervention. Our results suggest that behavioral safety evaluation alone provides an incomplete picture of model robustness, motivating representation-aware audits of latent vulnerability and observable behavior.
Enyi Jiang, Anders Gjølbye, Yibo Jacky Zhang +1
1Stanford University · University of Illinois Urbana-Champaign · 3Technical University of Denmark
Safety evaluation of multimodal large language models requires tracking not only whether an attack succeeds, but also how the interaction unfolds across turns and input modalities. We present MUSE (Multimodal Unified Safety Evaluation), an open-source, browser-based, run-centric platform for multimodal safety evaluation. MUSE treats each attack run as the persistent unit of execution, inspection, and analysis, preserving its configuration, multi-turn trajectory, delivered modalities and media, target responses, and safety judgments. A five-level response taxonomy further distinguishes full Compliance from Partial Compliance and refusal behavior, yielding hard ASR, soft ASR, and gray-zone width (GZW). Across 11,700 evaluations on six multimodal LLMs, direct text-only requests yield only 3.1% macro hard ASR and 4.4% soft ASR, while iterative attack procedures are substantially more effective. Attack effectiveness also varies substantially with the attacker backbone. In contrast, Inter-Turn Modality Switching (ITMS), evaluated as a controlled delivery-modality probe, does not consistently increase attack success. These results demonstrate the value of run-centric, fine-grained evaluation for characterizing multimodal safety behavior beyond a single binary success metric.
While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related annotations and rely on unidimensional metrics, hindering comprehensive robustness evaluation. To address this, we propose MME-Safety, a rigorously verified benchmark featuring a unique four-dimensional annotation schema that categorizes risk scenarios, harm severity, and modality-specific stealth levels. Furthermore, we introduce a hierarchical evaluation framework to assess fundamental response reliability, actual risk exposure, and the structural integrity of defensive behaviors. Extensive zero-shot evaluations across 17 state-of-the-art MLLMs provide a comprehensive safety profile of current multimodal systems. Our analysis systematically investigates cross-modal input configurations and uncovers safety implications associated with Chain-of-Thought (CoT) reasoning. These multifaceted findings underscore the urgent need for robust, reasoning-aware safety alignment in the multimodal landscape.