VLM safety is commonly evaluated through input- and output-level classification. Such classification is necessary, but it does not reveal whether a safety state is accessible or controllable inside the model. We argue that multimodal safety evaluation should therefore report a \emph{controllability profile} alongside behavioral classification, separating representation-level detectability, cross-modal specificity, intervention sensitivity, and benign-preserving selectivity. Using implicit toxicity as a stress case, we instantiate this profile on LlavaGuard and Qwen3.5 with sparse feature decompositions. LlavaGuard admits localized handles with a narrow benign-preserving intervention range and modest downstream safety gains, whereas Qwen3.5 supports strong representation-level readout but no comparable selective-control regime under the tested operators. These results show that internal readout and controllability can diverge. Future multimodal safety benchmarks should therefore report not only behavioral safety metrics, but also whether safety-relevant internal signals can be intervention-tested and controlled within a validated operating range.
Figures & tables
Figure 1: Two complementary views of multimodal safety evaluation. (A) Output classification measures whether harmful behavior is detected. (B) Controllability evaluation additionally asks whether candidate internal handles can be localized, intervened on, and validated without broad degradation on benign inputs. We argue that these axes should be reported separately rather than treating detection as evidence of control.
Work family
Primary question
Mod.
Behav.
Interv.
X-modal
Output guardrails
Is the case unsafe?
Text/VL
Yes
No
Part.
Cross-modal safety benchmarks
Does context induce harm?
VL
Yes
No
Yes
Mechanistic steering / revision
Can activations alter behavior?
Text/VL
Part.
Yes
Rare
Ours
Is the safety signal controllable?
VL
Down.
Yes
Yes
Table 1: Positioning relative to adjacent safety work. Existing multimodal benchmarks primarily evaluate behavioral safety failures, while mechanistic-control methods test interventions. Our position treats controllability itself as a reporting target and keeps it separate from behavioral classification.
Figure 2: Proximal intervention profiles diverge across models. (a,c) LlavaGuard shows strong proximal sensitivity and a narrow benign-preserving operating range in which top- k clamping suppresses pz(implicit) while preserving pz(benign) around k=4 – 8 . (b,d) Qwen3.5 remains intervention-sensitive but lacks a stable selective range. Small deletions suppress the signal but larger edits rebound, and top- k clamping either preserves benign behavior without suppressing pz(implicit) or enters an unstable collapse regime. The third curve in both clamp panels reports pz(explicit) on benign samples. These results separate intervention sensitivity from selective controllability.
ASR SSU ( by Judge )
α
GPT-5.2
Sonnet 4.5
RR SSS
Δ ASR
0 (baseline)
95.5%
91.0%
0.0%
–
0.5
94.5%
86.5%
0.0%
-1.0% / -4.5%
1.0
92.0%
83.5%
0.0%
-3.5% / -7.5%
Table 3: Output-level safety under the LlavaGuard feature-projection clamp. ASR SSU and RR SSS are evaluated by two independent VLM judges (GPT-5.2 and Claude Sonnet 4.5).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
Value
Notes
Total SAE latents
16,384
SAE sparsity KSAE=128 active latents
Interaction features selected
512
Top features by sf
Top-1 interaction score
42.54
Feature #15926
Top-10 mean score
32.53
Mean over ranks 1–10
Bottom-128 mean score
<0.01
Effectively zero interaction
Appendix
Table 4: LlavaGuard . Summary statistics of SAE interaction features ranked by sf .
Metric
Value
Notes
Total SAE latents
16,384
SAE sparsity KSAE=128 active latents
Interaction features selected
512
Top features by sf
Top-1 interaction score
6.19
Feature #12827
Top-10 mean score
4.14
Mean over ranks 1–10
Bottom-128 mean score
<0.001
Effectively zero interaction
Appendix
Table 5: Qwen3.5-9B . Summary statistics of SAE interaction features ranked by sf .
Rank
Feat. ID
Score
Rank
Feat. ID
Score
1
15926
42.54
11
2613
22.00
2
8180
32.53
12
1652
20.58
3
7410
32.17
13
15937
17.78
4
1462
28.70
14
1519
16.99
5
13199
28.19
15
3628
16.61
6
11594
27.75
16
758
16.39
Appendix
Table 6: LlavaGuard . Top-20 interaction features by interaction score sf .
Rank
Feat. ID
Score
Rank
Feat. ID
Score
1
12827
6.19
11
14744
2.55
2
5204
6.03
12
13061
2.50
3
13718
4.50
13
31
2.34
4
1778
4.47
14
6030
2.22
5
33
4.43
15
2413
1.90
6
5076
3.88
16
157
1.87
Appendix
Table 7: Qwen3.5-9B . Top-20 interaction features by interaction score sf .
Fraction modified
Deletion pz(implicit)
Insertion pz(implicit)
0% (baseline)
87.3%
9.4%
5%
0.0%
10.9%
10%
0.0%
16.7%
20%
0.0%
40.5%
30%
0.0%
83.7%
50%
0.0%
87.3%
Appendix
Table 8: LlavaGuard (layer 24) . Deletion/Insertion interventions measured by the SAE-latent internal classifier pz .
Fraction modified
Deletion pz(implicit)
Insertion pz(implicit)
0% (baseline)
88.7%
9.4%
5%
12.3%
47.5%
10%
37.3%
55.2%
20%
69.4%
79.2%
30%
69.4%
86.2%
50%
69.4%
87.9%
Appendix
Table 9: Qwen3.5-9B (layer 24) . Deletion/Insertion interventions measured by the SAE-latent internal classifier pz .
Clamp k
pz(benign)
pz(explicit)
pz(implicit) on benign
pz(implicit) on implicit
1
0.0%
100.0%
0.0%
0.0%
2
75.3%
24.7%
0.0%
0.0%
4
99.4%
0.6%
0.0%
0.0%
8
100.0%
0.0%
0.0%
0.0%
16
0.0%
100.0%
0.0%
0.0%
32
0.0%
0.0%
100.0%
100.0%
Appendix
Table 10: LlavaGuard . Side effects of clamping top- k SAE interaction features on benign samples, measured by the SAE-latent internal classifier pz .
Clamp k
pz(benign)
pz(explicit)
pz(implicit) on benign
pz(implicit) on implicit
1
82.6%
6.0%
11.5%
88.7%
2
81.5%
5.8%
12.8%
89.9%
4
80.2%
4.0%
15.9%
91.9%
8
12.3%
29.1%
58.6%
97.1%
16
0.1%
70.5%
29.4%
86.8%
32
5.6%
69.8%
24.6%
70.5%
Appendix
Table 11: Qwen3.5-9B . Side effects of clamping top- k SAE interaction features on benign samples, measured by the SAE-latent internal classifier pz .
Direction
Construction
Cosine
Behavioral implication
Raw steering
μimp−μben
0.264
Suppresses pz(implicit) and preserves pz(benign) .
SAE top-10 avg.
Mean of restricted top-10 decoders
0.870
Suppresses pz(implicit) and preserves pz(benign) .
SAE restricted top-1 (#8180)
Decoder of restricted rank-1 feature
0.129
Suppresses pz(implicit) , but collapses benign samples when ablated alone.
Appendix
Table 12: Geometric overlap of representative steering directions with the benign mean hidden vector μben , recomputed on the same SAE and classifier-aware feature set used in the handle-ablation study (LlavaGuard layer 24). The behavioral outcomes in Section A.4 show that cosine overlap alone is not predictive of selectivity.
Feature
Restricted rank
Cosine
Behavioral status
#8180
1
- 0.129
Suppresses pz(implicit) , but collapses benign samples when ablated alone.
#11594
2
- 0.315
Combined with #8180, reduces benign collapse but still causes benign-side errors.
#1519
6
- 0.833
Validated handle. Suppresses pz(implicit) while preserving pz(benign) .
#15926
–
- 0.213
Global top-1 interaction feature, but not read by the pz classifier; ablation has no measurable effect.
Appendix
Table 13: Representative SAE features from the classifier-aware handle-ablation set. Cosine is computed between the decoder vector wfdec (the f -th column of Wdec ) and μben . Behavioral validity is determined by Section A.4 . The validated handle #1519 has high benign overlap but selectively suppresses pz(implicit) while preserving pz(benign) , showing that geometric overlap alone is not a reliable selector.
Handle
Operation
Δpz(implicit)↑
pz(benign)↑
Benign → implicit ↓
Random k=1
zero 1 random latent
0.000
0.780
0.148
Random k=2
zero 2 random latents
0.000
0.780
0.148
SAE top-1 attribution
zero #8180
0.873
0.000
1.000
SAE top-2 attribution
zero #8180, #11594
0.873
0.753
0.238
SAE top-10 mean
project restricted top-10 mean
0.873
1.000
0.000
Raw mean-diff
project dense direction
0.873
1.000
0.000
Appendix
Table 14: LlavaGuard . Behavioral ablation of candidate intervention handles on MDIT group-holdout validation. All probabilities are measured by the SAE-latent classifier pz . We define Δpz(implicit)=pzbase(implicit)−pzedit(implicit) , so higher values indicate stronger suppression of the implicit signal, while higher pz(benign) and lower benign-to-implicit misfire indicate better benign preservation.
Handle
Operation
Δpz(implicit)↑
pz(benign)↑
Benign → implicit ↓
Random k=1
zero 1 random latent
0.000
0.826
0.128
Random k=2
zero 2 random latents
0.000
0.826
0.128
SAE top-1 attribution
zero #12827
−0.001
0.826
0.129
SAE top-2 attribution
zero #12827, #5204
−0.013
0.815
0.142
SAE top-10 mean
project restricted top-10 mean
0.514
0.015
0.985
Raw mean-diff
project dense direction
0.560
0.037
0.963
Appendix
Table 15: Qwen3.5-9B . Behavioral ablation of candidate intervention handles on MDIT validation. All probabilities measured by the SAE-latent classifier pz . In contrast to LlavaGuard, no single-feature or small-set handle simultaneously suppresses pz(implicit) and preserves pz(benign) ; the only strong suppressors cause full benign collapse, consistent with Qwen3.5’s unstable clamp regime ( Table 11 ).
Metric
Definition
Used for
ph(c)
Probability assigned to class c∈C by the pooled-hidden MLP on ho
Text-dominance diagnostic
pr(c)
Probability assigned to class c∈C by the proxy-residual MLP on r
Hidden-space interaction diagnostic
pz(c)
Probability assigned to class c∈C by the SAE-latent classifier
Latent intervention validation
ASRSSU
Fraction of SSU cases judged unsafe after generation
Output-level safety
RRSSS
Failure rate on SSS cases expected to be safe
SSS-side utility / over-refusal proxy
ΔASR
ASR after intervention minus baseline ASR
Mitigation effect
Appendix
Table 16: Metrics used in the case study. Internal metrics evaluate proximal representation-level signals, while output-level metrics evaluate generated behavior.