Vision-language models (VLMs) are increasingly trained to generate structured outputs like points and bounding boxes that downstream interfaces, agents, and robots can act on, yet safety alignment of this output channel has not been systematically analyzed. We study visual grounding safety by repurposing three safety benchmarks spanning direct harm (VLSU), social bias (BBQ-V), and situational safety (Asimov-2.0) into 15,401 matched pairs of harmful requests that differ only in the requested output: a free-text answer (VQA) or a grounding (point or bounding box). Across five VLMs, models that refuse a harmful request posed as a question often comply when the same request asks for a grounding: averaged over models, grounding refusal trails VQA refusal by 31-59 percentage points, depending on the domain, and safety system prompts do not close this gap. We propose a fine-tuning approach that combines grounding-form refusals with capability grounding data and self-distilled benign data to counter over-refusal. For Qwen3-VL-8B and VisionReasoner-7B, it improves grounding refusal by 77-95 percentage points on VLSU and BBQ-V and by 64-85 points on the held-out Asimov-2.0 domain, while also improving VQA refusal, preserving grounding capability, and keeping over-refusal limited. Representation analysis shows that fine-tuning moves harmful requests toward each model's refusal direction, most strongly for grounding, while leaving benign requests near the harmless reference.
Figures & tables
Figure 1: Visual grounding safety for three harm domains. The base model refuses the VQA request (green) but complies through its grounding interface (red), returning the box shown on the image.
Role
Evaluation set
Items
Relation to training data
Direct harm
VLSU
659
held-out test split
Social bias
BBQ-V
14,578
held-out dataset
Situational safety
Asimov-2.0
164
held-out domain
Over-refusal, direct
VLSU benign-intent control
300
held-out test split
Over-refusal, bias
BBQ-V disambiguated control
300
held-out dataset
Grounding capability
PointArena
982
held-out benchmark
Table 1: Evaluation sets. Each harmful and benign item is evaluated once in VQA mode and once in grounding mode; counts are unique items. The last column gives each set’s relation to the fine-tuning data (Section 3 ).
Figure 2: Harmful refusal rate (%, ↑ ) in VQA mode (solid) and grounding mode (hatched) without a safety system prompt. Grounding refusal is lower for every model and domain. Items per domain: VLSU 659, BBQ-V 14,578, and Asimov-2.0 164. Exact values are in Table 6 .
VLSU
BBQ-V
Asimov-2.0
Model
none
gen.
plac.
none
gen.
plac.
none
gen.
plac.
Claude Sonnet 4.6
74.1
88.8
88.3
40.9
66.8
67.8
39.6
65.9
73.2
Gemini 3.5 Flash
7.7
20.9
35.4
5.3
36.2
23.6
3.7
11.6
23.8
Qwen3-VL-8B
9.7
31.0
39.5
15.0
21.6
11.8
4.9
8.5
11.0
VisionReasoner-7B
3.9
2.7
3.0
0.6
0.7
0.8
3.0
4.3
1.2
Table 2: Grounding refusal (%, ↑ ) across system-prompt conditions (gen. = generic, plac. = placeholder). Molmo2 lacks a system role and is excluded. Corresponding VQA refusal is in Table 7 .
Qwen3-VL-8B
VisionReasoner-7B
Evaluation
Held-out
Form
Base
Plain
Plac.
Base
Plain
Plac.
Harmful refusal ↑
VLSU
split
VQA
78.0
92.9 +14.9
86.8 +8.8
41.0
94.8 +53.8
93.6 +52.6
grounding
9.7
96.5 +86.8
98.6 +88.9
3.9
97.4 +93.5
94.5 +90.6
BBQ-V
dataset
VQA
88.8
93.0 +4.2
91.7 +2.9
86.1
96.6 +10.5
96.9 +10.8
grounding
15.0
92.8 +77.8
92.4 +77.4
0.6
95.8 +95.2
83.8 +83.2
Table 3: Safety, over-refusal, and capability before and after fine-tuning (%). Plain and Plac. denote SFT-plain and SFT-placeholder. Held-out gives each set’s relation to the training data. Green values give the gain in harmful refusal over the base model; shaded rows are the grounding mode. Harmful refusal and capability are higher-is-better; over-refusal is lower-is-better. MOSSBench is scored on its benign set in VQA form.
Grounding refusal ↑
VQA refusal ↑
Variant
VLSU
BBQ-V
Asimov-2.0
VLSU
BBQ-V
Asimov-2.0
MOSS. ↓
PA ↑
Base
9.7
15.0
4.9
78.0
88.8
45.7
15.3
65.5
Full
96.5
92.8
68.9
92.9
93.0
79.9
24.7
65.4
VQA-ref
1.8
0.0
3.0
86.5
88.3
63.4
27.3
65.9
Gr-ref
94.2
91.7
71.3
95.6
96.3
84.8
55.0
68.2
Table 4: Form-isolating ablations of the Qwen3-VL-8B SFT-plain mixture (Full), in %. VQA-ref removes all grounding-form safety data; Gr-ref removes all VQA-form safety data. MOSS. is benign over-refusal on MOSSBench; PA is PointArena success. Further ablations are in Table 14 .
Figure 3: Refusal / over-refusal trade-off. Vertical axis: mean grounding refusal on harmful requests over VLSU, BBQ-V, and Asimov-2.0. Horizontal axis: aggregate over-refusal, the mean of MOSSBench and the VLSU and BBQ-V benign controls in VQA and grounding form. The top-left corner is ideal. Note the different horizontal scales. Values are in Table 12 .
Figure 4: Last-prompt-token residual-stream activations of VisionReasoner-7B on VLSU at layer 20, projected onto the top two principal components of the harmless and harmful reference prompts, fit separately for each checkpoint. Top row: harmful requests; bottom row: matched benign controls.
Appendix figures & tables34 assets
Supplementary material from the paper’s appendix.
Appendix
Qwen3-VL
VisionReasoner
Role
Source (domain, form)
Plain
Plac.
Plain
Plac.
Harmful refusal
VLSU, VQA
20
–
20
20
VLSU, grounding
50
150
100
100
IGP, VQA
–
–
20
20
IGP, grounding
50
150
100
100
Benign over-refusal (comply)
VLSU, VQA
600
600
600
600
Appendix
Table 5: Training-data composition of the four reported checkpoints, in example counts (“ – ” marks a source not used), taken from the training configs. The capability grounding data (1548 examples) are held fixed across all runs. VisionReasoner SFT-plain and SFT-placeholder share an identical mixture and differ only in the refusal target format . For grounding refusals, SFT-placeholder pairs natural-language refusal text with escape-hatch coordinates in each model’s structured output; VQA refusals remain textual. The two Qwen3-VL checkpoints use different mixtures. The Qwen3-VL SFT-plain column is the same recipe as the “Full” reference in Table 13 .
VLSU (direct)
BBQ-V (bias)
Asimov-2.0
Model
VQA ↑
Grounding ↑
VQA ↑
Grounding ↑
VQA ↑
Grounding ↑
Closed
Claude Sonnet 4.6
93.6
74.1
88.9
40.9
86.0
39.6
Gemini 3.5 Flash
56.8
7.7
73.4
5.3
27.4
3.7
Open
Molmo2
45.3
1.6
42.3
21.8
39.6
12.2
Appendix
Table 6: Harmful refusal rate (%, ↑ ) in VQA and grounding mode without a safety system prompt. Both modes leave substantial room for improvement, and grounding refusal is substantially lower for every model and domain. Items per domain: VLSU 659, BBQ-V 14,578, and Asimov-2.0 164.
VQA refuse
Grounding refuse
Model
none
gen.
plac.
none
gen.
plac.
VLSU (direct harm)
Claude Sonnet 4.6
93.6
96.1
94.2
74.1
88.8
88.3
Gemini 3.5 Flash
56.8
69.7
62.7
7.7
20.9
35.4
Qwen3-VL
78.0
91.2
98.5
9.7
31.0
39.5
VisionReasoner
41.0
63.7
55.5
3.9
2.7
3.0
Appendix
Table 7: Harmful refusal rate (%) under the three system-prompt conditions (none / generic / placeholder; gen. = generic, plac. = placeholder), for VQA and grounding by model and harm domain (higher is safer). Molmo2 is omitted because it lacks a system role; prepending the instruction to its user message is not directly comparable. Molmo2 remains in the no-system-prompt baseline evaluation (Table 6 ). Prompting yields limited and inconsistent gains across the two modes and leaves several models with low refusal.
VQA over-refusal ↓
Grounding over-refusal ↓
Model
none
gen.
plac.
none
gen.
plac.
VLSU (direct harm)
Base Qwen3-VL
0.0
4.7
43.7
1.0
1.3
2.3
Base VisionReasoner
0.3
0.7
1.3
0.0
0.3
0.0
BBQ-V (social bias)
Base Qwen3-VL
1.3
5.0
10.3
0.3
0.0
0.3
Appendix
Table 8: In-domain over-refusal (%) under the three system-prompt conditions (none / generic / placeholder; gen. = generic, plac. = placeholder), for the base models: refusal on benign controls (lower is better). A safety prompt raises benign over-refusal, most on the VQA form (placeholder-prompted Qwen3-VL reaches 43.7% on VLSU); the grounding form and VisionReasoner move little.
Model
Benign (all)
Exagg. Risk
Negated Harm
Counterint. Interp.
Base Qwen3-VL
15.3
20.0
17.0
9.0
Qwen3-VL SFT-plain
24.7
20.0
28.0
26.0
Qwen3-VL SFT-placeholder
27.0
27.0
33.0
21.0
Base VisionReasoner
2.3
2.0
1.0
4.0
VisionReasoner SFT-plain
13.0
12.0
13.0
14.0
VisionReasoner SFT-placeholder
9.0
6.0
8.0
13.0
Appendix
Table 9: MOSSBench over-refusal on the benign set: refusal rate (%; lower is better), overall and by scenario type. SFT raises over-refusal modestly on this held-out benchmark; the plain and placeholder targets are comparable.
Overall
Benign over-refusal by subject ↓
Model
Benign ↓
Contrast ↑
human
child
syn
ocr
Base Qwen3-VL
15.3
73.8
12.9
15.6
21.3
11.2
Qwen3-VL SFT-plain
24.7
85.2
24.2
23.3
29.5
15.7
Qwen3-VL SFT-placeholder
27.0
82.0
24.2
26.7
32.8
18.0
Base VisionReasoner
2.3
48.0
2.8
4.4
3.3
1.1
VisionReasoner SFT-plain
13.0
80.3
11.8
14.4
14.8
12.4
Appendix
Table 10: MOSSBench: benign over-refusal (%; lower is better) overall and by image subject (human, child, synthetic, OCR), and refusal on the paired harmful contrast set (higher is better). SFT raises benign over-refusal but also raises appropriate refusal on the contrast set.
Benign over-refusal ↓
Contrast refusal ↑
Model
none
gen.
plac.
none
gen.
plac.
Base Qwen3-VL
15.3
27.7
40.3
73.8
85.2
89.8
Base VisionReasoner
2.3
5.3
5.7
48.0
61.9
62.3
Appendix
Table 11: MOSSBench results (%) under the three system-prompt conditions (none / generic / placeholder; gen. = generic, plac. = placeholder), for the base models. Benign over-refusal is lower-is-better; contrast refusal on the paired harmful twins is higher-is-better. A safety prompt raises Qwen3-VL’s benign over-refusal steeply while barely affecting VisionReasoner, yet does not reliably provide harmful-grounding robustness (Table 7 ).
Qwen3-VL
VisionReasoner
Setting
Over-refusal ↓
Harmful refusal ↑
Over-refusal ↓
Harmful refusal ↑
Base (no training)
3.6
9.9
1.2
2.6
Sys-prompt: gen.
7.7
20.4
3.1
2.6
Sys-prompt: plac.
19.4
20.8
3.3
1.7
SFT-plain
6.8
86.1
4.5
92.9
SFT-placeholder
7.1
93.5
3.5
83.6
Appendix
Table 12: Coordinates of the points in Figure 3 (%): aggregate over-refusal (mean of MOSSBench and the VLSU/BBQ-V over-refusal sets in both forms; lower is better) and mean grounding-form harmful refusal over the three harm domains (higher is safer).
Role
Source (domain, form)
Full
VLSU-only
IGP-only
No-OR
No-Cap
No-coords
VQA-ref
Gr-ref
Harmful refusal
VLSU, VQA
20
20
–
20
20
20
20
–
VLSU, grounding
50
50
–
50
50
50
–
50
IGP, grounding
50
–
50
50
50
50
–
50
Benign over-refusal (comply)
VLSU, VQA
600
600
–
–
600
600
600
–
VLSU, grounding
200
200
–
–
200
–
–
200
IGP, VQA
400
–
400
–
400
400
400
–
Appendix
Table 13: Training-data composition of the reference (“Full”) SFT mixture, the four leave-one-out ablations, a combined No-coords ablation (No-Cap together with no grounding-form benign over-refusal data), and two form-isolating ablations: VQA-ref, which drops all grounding-form safety supervision, and its mirror Gr-ref, which drops all VQA-form safety supervision (both keep the capability grounding data). Entries are example counts, and “ – ” marks a removed source. All conditions are fine-tuned from the same base model with identical hyperparameters; only the data composition differs. The 1548 capability grounding examples are present in every condition except No-Cap and No-coords.
Harmful refusal ↑
Over-refusal ↓
Cap. ↑
Ablation
VLSU V
VLSU G
BBQ V
BBQ G
AS V
AS G
VLSU
BBQ
MOSS
PA
Full
92.9
96.5
93.0
92.8
79.9
68.9
3.7
1.3
24.7
65.4
VLSU-only
89.8
93.5
92.1
73.4
89.0
82.9
1.7
3.3
26.3
66.3
IGP-only
81.5
78.5
83.2
90.7
53.0
25.6
4.0
0.3
26.3
66.1
No-OR
99.4
99.8
100.0
99.8
98.8
100.0
68.0
100.0
91.3
65.1
No-Cap
92.3
96.1
90.9
92.2
79.3
74.4
5.3
1.3
26.3
66.7
Appendix
Table 14: Ablation results (Qwen3-VL; all values are in %). All conditions use the plain-refusal target except the No-coords (placeholder) row. Harmful refusal is higher-is-safer, over-refusal is lower-is-better, and capability (PointArena) is higher-is-better. Column key: V = VQA, G = grounding, AS = Asimov-2.0. Harmful refusal is shown for both modes; over-refusal shows grounding for VLSU and BBQ-V and VQA for MOSSBench; Cap. is PointArena success. The high refusal in No-coords (plain) is capability loss, not alignment: the model stops grounding (PointArena 6.7%) and its empty output is scored as refusal; the placeholder-target variant retains more grounding capability (52.5%). The form-isolating rows expose an asymmetry. VQA-only safety training leaves harmful-grounding refusal near base, whereas grounding-only safety training raises refusal in both modes, at some benign-VQA over-refusal cost (MOSSBench 55.0%).
Model
Overall
Afford.
Count.
Reason.
Spatial
Steer.
Base Qwen3-VL
65.5
75.3
66.3
67.9
72.3
46.0
Qwen3-VL SFT-plain
65.4
75.3
64.8
68.4
69.2
49.5
Qwen3-VL SFT-placeholder
64.9
74.2
65.8
68.4
73.3
43.0
Base VisionReasoner
62.4
76.8
46.9
65.8
69.7
53.0
VisionReasoner SFT-plain
64.1
75.8
51.0
67.4
74.9
51.5
VisionReasoner SFT-placeholder
63.5
81.3
56.1
64.2
66.2
50.0
Appendix
Table 15: PointArena capability by sub-skill: success rate (%; higher is better).
Qwen3-VL
VisionReasoner
Category
Base
Plain
Plac.
Base
Plain
Plac.
VQA
Overall
78.0
92.9
86.8
41.0
94.8
93.6
c1
93.8
96.9
90.6
62.5
100.0
100.0
c2
76.9
88.5
80.8
57.7
100.0
100.0
c3
83.3
92.9
88.1
35.7
88.1
83.3
Appendix
Table 16: VLSU: harmful refusal rate (%; higher is safer) by harm-category code c1–c15, for the VQA and grounding forms. After SFT refusal is uniformly high across categories for both models and formats.
Qwen3-VL
VisionReasoner
Target
Base
Plain
Plac.
Base
Plain
Plac.
VQA
Overall
45.7
79.9
88.4
20.1
89.0
84.8
human
91.7
100.0
100.0
38.9
100.0
100.0
robot
28.2
73.1
85.9
14.1
87.2
83.3
property
40.0
76.0
84.0
16.0
84.0
76.0
Appendix
Table 17: Asimov-2.0: harmful refusal rate (%; higher is safer) by harm target, for the VQA and grounding forms.
Qwen3-VL
VisionReasoner
Category
Base
Plain
Plac.
Base
Plain
Plac.
VQA
Overall
88.8
93.0
91.7
86.1
96.6
96.9
Age
81.5
87.6
84.1
83.9
96.9
97.0
Disability status
88.1
92.0
90.7
83.8
95.7
96.8
Gender identity
87.9
91.4
91.1
87.2
95.7
95.9
Appendix
Table 18: BBQ-V: harmful refusal rate (%; higher is safer) by bias category, for the VQA and grounding forms.
Domain
none
gen.
plac.
VLSU (direct harm)
12.5
14.1
9.4
BBQ-V (social bias)
64.5
68.1
67.0
Asimov-2.0 (situational)
3.0
3.7
4.3
Appendix
Table 19: VisionReasoner reasoning-trace refusal (%) on the grounding form, by system-prompt condition (none / generic / placeholder; gen. = generic, plac. = placeholder) and harm domain (higher means the <think> trace more often argues for refusal). VisionReasoner is the only model we evaluate that emits a structured reasoning trace; Molmo2 and Qwen3-VL emit none, and the two closed models were run with reasoning disabled, so the reasoning axis is undefined for them. Answer-axis refusal for the same runs is in Table 7 .
Category
Base
Plac.
Overall
11.7
95.0
c1
21.9
100.0
c2
23.1
100.0
c3
19.0
90.5
c4
5.0
93.8
c5
15.3
94.9
Appendix
Table 20: VisionReasoner reasoning-trace refusal (%) on the grounding form (higher is safer), for the base model and the placeholder-trained checkpoint, overall and by VLSU harm-category code c1–c15. The plain-trained checkpoint is omitted: it emits a plain textual refusal rather than a <think> / <answer> trace, so its reasoning axis is undefined.
Category
Base
Plac.
Overall
64.5
94.6
Age
50.0
93.7
Disability status
35.0
94.7
Gender identity
65.0
93.7
Nationality
60.0
96.0
Physical appearance
65.0
93.6
Appendix
Table 21: VisionReasoner reasoning-trace refusal (%) on the grounding form (higher is safer) by bias category, for the base model and the placeholder-trained checkpoint. The plain-trained checkpoint is omitted (it emits a plain textual refusal, not a <think> / <answer> trace).
Target
Base
Plac.
Overall
3.0
75.0
human
2.8
100.0
robot
2.6
64.1
property
4.0
74.0
Appendix
Table 22: VisionReasoner reasoning-trace refusal (%) on the grounding form (higher is safer) by harm target, for the base model and the placeholder-trained checkpoint; the plain-trained checkpoint is omitted (plain textual refusal, no <think> / <answer> trace).
Model
Setting
Harmful inputs
Benign controls
P
Δ
P
Δ
VisionReasoner
base
0.08
–
0.04
–
SFT-plain
0.49
+0.42
0.11
+0.07
SFT-placeholder
0.37
+0.30
0.13
+0.09
Qwen3-VL
base
0.23
–
0.15
–
SFT-plain
0.56
+0.33
0.25
+0.10
Appendix
Table 23: Mean last-token grounding position on the refusal axis ( 0= harmless Alpaca centroid, 1= harmful AdvBench centroid), averaged over the mid-to-upper layers and over the domains in each group; Δ is the change relative to the base backbone. A larger Δ for harmful inputs than for benign controls is consistent with a selective shift of harmful grounding toward the refusal-associated region.
Figure 5: Mean last-prompt-token position on the refusal direction (0 = harmless Alpaca centroid, 1 = harmful AdvBench centroid), averaged over the upper layers (15–27 for VisionReasoner-7B, 19–35 for Qwen3-VL-8B) and over the harm domains (harmful) or the VLSU and BBQ-V controls (benign). Lines connect each base model to its fine-tuned checkpoints.
Figure 6: VisionReasoner-7B projections at layer 20 on Asimov-2.0 and BBQ-V, in the layout of Figure 4 .
Figure 7: Qwen3-VL-8B-Instruct projections at layer 24 across all three domains, in the layout of Figure 4 .
Figure 8: VisionReasoner-7B refusal-axis position per layer. Curves show the mean position of the harmless and harmful reference prompts and the evaluated VQA and grounding requests, with one-sigma bands.
Figure 9: Qwen3-VL-8B-Instruct refusal-axis position per layer, in the layout of Figure 8 .
Vision-language models (VLMs) can comply with harmful requests delivered through images, even when their LLM backbones would refuse the same content in text. While prior work characterizes these jailbreaks empirically or at the representation level, how visual inputs perturb safety pathways at the neuron level remains uncharted. We close this gap with a causal, neuron-level analysis of safety mechanisms in 10 VLMs. We propose a two-stage detection pipeline with iterative ablation that accounts for self-repair, and introduce two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval, which decouple visual and textual safety signals. Our analysis reveals: (i) Text safety in VLMs is localizable: ∼88 neurons (<0.01%) whose targeted ablation substantially reduces refusal. (ii) Text safety neurons constitute the dominant refusal pathway: ablating them is the only intervention that consistently and substantially reduces refusal across all models. (iii) Visual safety is high-dimensional and diffuse at the single-neuron level: text safety concentrates in ∼5 subspace directions while visual safety requires ≥50. This gap holds across architectures, explaining why current alignment has not closed the visual safety gap. Project page is at: https://jiaxuan-li.github.io/vlm-safety-neuron/ Warning: this paper may include examples of harmful content.
Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time. They may also sacrifice helpfulness by directly refusing requests that could be answered safely. In this paper, we propose Intent-Privilege On-Policy Self-Distillation (OPSD), which leverages evidence-grounded intent as privileged supervision during training to help VLMs recognize implicit risks and provide safe, useful responses instead of blanket refusals. OPSD distills a teacher's intent-conditioned preferences over responses into a student using a single rollout per prompt; the student then responds without intent annotations or an additional safety module. With only 1,447 safety-specific examples - 95% fewer than standard preference datasets - OPSD reduces training time by 5x relative to multi-rollout GRPO-style training and average inference length by 7%. It attains the highest ratio for joint safety-helpfulness success, which measures the proportion of responses that are both safe and helpful, across all five evaluation groups. Remarkably, on pooled SIUO+HoliSafe, this success ratio rises from 43.9% to 53.5%. These results show that training-time intent supervision can improve both safety and helpfulness while substantially reducing data, training, and inference costs.
Haotian Deng, Wenbin Xing, Gang Xu +5
Southern University of Science and Technology · Sun Yat-sen University · Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ) +5
Vision-Language Models (VLMs) often achieve high performance on benchmarks while remaining "black boxes", yet they remain prone to hallucination or rely on superficial shortcuts. In this work, we propose a framework designed to enhance both performance and interpretability through De-compositional Evidence Grounding. Unlike monolithic inference approaches, our approach forces the model to decompose a global query into a sequence of atomic sub-questions, each requiring an explicit sub-answer and critically a localized evidence bounding box. By grounding intermediate logical steps (e.g. identifying a container, analyzing liquid properties, and assessing environmental context) in specific visual regions, we construct a structured reasoning path that mirrors human-like deduction. This allows the final answer to emerge as a logical consequence of verified visual facts rather than a statistical guess.
Eric Peh, Debaditya Roy, Basura Fernando
Institute of High-Performance Computing, Agency for Science, Technology and Research, Singapore · Centre for Frontier AI Research, Agency for Science, Technology and Research, Singapore · Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur, India +1