Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment
Authors: Haotian Deng, Wenbin Xing, Gang Xu, Tao He, Jinkai Zheng, Chun Li, Zheng Zhu, Ming Li
Organizations: Southern University of Science and Technology · Sun Yat-sen University · Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ) · University of Electronic Science and Technology of China · Hangzhou Dianzi University · Shenzhen MSU-BIT University · GigaAI · The Chinese University of Hong Kong (Shenzhen)
Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time. They may also sacrifice helpfulness by directly refusing requests that could be answered safely. In this paper, we propose Intent-Privilege On-Policy Self-Distillation (OPSD), which leverages evidence-grounded intent as privileged supervision during training to help VLMs recognize implicit risks and provide safe, useful responses instead of blanket refusals. OPSD distills a teacher's intent-conditioned preferences over responses into a student using a single rollout per prompt; the student then responds without intent annotations or an additional safety module. With only 1,447 safety-specific examples - 95% fewer than standard preference datasets - OPSD reduces training time by 5x relative to multi-rollout GRPO-style training and average inference length by 7%. It attains the highest ratio for joint safety-helpfulness success, which measures the proportion of responses that are both safe and helpful, across all five evaluation groups. Remarkably, on pooled SIUO+HoliSafe, this success ratio rises from 43.9% to 53.5%. These results show that training-time intent supervision can improve both safety and helpfulness while substantially reducing data, training, and inference costs.
Figures & tables
Figure 1: Motivation of Intent-Privilege OPSD. It achieves comprehensive data, training, and inference efficiency. A detailed analysis of defense collapse on implicit risks is provided in Sec. D .
Context
S↑
H↑
S=3↑
H≥2↑
Success ↑
FSR
No privilege
1.633
1.850
25.0
75.0
25.0
12.9
Source Privilege
1.767
1.683
11.7
66.7
11.7
45.0
Evidence-Grounded Privilege
1.917
2.283
41.7
95.0
41.7
12.1
Shuffled Evidence-Grounded Privilege
1.450
1.867
23.3
76.7
23.3
17.1
Table 1: Frozen-teacher Comparison. Safety S∈[−3,3] evaluates boundary handling and evidence calibration; Helpfulness H∈[0,3] measures useful assistance. Success requires S=3 and H≥2 . The last four columns are percentages. FSR means First-sentence refusal. New uses the intent annotation z produced by the pipeline in Sec. 3.2 ; Shuffled New uses another input’s annotation. Refusal is the sampled first-sentence frequency.
Figure 2: Qualitative comparison under different privileged intent conditions. Source Privilege (a) injects speculative threat assumptions, leading to a defensive blanket refusal. Conversely, Evidence-Grounded Privilege (b) preserves uncertainty and provides calibrated guidance.
Figure 3: Evidence-grounded intent re-annotation pipeline. Modality-isolated evidence is combined by a joint builder and audited against the original inputs to produce intent annotations z . Accepted annotations provide the Evidence-Grounded Privilege used in both frozen-teacher evaluation and OPSD.
Figure 4: Intent-Privilege OPSD. (a) Data Efficiency: achieves a 95% reduction using 1,447 training examples; (b) Training Efficiency: uses a single rollout to slash compute time by 5×; (c) Inference Efficiency: operates without extra components, reducing average inference length by 7%.
Dataset
N
Base
SPA-VL
VLGuard
TiS
SafeGRPO
Ours
SIUO+HoliSafe
767
43.9
50.3
40.8
45.9
45.4
53.5
BeaverTails-V
200
31.0
44.0
8.5
32.5
30.5
49.0
MSSBench
400
39.0
35.5
32.5
35.7
38.8
39.0
MOSSBench
300
29.3
36.3
8.0
29.0
32.0
37.0
MM-SafetyBench
390
30.3
32.3
0.3
25.6
28.7
34.1
Table 2: Overall safety–helpfulness alignment. Joint Success (%, ↑), requiring S=3 and H≥2 in the same response. SIUO+HoliSafe pools SIUO with the HoliSafe subsets. Bold indicates the highest success ratio in each row, including ties.
Figure 5: Qualitative comparison on an ambiguous safety-adjacent request. Several baselines either resolve the missing context toward a specific interpretation or refuse broadly. Intent-Privilege OPSD instead preserves the unresolved surface, medium, and permission conditions and provides conditional guidance without assuming unsupported user intent.
Model
SIUO+HoliSafe
BeaverTails-V
MSS
MOSS
MM-Safety
Base
43.9
31.0
39.0
29.3
30.3
No-Privilege
46.1
39.3
36.5
30.0
26.1
Source-Privilege
41.7
35.5
38.0
12.6
13.3
Shuffled EG-Privilege
43.6
40.5
35.5
31.0
31.8
Ours
53.5
49.0
39.0
37.0
34.1
Ours w/o General-Task Anchors
50.0
49.7
37.0
36.6
34.8
Table 3: Mechanism ablations. Joint Success (%, ↑) for controlled variants of INTENT-PRIVILEGE OPSD. EG denotes Evidence-Grounded Privilege.
Model
MMStar
MME-RealWorld
Aggregate
Base
64.0
32.0
53.3
SPA-VL
59.7
30.8
50.0
VLGuard
47.2
25.1
39.8
TiS
56.6
31.0
48.0
SafeGRPO
62.3
34.1
52.9
Source-Privilege OPSD
53.0
30.5
45.5
Table 4: General multimodal capability after safety post-training. Results on MMStar, MME-RealWorld, and the aggregate evaluation metric. The w/o-anchor ablation isolates the role of general-task anchors in preserving ordinary multimodal capability.
Figure 6: Efficiency–performance trade-offs. (a) Average Joint Success across the five safety evaluation groups versus wall-clock training time; circle area is proportional to the amount of safety-specific alignment data used by each method. Red annotations report the relative reduction in safety-specific training data and the relative performance difference of Intent-Privilege OPSD against each baseline. (b) Average generated response length (bars) and average Joint Success (line); numbers above the bars report inference latency under the matched serving configuration. Lower resource cost and higher Joint Success are preferred.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Evaluation group
Base
Intent-Privilege OPSD
Gain (pp)
SIUO+HoliSafe
40.9
44.1
+3.2
BeaverTails-V
24.49
32.65
+8.16
MSSBench
40.0
44.0
+4.0
MOSSBench
30.6
34.0
+3.4
MM-SafetyBench
12.3
12.8
+0.5
Appendix
Table 5: Safety–helpfulness alignment on LLaVA. Joint Success (%, higher is better), following the main-text metric convention. Gains are absolute percentage-point differences from Base. Bold marks the higher of the two reported model scores.
Model
MMStar
MME-RealWorld
Aggregate
Base
32
38
33.9
Intent-Privilege OPSD
33
35
33.6
Ours w/o General-Task Anchors
27
28
27.3
Appendix
Table 6: General multimodal capability on LLaVA. Higher is better. Aggregate is the reported aggregate score. Bold marks the highest success ratio in each column.
Existing safety alignment methods for vision-language models usually modify the model behavior globally: once the safety parameters are trained or loaded, they participate in both unsafe and already-safe generations. This always-on intervention can unnecessarily perturb the model's original reasoning path and degrade general multimodal capabilities. We argue that safety alignment should be an on-demand intervention rather than a permanent modification to every decoding trajectory. To this end, we propose a streaming recognition and gated LoRA framework for intrinsic VLM safety. During autoregressive generation, a lightweight recognizer estimates whether the current pre-token generation state is safe or unsafe. Its output updates the LoRA gate for the following decoding step; otherwise, generation follows the frozen-backbone policy. The LoRA module is trained from unsafe prefixes, transition statements, and safe continuations, so that it learns to redirect unsafe generations back to safe responses after activation. Experiments across multiple safety and general-purpose benchmarks demonstrate the effectiveness of our method in post-alignment settings.
Caoyuan Ma, Tian Gu, Wenpu Liu +11
The University of Tokyo · Wuhan University · Shanghai AI Laboratory +4
Vision-language models (VLMs) are increasingly trained to generate structured outputs like points and bounding boxes that downstream interfaces, agents, and robots can act on, yet safety alignment of this output channel has not been systematically analyzed. We study visual grounding safety by repurposing three safety benchmarks spanning direct harm (VLSU), social bias (BBQ-V), and situational safety (Asimov-2.0) into 15,401 matched pairs of harmful requests that differ only in the requested output: a free-text answer (VQA) or a grounding (point or bounding box). Across five VLMs, models that refuse a harmful request posed as a question often comply when the same request asks for a grounding: averaged over models, grounding refusal trails VQA refusal by 31-59 percentage points, depending on the domain, and safety system prompts do not close this gap. We propose a fine-tuning approach that combines grounding-form refusals with capability grounding data and self-distilled benign data to counter over-refusal. For Qwen3-VL-8B and VisionReasoner-7B, it improves grounding refusal by 77-95 percentage points on VLSU and BBQ-V and by 64-85 points on the held-out Asimov-2.0 domain, while also improving VQA refusal, preserving grounding capability, and keeping over-refusal limited. Representation analysis shows that fine-tuning moves harmful requests toward each model's refusal direction, most strongly for grounding, while leaving benign requests near the harmless reference.
Erfan Shayegani, Kundan Krishna, Yue Dong +3
University of California, Riverside · Work done during internship at Apple. · Apple
Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering. However, these methods show inconsistent improvements across models and evaluation distributions, and the intervention into model behavior or internal states introduce safety-utility tradeoffs by over-refusal. Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety. We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not. Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing. Across Qwen3-VL and Kimi-VL, the proposed detector substantially reduces safety errors on the HoliSafe benchmark and resoundingly generalizes to out-of-distribution safety benchmarks featuring different safety patterns, including MISHard and MM-SafetyBench. The success of the proposed router-logit detector also suggests a broader perspective on model internals: rather than focusing only on manipulating internal components to steer behavior, simply reading naturally emerging signals and linking them to an external safety mechanism can provide a simple, effective, and non-intrusive complement to existing safety interventions.