Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment
Authors: Haotian Deng, Wenbin Xing, Gang Xu, Tao He, Jinkai Zheng, Chun Li, Zheng Zhu, Ming Li
Organizations: Southern University of Science and Technology · Sun Yat-sen University · Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ) · University of Electronic Science and Technology of China · Hangzhou Dianzi University · Shenzhen MSU-BIT University · GigaAI · The Chinese University of Hong Kong (Shenzhen)
Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time. They may also sacrifice helpfulness by directly refusing requests that could be answered safely. In this paper, we propose Intent-Privilege On-Policy Self-Distillation (OPSD), which leverages evidence-grounded intent as privileged supervision during training to help VLMs recognize implicit risks and provide safe, useful responses instead of blanket refusals. OPSD distills a teacher's intent-conditioned preferences over responses into a student using a single rollout per prompt; the student then responds without intent annotations or an additional safety module. With only 1,447 safety-specific examples - 95% fewer than standard preference datasets - OPSD reduces training time by 5x relative to multi-rollout GRPO-style training and average inference length by 7%. It attains the highest ratio for joint safety-helpfulness success, which measures the proportion of responses that are both safe and helpful, across all five evaluation groups. Remarkably, on pooled SIUO+HoliSafe, this success ratio rises from 43.9% to 53.5%. These results show that training-time intent supervision can improve both safety and helpfulness while substantially reducing data, training, and inference costs.
Figures & tables
Figure 1: Motivation of Intent-Privilege OPSD. It achieves comprehensive data, training, and inference efficiency. A detailed analysis of defense collapse on implicit risks is provided in Sec. D .
Context
S↑
H↑
S=3↑
H≥2↑
Success ↑
FSR
No privilege
1.633
1.850
25.0
75.0
25.0
12.9
Source Privilege
1.767
1.683
11.7
66.7
11.7
45.0
Evidence-Grounded Privilege
1.917
2.283
41.7
95.0
41.7
12.1
Shuffled Evidence-Grounded Privilege
1.450
1.867
23.3
76.7
23.3
17.1
Table 1: Frozen-teacher Comparison. Safety S∈[−3,3] evaluates boundary handling and evidence calibration; Helpfulness H∈[0,3] measures useful assistance. Success requires S=3 and H≥2 . The last four columns are percentages. FSR means First-sentence refusal. New uses the intent annotation z produced by the pipeline in Sec. 3.2 ; Shuffled New uses another input’s annotation. Refusal is the sampled first-sentence frequency.
Figure 2: Qualitative comparison under different privileged intent conditions. Source Privilege (a) injects speculative threat assumptions, leading to a defensive blanket refusal. Conversely, Evidence-Grounded Privilege (b) preserves uncertainty and provides calibrated guidance.
Figure 3: Evidence-grounded intent re-annotation pipeline. Modality-isolated evidence is combined by a joint builder and audited against the original inputs to produce intent annotations z . Accepted annotations provide the Evidence-Grounded Privilege used in both frozen-teacher evaluation and OPSD.
Figure 4: Intent-Privilege OPSD. (a) Data Efficiency: achieves a 95% reduction using 1,447 training examples; (b) Training Efficiency: uses a single rollout to slash compute time by 5×; (c) Inference Efficiency: operates without extra components, reducing average inference length by 7%.
Dataset
N
Base
SPA-VL
VLGuard
TiS
SafeGRPO
Ours
SIUO+HoliSafe
767
43.9
50.3
40.8
45.9
45.4
53.5
BeaverTails-V
200
31.0
44.0
8.5
32.5
30.5
49.0
MSSBench
400
39.0
35.5
32.5
35.7
38.8
39.0
MOSSBench
300
29.3
36.3
8.0
29.0
32.0
37.0
MM-SafetyBench
390
30.3
32.3
0.3
25.6
28.7
34.1
Table 2: Overall safety–helpfulness alignment. Joint Success (%, ↑), requiring S=3 and H≥2 in the same response. SIUO+HoliSafe pools SIUO with the HoliSafe subsets. Bold indicates the highest success ratio in each row, including ties.
Figure 5: Qualitative comparison on an ambiguous safety-adjacent request. Several baselines either resolve the missing context toward a specific interpretation or refuse broadly. Intent-Privilege OPSD instead preserves the unresolved surface, medium, and permission conditions and provides conditional guidance without assuming unsupported user intent.
Model
SIUO+HoliSafe
BeaverTails-V
MSS
MOSS
MM-Safety
Base
43.9
31.0
39.0
29.3
30.3
No-Privilege
46.1
39.3
36.5
30.0
26.1
Source-Privilege
41.7
35.5
38.0
12.6
13.3
Shuffled EG-Privilege
43.6
40.5
35.5
31.0
31.8
Ours
53.5
49.0
39.0
37.0
34.1
Ours w/o General-Task Anchors
50.0
49.7
37.0
36.6
34.8
Table 3: Mechanism ablations. Joint Success (%, ↑) for controlled variants of INTENT-PRIVILEGE OPSD. EG denotes Evidence-Grounded Privilege.
Model
MMStar
MME-RealWorld
Aggregate
Base
64.0
32.0
53.3
SPA-VL
59.7
30.8
50.0
VLGuard
47.2
25.1
39.8
TiS
56.6
31.0
48.0
SafeGRPO
62.3
34.1
52.9
Source-Privilege OPSD
53.0
30.5
45.5
Table 4: General multimodal capability after safety post-training. Results on MMStar, MME-RealWorld, and the aggregate evaluation metric. The w/o-anchor ablation isolates the role of general-task anchors in preserving ordinary multimodal capability.
Figure 6: Efficiency–performance trade-offs. (a) Average Joint Success across the five safety evaluation groups versus wall-clock training time; circle area is proportional to the amount of safety-specific alignment data used by each method. Red annotations report the relative reduction in safety-specific training data and the relative performance difference of Intent-Privilege OPSD against each baseline. (b) Average generated response length (bars) and average Joint Success (line); numbers above the bars report inference latency under the matched serving configuration. Lower resource cost and higher Joint Success are preferred.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Evaluation group
Base
Intent-Privilege OPSD
Gain (pp)
SIUO+HoliSafe
40.9
44.1
+3.2
BeaverTails-V
24.49
32.65
+8.16
MSSBench
40.0
44.0
+4.0
MOSSBench
30.6
34.0
+3.4
MM-SafetyBench
12.3
12.8
+0.5
Appendix
Table 5: Safety–helpfulness alignment on LLaVA. Joint Success (%, higher is better), following the main-text metric convention. Gains are absolute percentage-point differences from Base. Bold marks the higher of the two reported model scores.
Model
MMStar
MME-RealWorld
Aggregate
Base
32
38
33.9
Intent-Privilege OPSD
33
35
33.6
Ours w/o General-Task Anchors
27
28
27.3
Appendix
Table 6: General multimodal capability on LLaVA. Higher is better. Aggregate is the reported aggregate score. Bold marks the highest success ratio in each column.