ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models
Authors: Tobia Poppi, Silvia Cappelletti, Samuele Poppi, Marcella Cornia, Lorenzo Baraldi, Diego Garcia-Olano, Rita Cucchiara
Organizations: University of Modena and Reggio Emilia, Modena, Italy. · University of Pisa, Pisa, Italy. · MBZUAI, Abu Dhabi, United Arab Emirates. · Meta Superintelligence Labs, Menlo Park, United States.
Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content at scale, existing datasets pair safe real samples with generated counterparts, but label every generated sample unsafe, even when one modality is individually safe. To address this, we introduce ShieldCLIP, the first framework to condition safety alignment on the observed safety state of each modality rather than the origin of a sample, preserving safe content while redirecting only what is unsafe. We also introduce ViSUv2, a 195k-quadruplet dataset with independent per-modality safety labels across 578 concepts and 28 categories. Using these labels, ShieldCLIP defines a four-way conditional objective beyond pair-level supervision: safe content is anchored, unsafe modalities are redirected to their safe counterparts, mixed pairs update only the unsafe branch, and coherence is enforced when both are unsafe. We evaluate ShieldCLIP on cross-modal retrieval, text-to-image generation with Stable Diffusion v1.4 and SDXL, and image-to-text generation with LLaVA. Across these settings, ShieldCLIP consistently reduces harmful outputs over prior safety-aligned encoders and strong mitigation baselines, while preserving the utility of the original embedding space. Extensive ablation studies further show that both modality-specific supervision and the selective alignment objective contribute to these gains. Source code, trained models, and ViSUv2 (under a controlled-access protocol) will be made publicly available at https://aimagelab.github.io/ShieldCLIP/.
Figures & tables
Figure 1: Examples of real (safe) pairs and their generated (safe or unsafe) counterparts with harmful tags.
Figure 2: ViSUv2 dataset generation. The pipeline used to generate unsafe text and images from real captions, followed by independent safety labeling.
Figure 3: Mapping of some CoPro granular NSFW concepts to the 28 NSFW categories in our taxonomy.
Figure 4: The ShieldCLIP architecture. Left: Real and generated text/image pairs are processed by trainable encoders to produce embeddings, while frozen oracle encoders ( T0,V0 ) generate corresponding anchor embeddings. Right: The conditional training objectives. Real and generated-safe samples (label 0) are preserved by aligning them with their anchors within designated “safe zones”. Generated-unsafe samples (label 1) are redirected according to their modality-specific safety configuration, with dedicated objectives for unsafe, mixed, and jointly unsafe pairs.
Text Div.
Text Harm. (%)
Img Div.
Img Harm. (%)
# Text
# Img
# Cat
Labels
Vendi-S ↑
Self-BLEU ↓
LLaMA-3.1 ↑
GPT-3.5 ↑
Vendi-S ↑
NN-Q16 ↑
GPT-4 ↑
SneakyPrompt
181
-
-
✗
6.1
0.355
62.5
59.5
-
-
-
Ring-A-Bell
1.1k (80)
-
-
✗
14.6
0.279
55.0
63.8
-
-
-
I2P
4.7k
-
7
✗
99.9
0.395
12.1
22.4
-
-
-
MMA-Diffusion
3k (1k)
80 (80)
-
✗
12.8
0.392
79.0
54.2
6.1
75.4
0.0
ViSU
165k (165k)
165k (165k)
20
✗
44.7
0.304
77.7
85.2
28.4
70.8
83.3
Table 1: Comparison of text and image diversity and harmfulness across ViSUv2 and existing safety datasets. Values in parentheses indicate the number of safe captions or images excluded from the score computation.
I2P
ViSUv2
Model
Hate
Harass
Viol
S-Harm
Sex
Shock
Ill Act
Avg
Hate
Harass
Viol
S-Harm
Sex
Shock
Ill Act
Avg
SD v1.4
40.5
32.4
42.0
40.0
24.1
50.8
36.7
38.1
27.5
26.3
30.2
27.0
17.4
21.9
27.8
25.4
SLD-Strong ( Schramowski et al., 2023 )
13.5
11.5
15.2
8.9
5.4
19.3
8.5
11.8
4.3
4.3
5.7
3.7
3.4
3.4
3.9
4.1
ESD ( Gandikota et al., 2023 )
39.1
32.3
43.1
40.2
21.5
48.1
34.6
37.0
25.6
21.4
29.4
26.7
15.1
19.4
26.4
23.4
SPM ( Lyu et al., 2024 )
23.0
21.0
34.0
25.7
15.2
37.4
22.7
25.6
15.5
13.5
18.7
16.1
9.0
11.1
17.0
14.4
UCE ( Gandikota et al., 2024 )
32.1
25.2
27.8
18.6
15.2
30.5
20.6
24.3
16.1
14.5
16.9
15.6
9.9
12.3
15.2
14.4
Table 2: Rate of generated harmful images using unsafe textual prompts from I2P ( Schramowski et al., 2023 ) and the proposed ViSUv2 dataset. Results are computed with both SD v1.4 and SDXL as text-to-image generators, combining predictions from NudeNet and Q16. Avg is the harmful rate across all individual generations regardless of its categories.
Figure 5: Qualitative examples generated with SD v1.4, SDXL, ShieldCLIP, and competing methods using unsafe prompts from I2P and ViSUv2.
I2P
ViSUv2
Model
Hate
Harass
Viol
S-Harm
Sex
Shock
Ill Act
Avg
Hate
Harass
Viol
S-Harm
Sex
Shock
Ill Act
Avg
SD v1.4
16.1
16.6
33.4
19.7
44.7
28.0
15.2
24.8
23.9
19.9
23.9
17.7
22.6
16.0
18.6
20.4
SLD-Strong ( Schramowski et al., 2023 )
3.6
4.6
13.0
3.5
13.6
8.6
5.2
7.4
5.8
3.6
6.6
3.0
5.9
3.3
4.5
4.7
Safe-CLIP ( Poppi et al., 2024 )
7.6
7.5
16.9
7.4
25.1
10.3
5.3
11.5
3.0
3.4
2.8
2.4
1.7
1.5
2.4
2.5
SafeR-CLIP ( Yousaf et al., 2026a )
7.6
7.2
11.1
6.9
13.8
9.0
5.5
8.7
2.9
3.7
3.3
3.1
2.7
2.8
3.5
3.1
SafetyDPO ( Liu et al., 2025 )
3.3
5.6
10.3
2.5
11.5
4.5
4.7
6.1
2.2
1.2
2.6
1.3
1.4
1.1
2.2
1.7
Table 3: Rate of generated harmful images using unsafe textual prompts from I2P ( Schramowski et al., 2023 ) and the proposed ViSUv2 dataset. Results are computed with both SD v1.4 and SDXL as text-to-image generators, using LlavaGuard ( Helff et al., 2025 ) as safety classifier.
Figure 6: Qualitative examples of image-to-text generation with the original LLaVA model, Safe-CLIP, and ShieldCLIP, using real NSFW images from different sources as input.
Figure 7: Qualitative examples of text-to-image (left) and image-to-text (right) retrieval using unsafe text or image queries, comparing ShieldCLIP with the original CLIP model and Safe-CLIP.
Retrieval (T2I)
Retrieval (I2T)
Generation
Generation Utility
Model
λ
% Harmful Content ( ↓ )
% Harmful Content ( ↓ )
% Harmful Content ( ↓ )
FID ( ↓ )
CLIP-Sim ( ↑ )
Loss formulation
Only cosine losses
–
0.0
100.0
2.8
80.3
0.096
Loss-weight trade-off
High preservation
(0.5,0.5,0.1,0.1,0.5)
18.9
34.5
3.3
15.1
0.258
Medium preservation
(0.25,0.25,0.1,0.1,0.25)
11.4
33.9
2.4
15.5
0.255
Table 4: Ablation of the loss formulation and safety-quality trade-off. We report harmful rates (%, ↓ ) for retrieval and generation, together with FID and CLIP-Sim to measure generation quality and semantic alignment. For the weight sweep, λ=(λreal,λsafe,λredir,λmix,λcoh) . Safety metrics are evaluated on ViSUv2 using T2I retrieval, I2T retrieval, and SD v1.4 generation, while FID and CLIP-Sim are computed on 30k samples from COCO.
Model
CIFAR-10
CIFAR-100
SUN-397
Food-101
Caltech-101
Imagenette
CLIP
94.5
60.7
63.6
88.1
74.8
99.7
Safe-CLIP ( Poppi et al., 2024 )
87.6
61.1
54.2
74.8
64.4
98.5
SafeR-CLIP ( Yousaf et al., 2026a )
92.1
61.5
58.4
77.2
74.3
98.8
ShieldCLIP (Ours)
91.9
64.0
61.8
79.5
74.8
99.3
Table 5: Preservation analysis of CLIP performance on zero-shot classification, reported in terms of top-1 accuracy ( ↑ ).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Prompt used to generate an NSFW caption from a safe caption.
Figure 9: Prompt used to classify a generated caption as safe (a) or NSFW (b).
Figure 10: Prompt used for image generator selection, detecting if a caption implies masculine nudity.
Figure 11: Prompt used to classify the CoPro concepts into ViSUv2 categories.
Figure 12: Examples from the ViSUv2 dataset. Each sample shows a real safe caption and image (green border) paired with a generated caption and image (red border for unsafe). We also report the granular tag used for conditioning and the associated high-level categories.
Figure 13: Prompt used for GPT-3.5 Turbo to classify NSFW captions as safe (0) or NSFW (1) based on the presence of offensive or unsafe content.
ViSU
Model
Hate
Harass
Viol
S-Harm
Sex
Shock
Ill Act
Avg
SD v1.4
26.2
18.1
31.3
19.9
23.5
29.9
23.5
24.6
SLD-Strong ( Schramowski et al., 2023 )
4.5
3.5
6.0
4.9
5.2
5.8
3.9
4.8
SalUn ( Fan et al., 2024 )
19.3
9.8
23.0
15.7
8.3
19.6
17.9
16.3
ESD ( Gandikota et al., 2023 )
23.4
15.5
30.4
19.3
20.9
27.7
21.8
22.7
SPM ( Lyu et al., 2024 )
15.8
9.7
19.3
11.4
13.1
17.9
13.7
14.4
Appendix
Table 6: Rate of generated harmful images using unsafe textual prompts from ViSU ( Poppi et al., 2024 ) , combining predictions from NudeNet and Q16 classifiers.
Figure 14: Additional qualitative examples generated with SD v1.4, SDXL, ShieldCLIP, and competing methods using unsafe prompts from I2P and ViSUv2.
Figure 15: Additional qualitative examples of image-to-text generation with the original LLaVA model, Safe-CLIP, and ShieldCLIP, using real NSFW images from different sources as input.
Figure 16: Representative failure cases for ShieldCLIP in text-to-image generation with SD v1.4 and SDXL on unsafe prompts from I2P and ViSUv2, shown alongside competing mitigation methods.
Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remains unclear whether and how training data composition directly drives model output safety or by other factors. We shed light on this question by isolating this variable: we train the same text-to-image model on datasets that differ \emph{only} in their fraction of unsafe images (0% to 9.6%), across several dataset scales (100K to 8M). Then we generate images with the resulting models, and evaluate them with four independent safety classifiers. Output unsafety rises monotonically from 16.6% at 0% contamination to 25.5% at 5%. A factorial design reveals that the \emph{proportion}, not the absolute count, of unsafe training images is the operative variable. The 16.6% irreducible baseline at zero contamination implicates the other components, e.g. frozen text encoder, as a residual safety risk -- confirmed by a text encoder ablation showing that SafeCLIP reduces this floor to 9.6%, while the dose-response effect persists across all three encoders tested. Critically, no quality degradation in terms of FID, CLIPscore and ImageReward accompanies safety filtering. These results establish that data curation and text encoder safety are complementary and independently effective interventions. At the same time, the remaining level of unsafety poses questions for future research about emerging capabilities and compositionality.
General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversarial nature of content and AI safety. We present Yuvion VL, a family of multimodal large language models purpose-built for content and AI safety, with both instruction-tuned and reasoning-oriented variants. Yuvion VL addresses this gap by treating safety as an inherently adversarial and multimodal problem and designing the entire pipeline around adversarial robustness. For data construction, we develop an automated pipeline integrating adversarial-aware data synthesis with multi-stage quality control, producing large-scale, high-quality multimodal samples augmented with domain knowledge and reasoning annotations. For training, we adopt a three-stage pipeline that includes continued pretraining for risk-concept cross-modal alignment, instruct post-training for production-grade safety tasks, and reasoning post-training for enhanced interpretability and performance in complex tasks. We further introduce Confuse-then-Contrast Fine-Tuning, a contrastive framework that mines model-specific confusions and constructs multi-image contrastive groups to enforce explicit discrimination of fine-grained visual-semantic elements, enabling the model to distinguish between visually similar cases with different safety implications in adversarial safety tasks. To support rigorous evaluation, we further introduce Yuvion VL RiskEval (YVRE), a collection of benchmarks covering diverse open and internal evaluations, with a focus on content and AI safety, adversarial robustness, and real-world capability requirements. Experiments show that Yuvion VL-32B achieves industry-leading safety performance, surpassing comparably sized open-source models and best closed-source commercial models, while maintaining comparable general capabilities.
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning. SafeCap trains a policy model to first generate a safety-relevant image caption and then produce a final answer; the caption is further optimized by whether it enables a frozen LLM to reach a safety-aligned decision. This caption-mediated objective encourages the policy to expose visual cues relevant to safe response generation rather than relying solely on direct refusal supervision. Across five multimodal safety benchmarks and six vision-utility benchmarks, SafeCap substantially improves aggregate safety performance under its intended DirectCap protocol, with gains of 3.7-19.0 points in safety average across four model settings while maintaining comparable or improved vision utility. Under controlled comparisons on matched backbones and data, SafeCap outperforms safety SFT, DPO, and SafeGRPO, demonstrating the effectiveness of caption-mediated reinforcement learning for multimodal safety alignment.
Caoyuan Ma, Wenpu Liu, Weichu Xie +12
The University of Tokyo · Peking University · Wuhan University +4