ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models
Authors: Tobia Poppi, Silvia Cappelletti, Samuele Poppi, Marcella Cornia, Lorenzo Baraldi, Diego Garcia-Olano, Rita Cucchiara
Organizations: University of Modena and Reggio Emilia, Modena, Italy. · University of Pisa, Pisa, Italy. · MBZUAI, Abu Dhabi, United Arab Emirates. · Meta Superintelligence Labs, Menlo Park, United States.
Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content at scale, existing datasets pair safe real samples with generated counterparts, but label every generated sample unsafe, even when one modality is individually safe. To address this, we introduce ShieldCLIP, the first framework to condition safety alignment on the observed safety state of each modality rather than the origin of a sample, preserving safe content while redirecting only what is unsafe. We also introduce ViSUv2, a 195k-quadruplet dataset with independent per-modality safety labels across 578 concepts and 28 categories. Using these labels, ShieldCLIP defines a four-way conditional objective beyond pair-level supervision: safe content is anchored, unsafe modalities are redirected to their safe counterparts, mixed pairs update only the unsafe branch, and coherence is enforced when both are unsafe. We evaluate ShieldCLIP on cross-modal retrieval, text-to-image generation with Stable Diffusion v1.4 and SDXL, and image-to-text generation with LLaVA. Across these settings, ShieldCLIP consistently reduces harmful outputs over prior safety-aligned encoders and strong mitigation baselines, while preserving the utility of the original embedding space. Extensive ablation studies further show that both modality-specific supervision and the selective alignment objective contribute to these gains. Source code, trained models, and ViSUv2 (under a controlled-access protocol) will be made publicly available at https://aimagelab.github.io/ShieldCLIP/.
Figures & tables
Figure 1: Examples of real (safe) pairs and their generated (safe or unsafe) counterparts with harmful tags.
Figure 2: ViSUv2 dataset generation. The pipeline used to generate unsafe text and images from real captions, followed by independent safety labeling.
Figure 3: Mapping of some CoPro granular NSFW concepts to the 28 NSFW categories in our taxonomy.
Figure 4: The ShieldCLIP architecture. Left: Real and generated text/image pairs are processed by trainable encoders to produce embeddings, while frozen oracle encoders ( T0,V0 ) generate corresponding anchor embeddings. Right: The conditional training objectives. Real and generated-safe samples (label 0) are preserved by aligning them with their anchors within designated “safe zones”. Generated-unsafe samples (label 1) are redirected according to their modality-specific safety configuration, with dedicated objectives for unsafe, mixed, and jointly unsafe pairs.
Text Div.
Text Harm. (%)
Img Div.
Img Harm. (%)
# Text
# Img
# Cat
Labels
Vendi-S ↑
Self-BLEU ↓
LLaMA-3.1 ↑
GPT-3.5 ↑
Vendi-S ↑
NN-Q16 ↑
GPT-4 ↑
SneakyPrompt
181
-
-
✗
6.1
0.355
62.5
59.5
-
-
-
Ring-A-Bell
1.1k (80)
-
-
✗
14.6
0.279
55.0
63.8
-
-
-
I2P
4.7k
-
7
✗
99.9
0.395
12.1
22.4
-
-
-
MMA-Diffusion
3k (1k)
80 (80)
-
✗
12.8
0.392
79.0
54.2
6.1
75.4
0.0
ViSU
165k (165k)
165k (165k)
20
✗
44.7
0.304
77.7
85.2
28.4
70.8
83.3
Table 1: Comparison of text and image diversity and harmfulness across ViSUv2 and existing safety datasets. Values in parentheses indicate the number of safe captions or images excluded from the score computation.
I2P
ViSUv2
Model
Hate
Harass
Viol
S-Harm
Sex
Shock
Ill Act
Avg
Hate
Harass
Viol
S-Harm
Sex
Shock
Ill Act
Avg
SD v1.4
40.5
32.4
42.0
40.0
24.1
50.8
36.7
38.1
27.5
26.3
30.2
27.0
17.4
21.9
27.8
25.4
SLD-Strong ( Schramowski et al., 2023 )
13.5
11.5
15.2
8.9
5.4
19.3
8.5
11.8
4.3
4.3
5.7
3.7
3.4
3.4
3.9
4.1
ESD ( Gandikota et al., 2023 )
39.1
32.3
43.1
40.2
21.5
48.1
34.6
37.0
25.6
21.4
29.4
26.7
15.1
19.4
26.4
23.4
SPM ( Lyu et al., 2024 )
23.0
21.0
34.0
25.7
15.2
37.4
22.7
25.6
15.5
13.5
18.7
16.1
9.0
11.1
17.0
14.4
UCE ( Gandikota et al., 2024 )
32.1
25.2
27.8
18.6
15.2
30.5
20.6
24.3
16.1
14.5
16.9
15.6
9.9
12.3
15.2
14.4
Table 2: Rate of generated harmful images using unsafe textual prompts from I2P ( Schramowski et al., 2023 ) and the proposed ViSUv2 dataset. Results are computed with both SD v1.4 and SDXL as text-to-image generators, combining predictions from NudeNet and Q16. Avg is the harmful rate across all individual generations regardless of its categories.
Figure 5: Qualitative examples generated with SD v1.4, SDXL, ShieldCLIP, and competing methods using unsafe prompts from I2P and ViSUv2.
I2P
ViSUv2
Model
Hate
Harass
Viol
S-Harm
Sex
Shock
Ill Act
Avg
Hate
Harass
Viol
S-Harm
Sex
Shock
Ill Act
Avg
SD v1.4
16.1
16.6
33.4
19.7
44.7
28.0
15.2
24.8
23.9
19.9
23.9
17.7
22.6
16.0
18.6
20.4
SLD-Strong ( Schramowski et al., 2023 )
3.6
4.6
13.0
3.5
13.6
8.6
5.2
7.4
5.8
3.6
6.6
3.0
5.9
3.3
4.5
4.7
Safe-CLIP ( Poppi et al., 2024 )
7.6
7.5
16.9
7.4
25.1
10.3
5.3
11.5
3.0
3.4
2.8
2.4
1.7
1.5
2.4
2.5
SafeR-CLIP ( Yousaf et al., 2026a )
7.6
7.2
11.1
6.9
13.8
9.0
5.5
8.7
2.9
3.7
3.3
3.1
2.7
2.8
3.5
3.1
SafetyDPO ( Liu et al., 2025 )
3.3
5.6
10.3
2.5
11.5
4.5
4.7
6.1
2.2
1.2
2.6
1.3
1.4
1.1
2.2
1.7
Table 3: Rate of generated harmful images using unsafe textual prompts from I2P ( Schramowski et al., 2023 ) and the proposed ViSUv2 dataset. Results are computed with both SD v1.4 and SDXL as text-to-image generators, using LlavaGuard ( Helff et al., 2025 ) as safety classifier.
Figure 6: Qualitative examples of image-to-text generation with the original LLaVA model, Safe-CLIP, and ShieldCLIP, using real NSFW images from different sources as input.
Figure 7: Qualitative examples of text-to-image (left) and image-to-text (right) retrieval using unsafe text or image queries, comparing ShieldCLIP with the original CLIP model and Safe-CLIP.
Retrieval (T2I)
Retrieval (I2T)
Generation
Generation Utility
Model
λ
% Harmful Content ( ↓ )
% Harmful Content ( ↓ )
% Harmful Content ( ↓ )
FID ( ↓ )
CLIP-Sim ( ↑ )
Loss formulation
Only cosine losses
–
0.0
100.0
2.8
80.3
0.096
Loss-weight trade-off
High preservation
(0.5,0.5,0.1,0.1,0.5)
18.9
34.5
3.3
15.1
0.258
Medium preservation
(0.25,0.25,0.1,0.1,0.25)
11.4
33.9
2.4
15.5
0.255
Table 4: Ablation of the loss formulation and safety-quality trade-off. We report harmful rates (%, ↓ ) for retrieval and generation, together with FID and CLIP-Sim to measure generation quality and semantic alignment. For the weight sweep, λ=(λreal,λsafe,λredir,λmix,λcoh) . Safety metrics are evaluated on ViSUv2 using T2I retrieval, I2T retrieval, and SD v1.4 generation, while FID and CLIP-Sim are computed on 30k samples from COCO.
Model
CIFAR-10
CIFAR-100
SUN-397
Food-101
Caltech-101
Imagenette
CLIP
94.5
60.7
63.6
88.1
74.8
99.7
Safe-CLIP ( Poppi et al., 2024 )
87.6
61.1
54.2
74.8
64.4
98.5
SafeR-CLIP ( Yousaf et al., 2026a )
92.1
61.5
58.4
77.2
74.3
98.8
ShieldCLIP (Ours)
91.9
64.0
61.8
79.5
74.8
99.3
Table 5: Preservation analysis of CLIP performance on zero-shot classification, reported in terms of top-1 accuracy ( ↑ ).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Prompt used to generate an NSFW caption from a safe caption.
Figure 9: Prompt used to classify a generated caption as safe (a) or NSFW (b).
Figure 10: Prompt used for image generator selection, detecting if a caption implies masculine nudity.
Figure 11: Prompt used to classify the CoPro concepts into ViSUv2 categories.
Figure 12: Examples from the ViSUv2 dataset. Each sample shows a real safe caption and image (green border) paired with a generated caption and image (red border for unsafe). We also report the granular tag used for conditioning and the associated high-level categories.
Figure 13: Prompt used for GPT-3.5 Turbo to classify NSFW captions as safe (0) or NSFW (1) based on the presence of offensive or unsafe content.
ViSU
Model
Hate
Harass
Viol
S-Harm
Sex
Shock
Ill Act
Avg
SD v1.4
26.2
18.1
31.3
19.9
23.5
29.9
23.5
24.6
SLD-Strong ( Schramowski et al., 2023 )
4.5
3.5
6.0
4.9
5.2
5.8
3.9
4.8
SalUn ( Fan et al., 2024 )
19.3
9.8
23.0
15.7
8.3
19.6
17.9
16.3
ESD ( Gandikota et al., 2023 )
23.4
15.5
30.4
19.3
20.9
27.7
21.8
22.7
SPM ( Lyu et al., 2024 )
15.8
9.7
19.3
11.4
13.1
17.9
13.7
14.4
Appendix
Table 6: Rate of generated harmful images using unsafe textual prompts from ViSU ( Poppi et al., 2024 ) , combining predictions from NudeNet and Q16 classifiers.
Figure 14: Additional qualitative examples generated with SD v1.4, SDXL, ShieldCLIP, and competing methods using unsafe prompts from I2P and ViSUv2.
Figure 15: Additional qualitative examples of image-to-text generation with the original LLaVA model, Safe-CLIP, and ShieldCLIP, using real NSFW images from different sources as input.
Figure 16: Representative failure cases for ShieldCLIP in text-to-image generation with SD v1.4 and SDXL on unsafe prompts from I2P and ViSUv2, shown alongside competing mitigation methods.