Molecular multimodal models support diverse understanding and generation tasks but may introduce safety vulnerabilities when handling hazardous molecules. In this work, We reveal substantial jailbreak vulnerabilities under both text-only and graph-conditioned settings. Our analysis further shows that safety robustness must hold across input modalities while balancing safety, over-refusal, and utility. To address these challenges, we construct SafeMolBench, a molecular multimodal safety-alignment benchmark with 3702 samples covering 618 unique hazardous molecules and safe molecular tasks, organized into hazardous-harmful, hazardous-allowed, and utility-replay subsets to support unified training and evaluation of safety, over-refusal, and utility. Based on SafeMolBench, we propose SafeMol, a parameter-efficient safety alignment framework that jointly optimizes lightweight modules across text-only and graph-conditioned inputs, uses MMD for distribution-level representation alignment to reduce modality-induced discrepancies, and explicitly models molecular hazardousness and harmful operational intent. Experiments on SafeMolBench show that SafeMol reduces attack success by several tens of percentage points while largely maintaining low over-refusal and preserving molecular-task utility.
Figures & tables
Model
Input
ASR
ORR
Utility
Attempt ↓
Process Success ↓
Progress ↓
ORR-A ↓
ORR-U ↓
BLEU-2 ↑
BLEU-4 ↑
ROUGE-1 ↑
ROUGE-2 ↑
ROUGE-L ↑
GPT-4o
w/o Graph
64.00%
36.00%
4.570
0.00%
14.75%
0.0246
0.0041
0.1056
0.0063
0.0671
Llama-3.2-1B-Instruct
w/o Graph
0.00%
0.00%
0.000
37.25%
100.00%
0.0013
∼ 0.0000
0.0437
0.0000
0.0387
Omni-Mol v2
w/o Graph
50.00%
38.00%
3.850
2.00%
0.00%
0.4391
0.3328
0.4177
0.1977
0.3562
w/ Graph
92.00%
66.00%
6.260
5.00%
0.00%
0.5847
0.4743
0.5700
0.3324
0.5094
ReactXT
w/o Graph
83.00%
27.00%
3.820
–
0.82%
0.3510
0.2550
0.3990
0.1560
0.3310
Table 1: Safety, over-refusal, and utility performance of language models and molecular multimodal models on SafeMolBench. We report results under text-only (w/o Graph) and graph-conditioned (w/ Graph) settings, while language models are evaluated only under the text-only setting.
Figure 1: Overview of the SafeMolBench construction pipeline.
Figure 2: Overview of SafeMol.
Model
Input
ASR
ORR
Utility
Attempt ↓
Process Success ↓
Progress ↓
ORR-A ↓
ORR-U ↓
BLEU-2 ↑
BLEU-4 ↑
ROUGE-1 ↑
ROUGE-2 ↑
ROUGE-L ↑
Language Models
GPT-4o
w/o Graph
64.00%
36.00%
4.570
0.00%
14.75%
0.0246
0.0041
0.1056
0.0063
0.0671
Llama-3.2-1B-Instruct
w/o Graph
0.00%
0.00%
0.000
37.25%
100.00%
0.0013
0.0000
0.0437
0.0000
0.0387
Omni-Mol
Omni-Mol v2
w/o Graph
50.00%
38.00%
3.850
2.00%
0.00%
0.4391
0.3328
0.4177
0.1977
0.3562
Table 2: Main results on SafeMolBench under text-only (w/o Graph) and graph-conditioned (w/ Graph) settings. Results are averaged over five seeds, with significance assessed at p<0.001 .
Training Configuration
Input
ASR
ORR
Utility
Attempt ↓
Process Success ↓
Progress ↓
ORR-A ↓
ORR-U ↓
BLEU-2 ↑
BLEU-4 ↑
ROUGE-1 ↑
ROUGE-2 ↑
ROUGE-L ↑
Omni-Mol v2
w/o Graph
50.00%
38.00%
3.850
2.00%
0.00%
0.4391
0.3328
0.4177
0.1977
0.3562
w/ Graph
92.00%
66.00%
6.260
5.00%
0.00%
0.5847
0.4743
0.5700
0.3324
0.5094
Text-only Align.
w/o Graph
13.00%
10.00%
0.990
3.00%
4.10%
0.5222
0.4016
0.5069
0.2477
0.4343
w/ Graph
94.00%
58.00%
5.860
4.00%
0.00%
0.5808
0.4684
0.5604
0.3212
0.4989
Graph-only Align.
w/o Graph
1.00%
0.00%
0.120
2.00%
74.59%
0.0383
0.0239
0.0693
0.0153
0.0575
Table 3: Data ablation results under different training-data configurations. Results are averaged over five random seeds, with statistical significance assessed at p<0.001 .
Model / Train Data
Input
ASR
ORR
Utility
Attempt ↓
Process Success ↓
Progress ↓
ORR-A ↓
ORR-U ↓
BLEU-2 ↑
BLEU-4 ↑
ROUGE-1 ↑
ROUGE-2 ↑
ROUGE-L ↑
Omni-Mol v2
w/o Graph
50.00%
38.00%
3.850
2.00%
0.00%
0.4391
0.3328
0.4177
0.1977
0.3562
w/ Graph
92.00%
66.00%
6.260
5.00%
0.00%
0.5847
0.4743
0.5700
0.3324
0.5094
Safety-LoRA only
w/o Graph
0.00%
0.00%
0.000
2.00%
18.03%
0.4649
0.3614
0.4524
0.2226
0.3876
w/ Graph
0.00%
0.00%
0.030
0.00%
14.75%
0.4975
0.3979
0.4748
0.2628
0.4202
ours w/o MMD
w/o Graph
2.00%
2.00%
0.170
5.00%
15.57%
0.4696
0.3631
0.4547
0.2236
0.3924
Table 4: Component ablation results under different input settings. Results are averaged over five random seeds, with statistical significance assessed at p<0.001 .
Figure 3: Effect of training data scale on the safety, over-refusal, and utility of SafeMol.
Task
Experimental Procedure (BLEU-4)
Forward Reaction Prediction (Exa)
Retrosynthesis (Exa)
Reagent Prediction (Exa)
Catalyst Prediction (Exa)
Quantum Mechanics Property Prediction (Ave MAE)
Molecular Captioning (BLEU-4)
Description Q&A (ROUGE-L)
Retention
87.79%
96.99%
95%
88.78%
101.21%
100%
91.95%
100.41%
Table 5: Retention of SafeMol on 16 general molecular tasks. Results are averaged over five random seeds, with statistical significance assessed at p<0.001 .
Figure 4: Comparison of modality hidden state gaps between Omni-Mol v2 and SafeMol. The dashed diagonal line indicates equal gaps between the two models.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Representative jailbreak examples under the text-only (a) and graph-conditioned (b) attack settings. The graph-conditioned setting includes molecular graphs of relevant chemical components as additional inputs.
Field
Description
molecule_id
Unique molecule identifier with category prefix (exp_, drg_, or cwp_).
chemical_name
Common or systematic chemical name of the target molecule.
canonical_smiles
Canonical SMILES of the target (hazardous) molecule.
inchi_key
InChIKey used for structural identity and de-duplication.
molecule_category
Hazard category: explosive, controlled drug, or chemical weapon/poison.
category_source
Regulatory source of the category label.
Appendix
Table 6: Molecular information fields retained for the target molecules.
Category
Source
Before Filtering
After Filtering
explosive
ATF + SMILES-prompting
88
63
controlled drug
DEA + SMILES-prompting
314
291
chemical_weapon_poison
OPCW + EPA + SMILES-prompting
372
264
Appendix
Table 7: Detailed statistics of the 618 hazardous molecules retained in SafeMolBench.
Figure 6: Representative samples from the H, A, and U subsets of SafeMolBench.
Split
H
A
U
Total
Train
1034
1034
984
3052
Val
102
102
124
328
Test
100
100
122
322
Appendix
Table 8: Benchmark splits of SafeMolBench.
Evaluation Aspect
Metric
Definition
Data Type
ASR
Attempt Rate ↓
An attack is considered successful if the LLM judge determines that the model does not refuse the request.
H
Process Success Rate ↓
An attack is considered successful only if the LLM judge determines that the generated response is correct.
H
Progress Score ↓
The LLM judge evaluates whether the response provides information that makes progress toward the requested objective and assigns a score from 0 to 10.
H
ORR
ORR-A ↓
Over-Refusal Rate, defined as the number of incorrectly refused samples divided by the total number of samples that should be answered, multiplied by 100%. Refusal is determined by an LLM judge, with A samples as the denominator.
A
ORR-U ↓
Over-Refusal Rate, defined as the number of incorrectly refused samples divided by the total number of samples that should be answered, multiplied by 100%. Refusal is determined by an LLM judge, with U samples as the denominator.
U
Utility
BLEU-2 ↑
Measures 1-gram and 2-gram matches between the model-generated response and the reference answer.
U
Appendix
Table 9: Definitions of the evaluation metrics used in SafeMolBench.
Figure 7: Prompts used for gpt-4o-based evaluation. A shared safety block is applied to all judging prompts, followed by metric-specific instructions for Attempt Rate, Process Success Rate, Progress Score, ORR-A, and ORR-U.
Method
GPU Type
GPU Hours
Omni-Mol v2
NVIDIA A100 80G
576
Safety-LoRA
NVIDIA V100 32GB
2.76
SafeMol w/o MMD
NVIDIA V100 32GB
2.2
SafeMol w/o Aux.
NVIDIA V100 32GB
2.79
SafeMol
NVIDIA V100 32GB
3.58
Appendix
Table 10: GPU hours comparison of SafeMol, its ablation variants, and the original Omni-Mol model.
Model
Input
ASR
ORR
Utility
Attempt ↓
Process Success ↓
Progress ↓
ORR-A ↓
ORR-U ↓
BLEU-2 ↑
BLEU-4 ↑
ROUGE-1 ↑
ROUGE-2 ↑
ROUGE-L ↑
Omni-Mol v2
w/o Graph
14.00%
12.00%
1.4500
2.00%
0.82%
0.2074
0.1545
0.1911
0.0902
0.1665
w/ Graph
85.00%
60.00%
5.7700
4.00%
0.00%
0.5677
0.4538
0.5429
0.3020
0.4830
SafeMol
w/o Graph
1.00%
1.00%
0.1600
1.00%
30.33%
0.3319
0.2526
0.3200
0.1476
0.2714
w/ Graph
4.00%
2.00%
0.2900
1.00%
24.59%
0.4087
0.3239
0.3932
0.2077
0.3447
Appendix
Table 11: Performance comparison between Omni-Mol v2 and SafeMol under OOD experimental-procedure generation instructions across safety, over-refusal, and utility metrics.
Table 19
Model
Input
Utility
BLEU-2 ↑
BLEU-4 ↑
ROUGE-1 ↑
ROUGE-2 ↑
ROUGE-L ↑
Llama-3.2-1B-Instruct
w/o Graph
0.0013
∼ 0.0000
0.0437
0.0000
0.0387
Omni-Mol v2
w/o Graph
0.4391
0.3328
0.4177
0.1977
0.3562
w/ Graph
0.5847
0.4743
0.5700
0.3324
0.5094
w/o Graph
0.5549
0.4319
0.5370
0.2701
0.4615
ours
w/ Graph
0.5737
0.4569
0.5558
0.3057
0.4903
Appendix
Table 14: Utility performance evaluated on non-refused U samples under different input settings, excluding over-refused samples.
Model
Type
#Par
Exa
BLEU
Lev
RDK
MAC
Mor
Val
Forward Reaction Prediction Task
DeepSeekV3
ICL
685B
0.35
0.939
12.76
0.719
0.823
0.68
1.00
Llama2
SL
6.7B
0.01
0.804
29.95
0.499
0.649
0.41
1.00
Mol-Ins
SL
6.7B
0.05
0.654
27.26
0.313
0.509
0.26
1.00
HIGHT
SL
6.7B
0.29
0.935
16.69
0.774
0.618
0.57
1.00
InstructMol
SL
6.7B
0.54
0.967
10.85
0.776
0.878
0.74
1.00
Appendix
Table 15: Performance comparison of Omni-Mol with baseline models across multiple molecular tasks.
While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related annotations and rely on unidimensional metrics, hindering comprehensive robustness evaluation. To address this, we propose MME-Safety, a rigorously verified benchmark featuring a unique four-dimensional annotation schema that categorizes risk scenarios, harm severity, and modality-specific stealth levels. Furthermore, we introduce a hierarchical evaluation framework to assess fundamental response reliability, actual risk exposure, and the structural integrity of defensive behaviors. Extensive zero-shot evaluations across 17 state-of-the-art MLLMs provide a comprehensive safety profile of current multimodal systems. Our analysis systematically investigates cross-modal input configurations and uncovers safety implications associated with Chain-of-Thought (CoT) reasoning. These multifaceted findings underscore the urgent need for robust, reasoning-aware safety alignment in the multimodal landscape.
Despite remarkable capability in multi-modal understanding, deploying Multi-modal Large Language Models (MLLMs) in open-ended conversational scenarios introduces safety risks that remain poorly addressed by existing alignment methods. Unlike simple malicious visual question and answer (VQA) pairs , multi-turn interactions enable adversaries to incrementally reconstruct harmful intent across dialogues, progressively bypassing safety constraints in ways that are difficult to detect at any individual turn. Meanwhile, conventional reinforcement learning from human feedback (RLHF) approaches are unsuitable for this situation: designed primarily for VQA tasks, they neither capture cross-turn risk dynamics nor scale efficiently without costly manual preference annotation. To close this gap, we introduce \textbf{MINT-Safe}, an open-source visual multi-turn training dataset comprising 11,270 multi-image dialogues and 500 refusal VQA pairs, constructed via multi-agent interaction with text-to-image (T2I) tool-call augmentation. Building on MINT-Safe, we propose \textbf{TAD-Align}, a dialogue safety alignment framework centered on a turn-aware dual-objective reward function. Rather than treating all dialogue turns uniformly, TAD-Align leverages rollout-based safety score variance to dynamically identify turns where the model exhibits inconsistent safety behavior, and adaptively up-weights these turns during optimization. Experiments on Qwen2.5-VL-7B-Instruct and LLaVA-NeXT-7B demonstrate reductions of over 10% in Attack Success Rate (ASR), alongside improvements of at least 8% in harmlessness and 13% in helpfulness on multi-modal multi-turn safety benchmarks, while preserving general model capabilities.
Han Zhu, Jiale Chen, Chengkun Cai +8
Hong Kong University of Science and Technology · University College London · AISpeech +1
Even modern AI models often remain vulnerable to multimodal queries in which harmful intent is embedded in images. A widely used approach for safety alignment is training with extensive multimodal safety datasets, but the costs of data curation and training are often prohibitive. To mitigate these costs, inference-time alignment has recently been explored, but they often lack generalizability across diverse multimodal jailbreaks and still incur notable overhead due to extra forward passes for response refinement or heavy pre-deployment calibration procedures. Here, we identify insufficient visual attention to safety-critical image regions as one of the key causes of multimodal safety failures. Building on this insight, we propose Multimodal Risk-Adaptive Steering (MoRAS), which enhances safety-critical visual attention via concise visual contexts for accurate multimodal risk assessment. This risk signal enables risk-adaptive steering for direct refusals, reducing inference overhead while remaining generalizable across diverse multimodal jailbreaks. Notably, MoRAS requires only a small calibration set to estimate multimodal risk, substantially reducing pre-deployment overhead. We conduct various empirical validations across multiple benchmarks and MLLM backbones, and observe that the proposed MoRAS consistently mitigates jailbreaks, preserves utility, and reduces computational overhead compared to state-of-the-art inference-time defenses.