SafeMol: Dual-Modality Safety Alignment for Molecular Multimodal Models
Organizations: Beihang University · China CITIC Bank · Stable AI · Department of Computer Science and Technology, Tsinghua University
Abstract
Molecular multimodal models support diverse understanding and generation tasks but may introduce safety vulnerabilities when handling hazardous molecules. In this work, We reveal substantial jailbreak vulnerabilities under both text-only and graph-conditioned settings. Our analysis further shows that safety robustness must hold across input modalities while balancing safety, over-refusal, and utility. To address these challenges, we construct SafeMolBench, a molecular multimodal safety-alignment benchmark with 3702 samples covering 618 unique hazardous molecules and safe molecular tasks, organized into hazardous-harmful, hazardous-allowed, and utility-replay subsets to support unified training and evaluation of safety, over-refusal, and utility. Based on SafeMolBench, we propose SafeMol, a parameter-efficient safety alignment framework that jointly optimizes lightweight modules across text-only and graph-conditioned inputs, uses MMD for distribution-level representation alignment to reduce modality-induced discrepancies, and explicitly models molecular hazardousness and harmful operational intent. Experiments on SafeMolBench show that SafeMol reduces attack success by several tens of percentage points while largely maintaining low over-refusal and preserving molecular-task utility.
Figures & tables
| Model | Input | ASR | ORR | Utility | |||||||
| Attempt | Process Success | Progress | ORR-A | ORR-U | BLEU-2 | BLEU-4 | ROUGE-1 | ROUGE-2 | ROUGE-L | ||
| GPT-4o | w/o Graph | 64.00% | 36.00% | 4.570 | 0.00% | 14.75% | 0.0246 | 0.0041 | 0.1056 | 0.0063 | 0.0671 |
| Llama-3.2-1B-Instruct | w/o Graph | 0.00% | 0.00% | 0.000 | 37.25% | 100.00% | 0.0013 | 0.0000 | 0.0437 | 0.0000 | 0.0387 |
| Omni-Mol v2 | w/o Graph | 50.00% | 38.00% | 3.850 | 2.00% | 0.00% | 0.4391 | 0.3328 | 0.4177 | 0.1977 | 0.3562 |
| w/ Graph | 92.00% | 66.00% | 6.260 | 5.00% | 0.00% | 0.5847 | 0.4743 | 0.5700 | 0.3324 | 0.5094 | |
| ReactXT | w/o Graph | 83.00% | 27.00% | 3.820 | – | 0.82% | 0.3510 | 0.2550 | 0.3990 | 0.1560 | 0.3310 |
| Model | Input | ASR | ORR | Utility | |||||||
| Attempt | Process Success | Progress | ORR-A | ORR-U | BLEU-2 | BLEU-4 | ROUGE-1 | ROUGE-2 | ROUGE-L | ||
| Language Models | |||||||||||
| GPT-4o | w/o Graph | 64.00% | 36.00% | 4.570 | 0.00% | 14.75% | 0.0246 | 0.0041 | 0.1056 | 0.0063 | 0.0671 |
| Llama-3.2-1B-Instruct | w/o Graph | 0.00% | 0.00% | 0.000 | 37.25% | 100.00% | 0.0013 | 0.0000 | 0.0437 | 0.0000 | 0.0387 |
| Omni-Mol | |||||||||||
| Omni-Mol v2 | w/o Graph | 50.00% | 38.00% | 3.850 | 2.00% | 0.00% | 0.4391 | 0.3328 | 0.4177 | 0.1977 | 0.3562 |
| Training Configuration | Input | ASR | ORR | Utility | |||||||
| Attempt | Process Success | Progress | ORR-A | ORR-U | BLEU-2 | BLEU-4 | ROUGE-1 | ROUGE-2 | ROUGE-L | ||
| Omni-Mol v2 | w/o Graph | 50.00% | 38.00% | 3.850 | 2.00% | 0.00% | 0.4391 | 0.3328 | 0.4177 | 0.1977 | 0.3562 |
| w/ Graph | 92.00% | 66.00% | 6.260 | 5.00% | 0.00% | 0.5847 | 0.4743 | 0.5700 | 0.3324 | 0.5094 | |
| Text-only Align. | w/o Graph | 13.00% | 10.00% | 0.990 | 3.00% | 4.10% | 0.5222 | 0.4016 | 0.5069 | 0.2477 | 0.4343 |
| w/ Graph | 94.00% | 58.00% | 5.860 | 4.00% | 0.00% | 0.5808 | 0.4684 | 0.5604 | 0.3212 | 0.4989 | |
| Graph-only Align. | w/o Graph | 1.00% | 0.00% | 0.120 | 2.00% | 74.59% | 0.0383 | 0.0239 | 0.0693 | 0.0153 | 0.0575 |
| Model / Train Data | Input | ASR | ORR | Utility | |||||||
| Attempt | Process Success | Progress | ORR-A | ORR-U | BLEU-2 | BLEU-4 | ROUGE-1 | ROUGE-2 | ROUGE-L | ||
| Omni-Mol v2 | w/o Graph | 50.00% | 38.00% | 3.850 | 2.00% | 0.00% | 0.4391 | 0.3328 | 0.4177 | 0.1977 | 0.3562 |
| w/ Graph | 92.00% | 66.00% | 6.260 | 5.00% | 0.00% | 0.5847 | 0.4743 | 0.5700 | 0.3324 | 0.5094 | |
| Safety-LoRA only | w/o Graph | 0.00% | 0.00% | 0.000 | 2.00% | 18.03% | 0.4649 | 0.3614 | 0.4524 | 0.2226 | 0.3876 |
| w/ Graph | 0.00% | 0.00% | 0.030 | 0.00% | 14.75% | 0.4975 | 0.3979 | 0.4748 | 0.2628 | 0.4202 | |
| ours w/o MMD | w/o Graph | 2.00% | 2.00% | 0.170 | 5.00% | 15.57% | 0.4696 | 0.3631 | 0.4547 | 0.2236 | 0.3924 |
| Task | Experimental Procedure (BLEU-4) | Forward Reaction Prediction (Exa) | Retrosynthesis (Exa) | Reagent Prediction (Exa) | Catalyst Prediction (Exa) | Quantum Mechanics Property Prediction (Ave MAE) | Molecular Captioning (BLEU-4) | Description Q&A (ROUGE-L) |
| Retention | 87.79% | 96.99% | 95% | 88.78% | 101.21% | 100% | 91.95% | 100.41% |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Field | Description |
| molecule_id | Unique molecule identifier with category prefix (exp_, drg_, or cwp_). |
| chemical_name | Common or systematic chemical name of the target molecule. |
| canonical_smiles | Canonical SMILES of the target (hazardous) molecule. |
| inchi_key | InChIKey used for structural identity and de-duplication. |
| molecule_category | Hazard category: explosive, controlled drug, or chemical weapon/poison. |
| category_source | Regulatory source of the category label. |
| Category | Source | Before Filtering | After Filtering |
| explosive | ATF + SMILES-prompting | 88 | 63 |
| controlled drug | DEA + SMILES-prompting | 314 | 291 |
| chemical_weapon_poison | OPCW + EPA + SMILES-prompting | 372 | 264 |
| Split | H | A | U | Total |
| Train | 1034 | 1034 | 984 | 3052 |
| Val | 102 | 102 | 124 | 328 |
| Test | 100 | 100 | 122 | 322 |
| Evaluation Aspect | Metric | Definition | Data Type |
| ASR | Attempt Rate | An attack is considered successful if the LLM judge determines that the model does not refuse the request. | H |
| Process Success Rate | An attack is considered successful only if the LLM judge determines that the generated response is correct. | H | |
| Progress Score | The LLM judge evaluates whether the response provides information that makes progress toward the requested objective and assigns a score from 0 to 10. | H | |
| ORR | ORR-A | Over-Refusal Rate, defined as the number of incorrectly refused samples divided by the total number of samples that should be answered, multiplied by 100%. Refusal is determined by an LLM judge, with A samples as the denominator. | A |
| ORR-U | Over-Refusal Rate, defined as the number of incorrectly refused samples divided by the total number of samples that should be answered, multiplied by 100%. Refusal is determined by an LLM judge, with U samples as the denominator. | U | |
| Utility | BLEU-2 | Measures 1-gram and 2-gram matches between the model-generated response and the reference answer. | U |
| Method | GPU Type | GPU Hours |
| Omni-Mol v2 | NVIDIA A100 80G | 576 |
| Safety-LoRA | NVIDIA V100 32GB | 2.76 |
| SafeMol w/o MMD | NVIDIA V100 32GB | 2.2 |
| SafeMol w/o Aux. | NVIDIA V100 32GB | 2.79 |
| SafeMol | NVIDIA V100 32GB | 3.58 |
| Model | Input | ASR | ORR | Utility | |||||||
| Attempt | Process Success | Progress | ORR-A | ORR-U | BLEU-2 | BLEU-4 | ROUGE-1 | ROUGE-2 | ROUGE-L | ||
| Omni-Mol v2 | w/o Graph | 14.00% | 12.00% | 1.4500 | 2.00% | 0.82% | 0.2074 | 0.1545 | 0.1911 | 0.0902 | 0.1665 |
| w/ Graph | 85.00% | 60.00% | 5.7700 | 4.00% | 0.00% | 0.5677 | 0.4538 | 0.5429 | 0.3020 | 0.4830 | |
| SafeMol | w/o Graph | 1.00% | 1.00% | 0.1600 | 1.00% | 30.33% | 0.3319 | 0.2526 | 0.3200 | 0.1476 | 0.2714 |
| w/ Graph | 4.00% | 2.00% | 0.2900 | 1.00% | 24.59% | 0.4087 | 0.3239 | 0.3932 | 0.2077 | 0.3447 | |
| Model | Input | Utility | ||||
| BLEU-2 | BLEU-4 | ROUGE-1 | ROUGE-2 | ROUGE-L | ||
| Llama-3.2-1B-Instruct | w/o Graph | 0.0013 | 0.0000 | 0.0437 | 0.0000 | 0.0387 |
| Omni-Mol v2 | w/o Graph | 0.4391 | 0.3328 | 0.4177 | 0.1977 | 0.3562 |
| w/ Graph | 0.5847 | 0.4743 | 0.5700 | 0.3324 | 0.5094 | |
| w/o Graph | 0.5549 | 0.4319 | 0.5370 | 0.2701 | 0.4615 | |
| ours | w/ Graph | 0.5737 | 0.4569 | 0.5558 | 0.3057 | 0.4903 |
| Model | Type | #Par | Exa | BLEU | Lev | RDK | MAC | Mor | Val |
| Forward Reaction Prediction Task | |||||||||
| DeepSeekV3 | ICL | 685B | 0.35 | 0.939 | 12.76 | 0.719 | 0.823 | 0.68 | 1.00 |
| Llama2 | SL | 6.7B | 0.01 | 0.804 | 29.95 | 0.499 | 0.649 | 0.41 | 1.00 |
| Mol-Ins | SL | 6.7B | 0.05 | 0.654 | 27.26 | 0.313 | 0.509 | 0.26 | 1.00 |
| HIGHT | SL | 6.7B | 0.29 | 0.935 | 16.69 | 0.774 | 0.618 | 0.57 | 1.00 |
| InstructMol | SL | 6.7B | 0.54 | 0.967 | 10.85 | 0.776 | 0.878 | 0.74 | 1.00 |