cs.LGSep 27, 2026

SafeMol: Dual-Modality Safety Alignment for Molecular Multimodal Models

Authors: Xinmiao Wang, Ruijie Wang, Menghui Wang, Jiawei Chen, Haoyue Deng, Ran Zhang, Xingxuan Zhang, Xiao Wang

Organizations: Beihang University · China CITIC Bank · Stable AI · Department of Computer Science and Technology, Tsinghua University

Abstract

Molecular multimodal models support diverse understanding and generation tasks but may introduce safety vulnerabilities when handling hazardous molecules. In this work, We reveal substantial jailbreak vulnerabilities under both text-only and graph-conditioned settings. Our analysis further shows that safety robustness must hold across input modalities while balancing safety, over-refusal, and utility. To address these challenges, we construct SafeMolBench, a molecular multimodal safety-alignment benchmark with 3702 samples covering 618 unique hazardous molecules and safe molecular tasks, organized into hazardous-harmful, hazardous-allowed, and utility-replay subsets to support unified training and evaluation of safety, over-refusal, and utility. Based on SafeMolBench, we propose SafeMol, a parameter-efficient safety alignment framework that jointly optimizes lightweight modules across text-only and graph-conditioned inputs, uses MMD for distribution-level representation alignment to reduce modality-induced discrepancies, and explicitly models molecular hazardousness and harmful operational intent. Experiments on SafeMolBench show that SafeMol reduces attack success by several tens of percentage points while largely maintaining low over-refusal and preserving molecular-task utility.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Towards Multi-modal Multi-turn Safety: From Agentic Interaction to Strategic Alignment

    Jan 8, 2026Han Zhu, Jiale Chen, Chengkun Cai +8Safety AlignmentMultimodal Large Language Models

  2. Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment

    Oct 15, 2025Jonghyun Park, Minhyuk Seo, Chaewon Yeo +1Safety AlignmentInference-Time Defense