cs.LGJun 16, 2026

MOLAR: Learning Multimodal Molecular Representations from Noisy Labels

Authors: Yingxu WangKunyu ZhangNan YinYu LiEran Segal

Organizations: Department of Machine Learning, Mohamed bin Zayed University of Artificial Intelligence, AI Diyafah St, 7909, Abu Dhabi, United Arab Emirates · Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong, China · International College, Zhengzhou University, Daxue North Road, 450000, Henan, China · The Education University of Hong Kong, Hong Kong, China · Department of Molecular Cell Biology, Weizmann Institute of Science, Rehovot, Israel

Abstract

Motivation: Noisy labels are a common challenge in molecular property prediction because molecular annotations are often obtained from assays, curated databases, or weak annotation pipelines rather than directly observed clean biological states. Treating recorded labels as reliable supervision can cause models to memorize corrupted observations and learn misleading molecular evidence. In multimodal molecular representation learning, this issue can be amplified by graph-text fusion or alignment, which may propagate label-induced errors across modalities. Results: We propose MOLAR, a noise-aware framework for learning multimodal molecular representations from noisy labels. MOLAR separates latent clean-property inference from recorded-label observation: graph and text views contribute residual evidence to a clean-property distribution, and a categorical label-observation channel maps this distribution to recorded labels for training. This formulation derives posterior label reliability and modality-specific molecular evidence from the model. Experiments on naturally noisy molecular benchmarks and controlled label-flipping benchmarks show that MOLAR consistently outperforms representative baselines. Visualization analyses further show that MOLAR provides interpretable reliability and modality-evidence diagnostics.

Explore similar work

CardsList