cs.LGSep 13, 2026

Robust small-molecule identification from incomplete, degraded, and inconsistent spectra using multimodal mixed-condition training

Authors: Bowen GaoLei ZhuYiying WangWenjie Yu

Abstract

Reliable small-molecule identification often requires complementary evidence from multiple spectroscopic measurements. In practice, however, spectra may be unavailable, degraded by measurement-related variations, or even incorrectly associated with a sample, thereby hindering accurate molecular identification. Herein, we propose a multimodal mixed-condition training strategy that accommodates missing, degraded, and mismatched measurements for small-molecule structure identification. The strategy incorporates chemical and spectroscopic knowledge through predefined missing-input configurations, modality-specific spectral perturbations, and chemically informed spectrum replacements. Models were trained on 635,441 samples comprising mass spectrometry (MS), infrared (IR), and nuclear magnetic resonance (NMR) simulated spectra from the Multimodal Spectroscopic Dataset (MSSD). They were then systematically evaluated on 79,462 held-out samples across 30 views designed to represent variations in spectra. A controlled comparison of complete-input and mixed-condition training under concatenation and mixture-of-experts (MoE) fusion showed that the training strategy was the principal source of improvement. For MoE, mixed-condition training increased the mean reciprocal rank (MRR) by 6.08% (from 0.9203 to 0.9763) and the top-1 molecular identification rate by 7.67% (from 89.50% to 96.36%). Notably, under single-modality inputs, IR MRR increased 2.15-fold (from 0.4337 to 0.9307), while MS MRR increased 2.31-fold (from 0.3711 to 0.8575). With the proposed strategy, complete-input performance remained high, while sample-level mismatch detection also improved. Together, these results highlight the potential of multimodal mixed-condition training for practical molecular identification by explicitly addressing incomplete, degraded, and mismatched measurements encountered in real-world analysis.

Explore similar work

May 11, 2026eess.IV

SpecX: A Large-Scale Benchmark for Multi-Modal Spectroscopy and Cross-Paradigm Evaluation

Existing spectral benchmarks are limited in scale, modality alignment, and evaluation scope, and typically focus on either specialized models or multimodal language models (MLLMs). We introduce SpecX, a large-scale benchmark for multi-modal spectroscopy with cross-paradigm evaluation. SpecX contains 1.7M molecules with diverse spectral modalities, including NMR (1H, 13C, HSQC), IR, MS,UV,Raman and FL, and is organized into three tiers: a large-scale dataset for pretraining, an aligned multi-spectral subset for benchmarking, and a high-quality experimental subset for evaluation. SpecX supports a range of tasks such as molecular elucidation, spectrum simulation, and spectral understanding, and enables unified evaluation across both specialized spectral models and MLLMs. Experiments show that specialized models excel at signal-level modeling, while MLLMs exhibit strengths in high-level reasoning but lack precise spectral grounding. SpecX establishes a unified benchmark for spectral intelligence and highlights the need for spectrum-native foundation models.
Chengrui Xiang, Tengfei Ma, Yujie Chen +3
Jul 22, 2026physics.chem-ph

Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data

Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply structure-resolving NMR signals for multimodal learning, we construct \textbf{QM9SPIN}, a DFT-derived dataset comprising diverse 1D and 2D spectra, including J-coupling, DEPT experiments, and explicit spin--spin interactions. On this foundation, we introduce \textbf{SpectroMol}, a spectrum-to-structure model that proposes chemically valid molecular hypotheses conditioned on multimodal spectral inputs. Complementarily, we develop \textbf{MS-Mol2Mol}, a high-resolution mass-constrained molecular generator that integrates molecular formula, exact mass, and degree of unsaturation within a conditional generative prior trained on 400 million molecules, ensuring global compositional consistency and chemically realistic refinement. The integrated system achieves 93.8% top-1 accuracy on the simulated benchmark, adapts effectively from simulated to experimental spectra with limited experimental fine-tuning, and further improves experimental predictions through mass-guided refinement, establishing a scalable route toward automated, data-driven organic structure elucidation.
Chengchun Liu, Zhiyuan Yan, Li Yuan +5
Dec 17, 2025physics.chem-ph

NMIRacle: Multi-modal Generative Molecular Elucidation from IR and NMR Spectra

Molecular structure elucidation from spectroscopic data is a long-standing challenge in Chemistry, traditionally requiring expert interpretation. We introduce NMIRacle, a two-stage generative framework that builds upon recent paradigms in AI-driven spectroscopy with minimal assumptions. In the first stage, NMIRacle learns to reconstruct molecular structures from count-aware fragment representations, capturing both fragment identities and their occurrences. In the second stage, a spectral encoder maps input spectra (IR, 1H-NMR, 13C-NMR) into a latent embedding used to condition the pre-trained generator, which is fine-tuned for direct spectra-to-molecule generation. This formulation bridges fragment-level chemical modeling with spectral evidence, yielding accurate molecular predictions. Empirical results demonstrate that NMIRacle outperforms existing baselines on molecular elucidation, while maintaining robust performance across increasing levels of molecular complexity.
Federico Ottomano, Yingzhen Li, Alex M. Ganose