cs.IRAug 3, 2025

ChEmbed: Enhancing Chemical Literature Search Through Domain-Specific Text Embeddings

Authors: Ali Shiraee Kasmaee, Mohammad Khodadad, Mahdi Astaraki, Mohammad Arshi Saloot, Nicholas Sherck, Hamidreza Mahyar, Soheila Samiee

Organizations: Department of Computational Science and Engineering, McMaster University, Canada · BASF Canada Inc., Canada · BASF Corporation, USA

Abstract

Retrieval-Augmented Generation (RAG) systems in chemistry heavily depend on accurate and relevant retrieval of chemical literature. However, general-purpose text embedding models frequently fail to adequately represent complex chemical terminologies, resulting in suboptimal retrieval quality. Existing embedding models for chemistry are outdated, and none is tailored to chemical literature retrieval, leaving a substantial performance gap. To address this challenge, we introduce ChEmbed, the first purpose-built family of domain-adapted text embedding models engineered for chemical literature retrieval. These models are fine-tuned via contrastive learning on a dataset comprising chemistry-specific text from the PubChem, Semantic Scholar, and ChemRxiv corpora. To create effective training data, we employ large language models to synthetically generate queries, resulting in approximately 1.7 million high-quality query-passage pairs. Additionally, we augment the tokenizer by adding 900 chemically specialized tokens to previously unused slots, which reduces the fragmentation of chemical entities, such as IUPAC names. ChEmbed also maintains an 8192-token context length, enabling retrieval of longer passages than many open-source embedding models allow. Evaluated on our newly introduced ChemRxiv Retrieval benchmark, ChEmbed outperforms state-of-the-art general embedding models, raising MRR@10 from 0.781 to 0.882 (+10.1 pp). It also substantially outperforms domain-specific embedding models such as Chemical-BERT, improving MRR@10 from 0.096 to 0.882. A role-based retrieval analysis using PubChem descriptions and ChEBI annotations shows that the improvement extends to chemical-role queries. ChEmbed represents a practical, lightweight, and reproducible embedding solution that effectively improves chemical literature retrieval.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Chunk Twice, Embed Once: A Systematic Study of Segmentation and Representation Trade-offs in Chemistry-Aware Retrieval-Augmented Generation

    Jun 13, 2025Mahmoud Amiri, Thomas BocklitzChemistryChunk

  2. Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural Language

    Aug 4, 2026David Ming Segura, Jeremy Goumaz, Joshua W. Sin +2Molecular Representation LearningLanguage Modeling

  3. MolE-RAG: Molecular Structure-Enhanced Retrieval-Augmented Generation for Chemistry

    Jun 4, 2026Joey Chan, Wonbin Kweon, Ashley Shin +4Molecular Property PredictionChemical Structures