cs.LGAug 19, 2025

A Large Scale Investigation of Scaling Limits in Chemical Language Models

Authors: Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar

Organizations: BioSystems Engineering and Control Lab · Wadhwani School of Data Science and AI, IIT Madras · The Centre for Integrative Biology and Systems medicinE (IBSE) · Chandar Research Lab · Mila – Quebec AI Institute · IIT Madras Zanzibar · Polytechnique Montréal · Canada CIFAR AI Chair

Abstract

Chemical Language Models (CLMs) are increasingly used in de novo drug design, driven by recent growth in model scale, compute, and dataset size. However, the relationship between design choices, training dynamics, and downstream generation quality remains poorly understood. We present a compute-controlled scaling study of CLMs comprising more than 30,000 experiments across molecular representations (SMILES, SELFIES, SAFE), tokenizations (atom-level and byte-pair encoding), model scales (0.5M-1B parameters), leakage-controlled datasets (MOSES, ChEMBL, PubChem, ZINC-22), and architectures (decoder-only and encoder-decoder). By fitting IsoFLOP profiles, we establish clear scaling trends in pretraining loss, but find that these improvements do not translate into comparable gains in goal-directed molecular design. Layer-wise probing and sparse autoencoder analysis reveal continued development of chemical representations: chemical syntax saturates early, while semantic properties emerge more slowly and become increasingly accessible with further training and model scale. These representational gains coexist with diminishing improvements in goal-directed generation under the evaluated protocols and oracle budgets. Our resulting suite of models, NovoMolGen, achieves state-of-the-art results, outperforming prior CLMs and specialized generative models in goal-directed molecular generation across drug discovery tasks. These findings expose a disconnect between chemical representation learning and downstream molecular design, motivating the development of pretraining paradigms that more directly learn chemical semantics.

Figures & tables

Appendix figures & tables40 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Probing Chemical Language Models: Effects of Pre-training and Fine-tuning

    Jul 2, 2026Anna Karnysheva, Dietrich Klakow, Ji-Ung LeeMoleculenetMolecules

  2. What Does a Chemical Language Model Know About Molecules?

    Jun 22, 2026Christian Kenneth, Etowah Adams, Liam Bai +1MoleculesLanguage Modeling