A Large Scale Investigation of Scaling Limits in Chemical Language Models
Organizations: BioSystems Engineering and Control Lab · Wadhwani School of Data Science and AI, IIT Madras · The Centre for Integrative Biology and Systems medicinE (IBSE) · Chandar Research Lab · Mila – Quebec AI Institute · IIT Madras Zanzibar · Polytechnique Montréal · Canada CIFAR AI Chair
Abstract
Chemical Language Models (CLMs) are increasingly used in de novo drug design, driven by recent growth in model scale, compute, and dataset size. However, the relationship between design choices, training dynamics, and downstream generation quality remains poorly understood. We present a compute-controlled scaling study of CLMs comprising more than 30,000 experiments across molecular representations (SMILES, SELFIES, SAFE), tokenizations (atom-level and byte-pair encoding), model scales (0.5M-1B parameters), leakage-controlled datasets (MOSES, ChEMBL, PubChem, ZINC-22), and architectures (decoder-only and encoder-decoder). By fitting IsoFLOP profiles, we establish clear scaling trends in pretraining loss, but find that these improvements do not translate into comparable gains in goal-directed molecular design. Layer-wise probing and sparse autoencoder analysis reveal continued development of chemical representations: chemical syntax saturates early, while semantic properties emerge more slowly and become increasingly accessible with further training and model scale. These representational gains coexist with diminishing improvements in goal-directed generation under the evaluated protocols and oracle budgets. Our resulting suite of models, NovoMolGen, achieves state-of-the-art results, outperforming prior CLMs and specialized generative models in goal-directed molecular generation across drug discovery tasks. These findings expose a disconnect between chemical representation learning and downstream molecular design, motivating the development of pretraining paradigms that more directly learn chemical semantics.
Figures & tables
Appendix figures & tables40 assets
Supplementary material from the paper’s appendix.
Appendix
| Corpus | Molecules | Median heavy atoms | Median SMILES tokens | Character |
|---|---|---|---|---|
| MOSES | 1.9M | 21 | 28 | Narrow, drug-like, restricted atoms |
| ChEMBL | 1.9M | 27 | 48 | Bioactive, literature-derived |
| PubChem | 100M | 24 | 47 | Broad drug-like (reference corpus) |
| ZINC-22 | 1.5B (sampled) | 29 | 52 | Make-on-demand, largest coverage |
| Axis | Variants | Model sizes |
|---|---|---|
| Reference | PubChem, SMILES, atom-level, Qwen3 | 0.5M, 5M, 32M, 157M, 300M, 1B |
| Representation | SELFIES, SAFE | 0.5M, 5M, 32M |
| Tokenisation | Byte-pair encoding | 0.5M, 5M, 32M, 157M, 300M, 1B |
| Architecture | T5Gemma (encoder–decoder) | 0.5M, 5M, 32M, 157M, 300M, 1B |
| Pretraining corpus | MOSES, ChEMBL, ZINC-22 | 0.5M, 5M, 32M, 157M, 300M, 1B |
| Components | NovoMolGen-32M | NovoMolGen-157M | NovoMolGen-300M |
|---|---|---|---|
| Attention Heads | 8 | 10 | 12 |
| Hidden Layers | 12 | 24 | 32 |
| Hidden Size | 512 | 640 | 768 |
| Intermediate Size | 1024 | 2560 | 3072 |
| Hyperparameter | Value | Description |
|---|---|---|
| Dictionary Multiplier | Expansion factor ( features) | |
| Sparsity ( ) | Top- active features allowed | |
| Optimizer | Adam | Standard configuration |
| Learning Rate | Linear warmup over first 2,500 steps | |
| Batch Size (SAE) | 4,096 | Activation vectors per optimization step |
| Training Steps | 50,000 | Total updates ( tokens) |
| Functional Group | SAE latents | Neurons | |
|---|---|---|---|
| Carbonyl | 0.999 | 0.982 | |
| Carbonyl (non-acid) | 0.997 | 0.965 | |
| Bicyclic | 0.997 | 0.984 | |
| Ether | 0.995 | 0.968 | |
| Secondary amine | 0.993 | 0.963 | |
| Amide | 0.993 | 0.960 |
| Budget (FLOPs) | Evaluated sizes | (tokens) | |
|---|---|---|---|
| 0.5M, 5M, 32M | 2.5M | ||
| 5M, 32M, 157M | 12M | ||
| 5M, 32M, 157M | 55M | ||
| 32M, 157M, 300M | 130M | ||
| 157M, 300M, 1B | 700M |
| Axis | Setting | Bits / molecule ( ) | Bits / token | Tokens / molecule |
| Representation | SMILES (atomwise) | 48.1 | 1.02 | 47 |
| SELFIES (atomwise) | 50.9 | 0.33 | 155 | |
| SAFE (atomwise) | 50.2 | 0.95 | 53 | |
| Tokenisation | SMILES (atom-level) | 48.1 | 1.02 | 47 |
| SMILES (byte-pair) | 46.4 | 1.86 | 25 | |
| Corpus (SMILES, atomwise) | MOSES | 25.6 | 0.91 | 28 |
| Model size | 0.5M | 5M | 32M | 157M | 300M | 1B |
|---|---|---|---|---|---|---|
| Bits / molecule ( ) | 71.5 | 54.1 | 48.1 | 46.0 | 45.5 | 45.0 |
| Split | 0.5M | 5M | 32M | 157M | 300M | 1B |
|---|---|---|---|---|---|---|
| Random | 71.5 | 54.1 | 48.1 | 46.0 | 45.5 | 45.0 |
| Scaffold | 72.9 | 55.2 | 49.1 | 46.9 | 46.4 | 45.9 |
| Molecular weight | 422 | 243 | 168 | 143 | 132 | 126 |
| Maximum dissimilarity | 172 | 92 | 68 | 62 | 60 | 59 |
| Max-dissim. / random | 2.4 | 1.7 | 1.4 | 1.35 | 1.32 | 1.3 |
| Axis | Setting | OOD gap (%) |
|---|---|---|
| Representation | SMILES (atomwise) | 41 |
| SELFIES (atomwise) | 57 | |
| SAFE (atomwise) | 50 | |
| Tokenisation | Atom-level | 41 |
| Byte-pair | 32 | |
| Architecture | Qwen3 (decoder) | 41 |
| Axis | Setting | Size | Valid ( ) | IntDiv ( ) | Novelty ( ) | FCD ( ) | SNN ( ) | Frag ( ) | Scaf ( ) |
|---|---|---|---|---|---|---|---|---|---|
| Ref. size | PubChem, SMILES | 0.5M | 0.310 | 0.892 | 0.989 | 3.92 | 0.386 | 0.948 | 0.469 |
| 5M | 0.901 | 0.894 | 0.970 | 2.84 | 0.402 | 0.892 | 0.722 | ||
| 32M | 0.961 | 0.884 | 0.979 | 5.12 | 0.395 | 0.901 | 0.314 | ||
| 157M | 0.968 | 0.883 | 0.980 | 5.05 | 0.398 | 0.905 | 0.330 | ||
| 300M | 0.972 | 0.882 | 0.980 | 4.95 | 0.399 | 0.908 | 0.335 | ||
| 1B | 0.978 | 0.882 | 0.981 | 4.82 | 0.401 | 0.910 | 0.340 |
| Representation (32M) | Random | Scaffold | Molecular weight | Maximum dissimilarity |
|---|---|---|---|---|
| SMILES (atomwise) | 5.12 | 13.9 | 23.4 | 29.0 |
| SELFIES (atomwise) | 2.43 | 12.6 | 22.9 | 27.4 |
| SAFE (atomwise) | 2.93 | 12.5 | 22.8 | 26.2 |
| Representation | Size | FCD ( ) | SNN ( ) | Scaf ( ) | |||
|---|---|---|---|---|---|---|---|
| Test | TestSF | Test | TestSF | Test | TestSF | ||
| SMILES (atomwise) | 0.5M | 3.92 | 15.03 | 0.386 | 0.333 | 0.469 | 0.064 |
| 5M | 2.84 | 13.17 | 0.402 | 0.347 | 0.722 | 0.063 | |
| 32M | 5.12 | 13.91 | 0.395 | 0.345 | 0.314 | 0.028 | |
| SELFIES (atomwise) | 32M | 2.43 | 12.61 | 0.400 | 0.345 | 0.662 | 0.041 |
| SAFE (atomwise) | 32M | 2.93 | 12.53 | 0.383 | 0.336 | 0.595 | 0.092 |
| Oracle | -RAG | REINVENT | NovoMolGen‑32M (AtomWise) | NovoMolGen‑157M (AtomWise) | NovoMolGen‑300M (AtomWise) |
|---|---|---|---|---|---|
| albuterol_similarity | 0.977 (0.002) | 0.968 (0.002) | |||
| amlodipine_mpo | 0.749 (0.019) | 0.732 (0.024) | |||
| celecoxib_rediscovery | 0.778 (0.007) | 0.904 (0.022) | |||
| drd2 | 0.973 (0.001) | 0.983 (0.002) | |||
| deco_hop | 0.992 (0.000) | 0.945 (0.007) | |||
| fexofenadine_mpo | 0.856 (0.016) | 0.803 (0.018) |
| Oracle | -RAG | REINVENT | NovoMolGen-32M (BPE) | NovoMolGen-157M (BPE) | NovoMolGen-300M (BPE) |
|---|---|---|---|---|---|
| albuterol_similarity | 0.977 (0.002) | 0.981 (0.001) | |||
| amlodipine_mpo | 0.749 (0.019) | 0.739 (0.021) | |||
| celecoxib_rediscovery | 0.872 (0.061) | 0.890 (0.018) | |||
| drd2 | 0.987 (0.001) | 0.967 (0.000) | |||
| deco_hop | 0.992 (0.000) | 0.945 (0.007) | |||
| fexofenadine_mpo | 0.856 (0.016) | 0.837 (0.022) |
| Method | parp1 | fa7 | 5ht1b | braf | jak2 | Sum |
|---|---|---|---|---|---|---|
| REINVENT | ||||||
| MOOD | ||||||
| GEAM | ||||||
| -RAG | ||||||
| NovoMolGen-0.5M | ||||||
| NovoMolGen-5M |
| Method | parp1 | fa7 | 5ht1b | braf | jak2 |
|---|---|---|---|---|---|
| REINVENT | |||||
| MORLD | |||||
| HierVAE | |||||
| FREED | |||||
| MOOD | |||||
| GEAM |