Zatom-2: Multitask Pretraining on Atomistic Data for Generative Modeling across Domains
Organizations: University of Cambridge · Yale University · MIT · Lawrence Berkeley National Lab · International Computer Science Institute · UC Berkeley
Abstract
Unified atomistic modeling has the potential to accelerate discovery in chemistry, materials science, and biology by bridging data-rich chemical domains and data-scarce biological contexts. However, existing generative approaches to atomistic modeling remain highly specialized to scientific disciplines (chemistry vs. biology) or do not leverage both high-volume organic (molecule) and inorganic (material) data for general-purpose pretraining. To this end, we introduce Zatom-2, an atomistic generative model pretrained on approximately five million structures from the OMol25 and OMat24 electronic structure datasets. Zatom-2 features a multiscale Transformer architecture coupled with conditional flow matching that supports force conditioning and foundational pretraining tasks such as generation, structure prediction, and prediction of molecular and material energies and forces. Empirically, Zatom-2 achieves better molecular distribution fidelity than Zatom-1 and achieves strong performance on existing molecule and material generation benchmarks. Zatom-2 demonstrates the ability to control sample generation across low- and high-force regimes, and enhances protein generation in a low-data setting through joint generative-predictive pretraining and transfer learning, increasing protein backbone designability in a length extrapolation setting from 67.8% without pretraining to 74.8% after finetuning on 2,000 protein domains.
Figures & tables
| Datasets | Tasks | ||||||
| Model | Size | OMol25 | OMat24 | Generation | Structure prediction | Force & energy prediction | Force conditioning |
| Zatom-1 (released) | Large | – | – | – | – | ||
| Zatom-2 (I) | Small | – | – | – | – | ||
| Zatom-2 (II) | Small | – | – | – | |||
| Zatom-2 (III) | Small | – | – | ||||
| Zatom-2 (IV) | Base | – | |||||
| Model | Valid | Unique | PB-valid | Valid & unique | Valid & PB-valid |
|---|---|---|---|---|---|
| EQGAT-diff | 94.6 | 100.0 | 59.7 | 94.60 | 56.48 |
| SemlaFlow | 93.9 | 100.0 | 87.5 | 93.90 | 82.16 |
| ADiT | 95.3 | 100.0 | 85.3 | 95.30 | 81.29 |
| TABASCO | 97.6 | 99.03 | 91.6 | 96.65 | 89.40 |
| Zatom-1 | 93.6 | 99.93 | 94.1 | 93.53 | 88.08 |
| GEOM-Drugs (reference) | – | – | 94.0 | – | – |
| Model / Zatom-2 initialization | Designability (%) | Per-sequence success (%) | Diversity (#clusters) | Novelty (%) |
|---|---|---|---|---|
| (a) In-distribution: 50–128 residues | ||||
| RFdiffusion3 | ||||
| None | ||||
| Zatom-2 (IV): Generation + structure (w/o FC) | ||||
| Zatom-2 (IV): Generation + structure | ||||
| Zatom-2 (V): Generation + structure + F/E pred. | ||||
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Source data terms |
|---|---|
| OMol25 | Creative Commons Attribution 4.0 (CC BY 4.0) |
| OMat24 | CC BY 4.0 |
| QM9 | CC0 1.0 (original Figshare data deposit) |
| MP20 | MIT for the processed ADiT release used by Zatom-1; underlying Materials Project data are CC BY 4.0 |
| GEOM-Drugs | CC0 1.0 (Harvard Dataverse deposit) |
| SCOPe / ASTRAL | Provider states that all data are freely available to all users; no standard named license is specified |
| RFdiffusion3 | Emyx | Zatom-2 | |
|---|---|---|---|
| 1 | Build local atom neighborhoods (128 keys); use global token attention. | Build sparse atom/token graphs prioritizing sequence, bonds, ligand, and motif edges, then spatial neighbors. | Build local atom neighborhoods (128 keys, token-index/spatial neighbors); use global token attention. |
| 2 | Embed atom/token metadata; mix dense pair features with two triangle-free Pairformer blocks; pool atom pairs into token pairs. | Bottleneck-embed atom14 atom/token and global features; project relative positions and bonds to edge biases. | Bottleneck-embed atom1/atom14 metadata, domain, and charge/spin; residually-embed mean force norm; map bonds, relative positions, and any available fixed/reference geometry to head biases. |
| 3 | Add linear embeddings of ; condition on Fourier features of . | Add Fourier coordinate embeddings; average geometry-aware conditioning with sinusoidal time features. | Add linear atom and mean-pooled linear token embeddings of ; average each conditioning stream with Fourier features of . |
| 4 | Cache and before recycling. | Begin recycling before the atom encoder. | Begin recycling before the atom encoder. |
| 5 | Reset ; concatenate , noisy-input and recycled distograms; apply transitions and two triangle-free Pairformer blocks to obtain . | Reset to initial states plus projected, normalized ; set on sparse edges. | Reset ; set for all token pairs. No node-state feedback. |
| 6 | Reuse cached encoder/downcast outputs. | ; . | ; . |
| Hyperparameter | Setting |
|---|---|
| Approximate model size (S/base/L) | 70M / 140M / 270M parameters |
| Token-trunk blocks (S/base/L) | 9 / 18 / 36 |
| Atom encoder / decoder blocks | 3 / 3 |
| Atom / token / token-pair widths | 128 / 768 / 256 |
| Atom / token attention heads | 4 / 12 |
| Attention | Local atom attention (128 keys); global token attention |
| Model training configuration | Budget / endpoint | GPUs | GPU hours |
| Small, OMol25 G, no FC | 80 epochs | 16 | 1,830 |
| Small, OMol25 G, FC | 80 epochs | 16 | 1,820 |
| Small, OMol25 G+S, FC | 80 epochs | 64 | 2,280 |
| Base, OMol25+OMat24 G+S, FC | 80 epochs | 64 | 2,740 |
| Base, OMol25+OMat24 G+S+P, FC | 80 epochs | 64 | 2,850 |
| Large, OMol25+OMat24 G+S+P, FC | 80 epochs | 64 | 4,090 |
| OMol25 4m / ct-scd-omol25 | OMat24 1m / ct-scd-amp20 | |
|---|---|---|
| 100 | ||
| 500 | ||
| 1,000 | ||
| 5,000 |
| OMol25 | |||||
|---|---|---|---|---|---|
| Model | Mean RMSD | Median RMSD | Mean | Target MAE | AUROC |
| Zatom-2 (III) | 1.739 | 1.714 | 2.271 | 0.725 | 0.798 |
| Zatom-2 (IV) | 1.708 | 1.675 | 2.138 | 0.787 | 0.787 |
| Zatom-2 (V) | 1.762 | 1.729 | 1.953 | 0.637 | 0.801 |
| Zatom-2 (VI) | 1.723 | 1.743 | 1.211 | 1.214 | 0.818 |
| OMol25 | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Generation | Structure prediction | ||||||||
| Model | AFD | Mean | Target MAE | AUROC | Mean RMSD | Median RMSD | Mean | Target MAE | AUROC |
| Zatom-2 (IV) + FT | 0.017 | 1.205 | 0.540 | 0.807 | 1.771 | 1.747 | 1.570 | 0.632 | 0.890 |
| Zatom-2 (V) + FT | 0.016 | 1.137 | 0.530 | 0.809 | 1.729 | 1.706 | 1.458 | 0.690 | 0.828 |
| Zatom-2 (V) + FT (extended) | 0.015 | 1.040 | 0.526 | 0.840 | 1.720 | 1.713 | 1.178 | 0.871 | 0.882 |
| Model | Validity |
|---|---|
| Equivariant Diffusion | 91.90 |
| Symphony | 83.50 |
| GeoLDM | 93.80 |
| ADiT (QM9-only) | 92.19 |
| ADiT (joint QM9+MP20) | 94.45 |
| Zatom-1 (QM9-only, 80M) | 92.88 |
| Model | Overall valid | Unique | Novel | MetaSUN |
|---|---|---|---|---|
| ADiT (joint QM9+MP20) | 90.6 | 87.8 | 26.0 | 1.0 |
| Crystal-GFN | 51.7 | 51.7 | 51.7 | 0.0 |
| Crystalformer | 69.9 | 69.4 | 31.8 | 3.1 |
| LLaMat2-CIF | 84.4 | 81.4 | 30.0 | 2.1 |
| LLaMat3-CIF | 15.4 | 15.2 | 10.5 | 0.2 |
| SymmCD | 73.4 | 73.0 | 47.0 | 2.4 |
| Model | Atoms connected | Bond angles | Bond lengths | Aromatic ring flat | Double bond flat | Internal energy | No steric clash |
|---|---|---|---|---|---|---|---|
| EQGAT-diff | 84.4 | 86.9 | 87.0 | 87.0 | 87.0 | 86.8 | 82.9 |
| SemlaFlow | 92.3 | 94.8 | 94.6 | 94.9 | 94.2 | 94.8 | 92.0 |
| ADiT | 93.0 | 92.3 | 92.5 | 95.4 | 95.3 | 91.3 | 91.8 |
| TABASCO | 99.9 | 99.2 | 99.4 | 100.0 | 99.8 | 99.4 | 94.3 |
| Zatom-1 | 99.6 | 99.6 | 99.7 | 100.0 | 100.0 | 99.3 | 96.4 |
| GEOM-Drugs (reference) | – | – | – | – | – | – | – |
| Model | COV-R mean | MAT-R mean | COV-P mean | MAT-P mean |
|---|---|---|---|---|
| GO-Flow | 94.82 | 0.7971 | 70.13 | 1.1068 |
| Zatom-2 | 86.09 | 0.8459 | 83.07 | 0.9083 |
| Model | MR (%) | RMSE |
|---|---|---|
| Crystalite | 66.09 | 0.0337 |
| Zatom-2 | 64.00 | 0.1156 |