Information-Dense Synthesis for Molecular Discovery
Organizations: Department of Chemistry, Technical University of Denmark, Kgs. Lyngby, DK.
Abstract
Machine learning can accelerate molecular discovery by designing molecules and planning experiments. However, many scientific challenges demand molecules with very rare properties, and in this sparse setting, existing algorithms offer little gain over random guessing. We propose a method to efficiently search large regions of molecular space using algorithmically controlled stochastic synthesis. Rather than design, make and test individual molecules, we design and make complex mixtures, test them as a pool, then deconvolute the molecule-activity map. We optimize synthesis to encode maximal information. Theoretically, this approach can reduce the number of experiments required to find the optimal molecule among candidates from to or . In simulation, on estimated protein fitness landscapes, it finds active molecules with an order of magnitude fewer experiments than existing Bayesian optimization methods.
Figures & tables
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Notation | Description |
| The molecular space, usually a Hamming space | |
| The set of probability distributions on | |
| The molecule-activity map | |
| The design - can be a mixture or a delta. The restriction on comes from the synthesis model. | |
| The normalized design, where is the yield, the total number of molecules produced by a synthesis. | |
| The chemical yield. The total number of molecules produced by the synthesis model. |