Biological machine learning was long bottlenecked by the ability to synthesize designed DNA. Variational synthesis models control chemical reactions to physically manufacture quadrillions of designed sequences in DNA. However, training these generative models is challenging: constraints on chemical synthesis can force many parameters into a discrete space, limiting the ability to pre-train and fine-tune. In this article we train ``free'' variational synthesis models using stochastic gradient descent in continuous space, and then discretize with post-training quantization to impose hardware and wetware constraints. This enables variational synthesis models to satisfy stringent reward criteria, while still synthesizing diverse designs, achieving a strictly dominating quality-diversity Pareto frontier. We demonstrate by training variational synthesis models of enzymes, peptides, antibody CDRH3s, and regulatory DNA elements. In silico performance is maintained in vitro.
Figures & tables
Figure 1: Enzyme design (Cytochrome P450). (a) Quality: ESMFold’s pTM and pLDDT, and TM-score to the predicted CYP2C9 structure. We compare cVS pre- and post-quantization (free cVS, cVS) to EM VS and data. (b) Diversity: entropy of the library (left), and mean Hamming distance (center); higher is more diverse on both. The EM library had many internal stop codons (right); the other evaluations are on samples with no stops. (c) Predicted structure of independent samples from the cVS library, compared to that of the human cytochrome P450 2C9.
Figure 2: Quality-diversity Pareto frontier on benchmark peptide design. Y-axis: mean binding quantile of peptides in the designed library, compared to a reference distribution. X-axis: "theoretical diversity". We compare cVS (blue) to PGLD (gray). We show the raw data from ( Sussex et al., 2026 ) (published frontier) and our reproduction (stars), with learning rate (LR) optimized and maximized NMC . We compare to cVS designs before (free cVS) and after quantization (cVS). Left: Reproduction and hyperparameter improvement of PGLD compared to cVS. Right: Pareto frontiers for PGLD (published, optimized), and for cVS (post-quantization).
Figure 3: scFv CDRH3 library design (TCR mimics). (a) Pareto frontier of quality (mean predicted binding counts against a pHLA target) versus diversity (KL to human repertoire prior). (b) Performance with changing synthesis hardware and wetware, increasing the number of pre-set mixes available in Uψ1:K . (c) Performance of free cVS with increasing wells M .
Figure 4: Regulatory DNA design (EF1 α ). (a) Quality-diversity Pareto frontier, evaluating the average difference in cell-type accessibility versus the KL to the human genome prior. (b) Predicted accessibility of sampled sequences (above) from the learned synthesis models (below). Positions in each well are colored by the nucleotide mixture.
Figure 5: In vitro validation of the EF1 α promoter library : (a) Estimated position of the in vitro library on the Pareto frontier. (b) KSD-B goodness-of-fit to the reverse-KL target π(x)exp(αr(x)) . Mean and standard error (SEM) from five independent redraws of 1000 samples from the in silico model or the in vitro sequencing data. (c) Average reward and SEM. (d) Average prior log likelihood and SEM. (e-g) Low-dimensional representation of the samples from the prior model (pink) overlaid with the samples from each of the evaluated models.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Training loss as a function of the number of reward evaluations (in millions), for PGLD versus free cVS on the peptide benchmark . Left: overview, with dashed line showing the results plotted in Figure 2 . Middle, Right: zoom in, with different learning rates marked.
Figure 7: Hyperparameter sensitivity analysis on the peptide benchmark, for PGLD versus free cVS . x-axis scans each hyperparameter relative to the setting used in Figure 2 . First column: learning rate. Second and third columns: Adam hyperparameters. Fourth column: rate of the exponential moving average used to construct the baseline control variate. Fifth column: training batch size.
Figure 8: Hyperparameter sensitivity analysis on scFv CDRH3 designs, for PGLD versus free cVS . x-axis scans each hyperparameter relative to the setting used in Figure 3(a) . First column: learning rate. Second and third columns: Adam hyperparameters. Fourth column: rate of the exponential moving average used to construct the baseline control variate. Fifth column: training batch size.
Figure 9: Training loss as a function of the number of reward evaluations (in millions), for PGLD versus free cVS, on the scFv CDRH3 designs . Left: overview, with dashed line showing the results plotted in Figure 3(a) . Middle, Right: zoom in, with different learning rates marked.
Figure 10: Performance of cVS and PGLD as the number of wells M increases, on the scFv CDRH3 designs. Left: free cVS. Middle: PGLD. Right: post-quantization cVS for increasing M , compared to PGLD with increasing M .
Figure 11: Performance of cVS and PGLD with and without pretraining on the forward KL objective, on the scFv CDRH3 designs.
Figure 12: Performance on the scFv CDRH3 designs with total reward evaluations held fixed. Here each method is trained from scratch, using 5 million evaluations of the reward and prior, rather than including an added pre-training phase for cVS and PGLD as in Figure 3(a) .
Figure 13: Pareto frontier for cVS versus PGLD, measuring diversity by the Shannon entropy, on the scFv CDRH3 designs . X-axis is the exponential of the Shannon entropy of qθ(x) , a measure of the effective library size. These designs are trained on a uniform prior rather than a human prior, to optimize for this diversity measure.
Figure 14: Diversity measured by the exponential of the kernel 2-Renyi entropy, on the scFv CDRH3 designs. We compare a cVS design to a PGLD design with the same constraints and same average reward, which achieves similar mean Hamming distance among sequences (left). We find similar diversity at high RBF kernel bandwidths, but cVS reaches much higher diversity at low bandwidths, indicating PGLD finds spread out but narrow modes compared to cVS (right).
Figure 15: Ablating gradient variance reduction strategies, on the scFv CDRH3 designs. Removing the prior Rao-Blackwellization or the REINFORCE control variate reduces variance (left) but leads to only a minor change in final performance (right).
Figure 16: Ablating split-merge resampling of wells, on the scFv CDRH3 designs. Removing the split-merge step ( Appendix A ) decreases performance.
Figure 17: Designed regulatory DNA sequences in a genomic context . Same as Figure 4 but the designs are done in a random genomic context, rather than EF1 α . (a) Quality-diversity Pareto frontier, evaluating the average difference in cell-type accessibility versus the KL to the human genome prior. (b) Predicted accessibility of sampled sequences (above) from the learned synthesis models (below). Positions in each well are colored by the nucleotide mixture.
Figure 18: KSD-B of EF1 α libraries computed with with exponential Hamming kernels (a) KSD-B as in Figure 5(b) but computed with an exponential Hamming kernel with a scanned bandwidth σ , rather than an IMQ Hamming. (b) Same as in (a) but with a fixed σ=3 and five independent redraws for each model (error bars: SEM across redraws)
Jul 1, 2026·Miruna Cretu, John Bradshaw, Patricia Suriana +6MoleculesSynthesis
University of Cambridge, Cambridge, UK · Prescient Design (AI for Drug Discovery), Genentech, South San Francisco, USA · Work done during an internship at Prescient Design