Evolutionary One-Step Generators: Fast and Diverse Sampling for Discrete Design
Organizations: University of Trento, Italy · UiT The Arctic University of Norway, Tromsø, Norway
Abstract
Several discrete design tasks, such as molecular discovery, require diverse collections of useful candidates at low computational cost. High validity alone does not guarantee a useful candidate library: repeatedly generating the same valid structures leaves few distinct alternatives. Training for both feasibility and diversity is challenging because many relevant criteria can only be evaluated after hard decoding. To address this challenge, we propose EGO (Evolutionary Generators with One-step inference), a framework for training compact generators directly on discrete outputs. The method combines distribution matching with structural constraints and optional diversity or history-dependent rewards, using antithetic low-rank evolution strategies without requiring criterion-specific differentiable surrogates. Once trained, the generator produces the entire graph in a single neural-network evaluation. On molecular generation benchmarks, our compact generator achieves over the valid-and-unique yield per estimated dense operation compared to recent one-step flow-map baselines while retaining high chemical validity. In scaffold completion, EGO achieves an observed speedup over MoLeR in generation to SMILES and produces approximately as many filter-passing proposals within matched time budgets for generation and screening. Beyond chemistry, EGO produces as many distinct held-out elite architectures as relaxed gradient training on NAS-Bench-101. The low generation cost may enable real-time candidate generation across discrete design tasks, supporting interactive exploration of constrained design spaces and rapid construction of candidate sets for downstream evaluation.
Figures & tables
| Metric | EGO (ES) | STE | Soft | |||
|---|---|---|---|---|---|---|
| Validity (%) | ||||||
| Valid-and-unique yield (%) | ||||||
| Held-out elite yield (%) | ||||||
| Mean test accuracy (%) | ||||||
| Feature MMD ( ) | ||||||
Appendix figures & tables37 assets
Supplementary material from the paper’s appendix.
Appendix
| Experiment | Canvas size | Hidden layers |
|---|---|---|
| QM9 benchmark and matched-reward control | 9 | |
| ZINC250K, primary experiment | 38 | |
| Full GuacaMol | 88 | |
| Controlled molecular size sweep | 32 | |
| ZINC250K active-time control | 38 |
| Feature family | Information matched to training molecules |
|---|---|
| Atom states | Composition, including charge and hydrogen states. |
| Valence | Distribution of bond-order sums at atoms, in bins 0–8. |
| Bond types | Relative counts of single, double, and triple bonds. |
| WL1 | Local atom neighborhoods after one refinement round; 64 bins. |
| WL2 | Larger neighborhoods after two refinement rounds; 64 bins. |
| Spectrum | Global connectivity, using 10 bins of normalized-Laplacian eigenvalues. |
| Experiment | Seeds | Reported training budget |
|---|---|---|
| QM9 benchmark | 3 | 120,000 updates |
| Full GuacaMol | 3 | 240,000 updates |
| Molecular size sweep | 5 | 30,000 updates |
| ZINC250K active-time control | 3 | 30 active minutes |
| QM9 matched-reward control | 5 | 11.52M graph evaluations |
| Coefficient | Value used in the experiments |
|---|---|
| Core feature weight | |
| Risk weights | |
| Diversity thresholds | |
| Rank weights |
| Feature family | Dimension | Purpose |
|---|---|---|
| Pixels | 784 | Fine image detail. |
| pooled means | 196 | Local stroke patterns. |
| pooled means | 49 | Coarse digit shape. |
| Row/column means | 56 | Horizontal and vertical ink placement. |
| Neighbor differences | 56 | Stroke transitions along rows and columns. |
| Global statistics | 8 | Ink amount, position, spread, and mirror asymmetry. |
| Method | Steps | Valid (%) | Unique (%) | FCD | VU/s |
|---|---|---|---|---|---|
| Autoregressive generation | |||||
| GraphARM [ Kong et al., 2023 ] | 90.25 | 95.62 | 1.22 | — | |
| G2PT-base [ Chen et al., 2025 ] | tokens | 99.0 | 96.8 | 0.06 | — |
| Iterative diffusion / flow | |||||
| GDSS a [ Jo et al., 2022 ] | 1,000 | 95.72 | 98.46 | 2.900 | |
| DiGress [ Vignac et al., 2023 ] | 500 | 99.0 | 96.2 | — | — |
| Architecture | Params. (M) | Full valid | Strict valid | Strict VUN | FCD |
|---|---|---|---|---|---|
| 0.866 | |||||
| 1.456 | |||||
| 2.047 | |||||
| 0.433 | |||||
| 2.466 |
| Architecture | 30,000 updates | 60,000 updates | 120,000 updates |
|---|---|---|---|
| 63.04 | 66.23 | 68.70 | |
| 63.19 | 66.19 | 68.31 | |
| 62.54 | 64.33 | 67.10 | |
| 61.02 | 63.40 | 67.06 | |
| 62.16 | 64.66 | 66.50 |
| EGO architecture | Dense MFLOP | Time (s) | VU/s | |
|---|---|---|---|---|
| (primary) | ||||
| Method | Full valid | Strict valid | Strict unique yield | Strict VUN | FCD |
|---|---|---|---|---|---|
| EGO | |||||
| RLOO | |||||
| RLOO + rarity |
| Method | Full valid | Strict valid | Strict unique yield | Strict VUN | FCD |
|---|---|---|---|---|---|
| EGO | |||||
| RLOO |
| Method | Steps | Valid (%) | Unique (%) | FCD | VU/s |
|---|---|---|---|---|---|
| Autoregressive generation and autoregressive diffusion | |||||
| GraphARM [ Kong et al., 2023 ] | 88.23 | 99.46 | 16.26 | — | |
| LO-ARM-st-sep [ Wang et al., 2025 ] | 96.26 | 100.00 | 3.229 | — | |
| PARD [ Zhao et al., 2024 ] | blockwise | 95.23 | 99.99 | 1.98 | — |
| Iterative diffusion and flow | |||||
| GDSS a [ Jo et al., 2022 ] | 1,000 | 97.01 | 99.64 | 14.656 | |
| Architecture | Full valid | Strict valid | Connected | Unique given valid | Full VU | Strict VUN | FCD |
|---|---|---|---|---|---|---|---|
| — | — | ||||||
| — | — | ||||||
| Architecture | 30,000 updates | 60,000 updates | 120,000 updates |
|---|---|---|---|
| Full validity (%) | |||
| Architecture | VU/s (graph generation only) |
|---|---|
| — | |
| Timing boundary | Time (s) | Full VU/s |
|---|---|---|
| Graph generation only | ||
| Generation through SMILES/CSV |
| Method | Full | Strict | Strict VUN | Conn. | ||
|---|---|---|---|---|---|---|
| 16 | EGO | 97.33 1.11 | 94.91 1.57 | 90.84 0.77 | 97.55 0.94 | 0.0084 0.0007 |
| 16 | EGO 0 | 86.60 2.81 | 83.82 4.18 | 77.39 3.75 | 97.07 1.85 | 0.0052 0.0009 |
| 16 | STE (feature-ST) | 99.21 0.91 | 96.52 1.75 | 7.51 4.83 | 97.31 2.29 | 0.0538 0.0351 |
| 16 | Soft | 95.01 1.74 | 89.33 7.19 | 10.81 4.21 | 94.29 6.82 | 0.0430 0.0313 |
| 24 | EGO | 95.48 2.98 | 91.69 3.19 | 89.45 2.67 | 96.18 2.12 | 0.0099 0.0015 |
| 24 | EGO 0 | 75.58 11.00 | 69.38 14.01 | 63.30 10.96 | 93.17 4.79 | 0.0065 0.0012 |
| Contrast | Validity | VUN | ||
|---|---|---|---|---|
| 16 | EGO EGO 0 | |||
| 16 | EGO 0 Soft | |||
| 16 | EGO 0 STE | |||
| 16 | STE Soft | |||
| 24 | EGO EGO 0 | |||
| 24 | EGO 0 Soft |
| Method | Full valid | Strict valid | Strict VUN | Updates ( ) |
|---|---|---|---|---|
| EGO | ||||
| EGO 0 | ||||
| STE | ||||
| Soft |
| Method | Steps | Valid (%) | VU (%) | FCD | VU/s |
| Autoregressive generation | |||||
| SMILES LSTM [ Brown et al., 2019 ] | tokens | 95.9 | — | ||
| G2PT-base [ Chen et al., 2025 ] | tokens | 94.6 | — | ||
| G2PT-large [ Chen et al., 2025 ] | tokens | 95.3 | — | ||
| Iterative diffusion and flow | |||||
| DiGress [ Vignac et al., 2023 ] | 500 | 85.2 | — | ||
| Method | Full | Strict | Strict VUN | Conn. | ||
|---|---|---|---|---|---|---|
| 16 | EGO | 94.57 2.66 | 90.27 2.34 | 82.12 2.19 | 95.68 1.43 | 0.0091 0.0024 |
| 16 | EGO 0 | 86.71 5.49 | 82.60 6.82 | 66.01 4.08 | 95.79 1.59 | 0.0064 0.0016 |
| 16 | STE (feature-ST) | 99.72 0.38 | 97.88 1.95 | 6.39 1.38 | 98.16 1.63 | 0.0364 0.0189 |
| 16 | Soft | 97.75 2.09 | 92.52 4.41 | 7.74 7.53 | 94.77 2.62 | 0.0387 0.0162 |
| 24 | EGO | 92.80 5.42 | 86.27 6.47 | 81.53 6.81 | 93.44 1.54 | 0.0109 0.0037 |
| 24 | EGO 0 | 74.42 7.91 | 68.65 7.93 | 58.68 8.22 | 93.99 3.17 | 0.0088 0.0024 |
| Contrast | Validity | VUN | ||
|---|---|---|---|---|
| 16 | EGO EGO 0 | |||
| 16 | EGO 0 Soft | |||
| 16 | EGO 0 STE | |||
| 16 | STE Soft | |||
| 24 | EGO EGO 0 | |||
| 24 | EGO 0 Soft |
| Metric | CFM-ECLD pilot |
|---|---|
| Full validity (%) | 19.67 |
| Strict validity (%) | 4.64 |
| Uniqueness among full valid molecules (%) | 100.00 |
| Full VU (%) | 19.67 |
| Strict VU (%) | 4.64 |
| Full VUN (%) | 19.66 |
| Dataset | Method | VU (%) | Dense MFLOP | |
| QM9 | GraphARM | — | — | |
| G2PT-base | — | |||
| CatFlow (100) | — | |||
| DeFoG (500) | 86,234.240 | 11.1 | ||
| MELD | — | |||
| MolGAN | 0.714 | 142,945 |
| Method | QM9 | ZINC250K | GuacaMol |
|---|---|---|---|
| EGO | 2.0983 | ||
| DeFoG [ Qin et al., 2025 ] | 6.5 | 14 | 141 |
| Cometh [ Siraudin et al., 2025 ] | 6 | — | |
| CFM [ Roos et al., 2026 ] | to |
| Metric | 42 / 41 | 43 / 42 | 44 / 43 |
| Mean deadline per query (s) | 6.176 | 6.105 | 5.974 |
| Mean proposals screened, EGO | 1,822.94 | 1,947.70 | 1,907.35 |
| Mean proposals screened, MoLeR | 256.0 | 256.0 | 256.0 |
| Filter-passing proposals, EGO | 55,152 | 52,975 | 56,064 |
| Filter-passing proposals, MoLeR | 5,434 | 5,485 | 5,417 |
| Passing fraction (%), EGO | 30.25 | 27.20 | 29.39 |
| Metric | EGO | MoLeR |
|---|---|---|
| Usable completions (%) | ||
| Distinct usable yield (%) | ||
| Filter passing fraction (%) | ||
| Distinct filtered completions | ||
| Scaffold coverage (/100) | ||
| Generation to SMILES (s) |
| Method | Seed type | Seed | Distinct filtered count |
|---|---|---|---|
| EGO | Training | 42 | |
| EGO | Training | 43 | |
| EGO | Training | 44 | |
| MoLeR | Sampling | 41 | |
| MoLeR | Sampling | 42 | |
| MoLeR | Sampling | 43 |
| Method | Update time | Compilation | Validation |
|---|---|---|---|
| EGO | |||
| STE | |||
| Soft |
| Method | LeNet Fréchet | Predicted classes / 10 |
|---|---|---|
| EGO | ||
| STE | ||
| Soft |
| Metric | EGO | STE | Soft |
|---|---|---|---|
| LeNet KID | |||
| Pooled-image MMD 2 | |||
| Class-balance TV | |||
| Normalized class entropy |
| Method | Selected update | Predicted classes / 10 |
|---|---|---|
| EGO | 29,000 / 30,000 / 30,000 | 6 / 6 / 7 |
| STE | 500 / 30,000 / 500 | 8 / 8 / 7 |
| Soft | 1,500 / 28,000 / 1,000 | 8 / 9 / 8 |