Initialization Improves LLM-Driven Discovery
Organizations: University of Chicago · Vector Institute · University of Toronto
Abstract
Large Language Models (LLMs) have been used for novel discovery of algorithms, theorems, drugs, and other tasks through the use of harnesses that prompt an LLM to iteratively optimize an objective. In this work, we study the relationship between the population of previous iterates and eventual discovery success. We generalize past work on harness design to develop a suite of 12 harnesses called 'Modular' and characterize their performance across 5 diverse discovery tasks, finding that discovery success is brittle and sensitive to harness design. We uncover mode collapse, characterized by a dramatic drop in the diversity of iterates, as a common failure mode. We find that popular state-of-the-art harnesses and diversity-inducing harness interventions, which aim to prolong this collapse, yield inconsistent gains. Our results instead uncover that the performance of early discoveries is predictive of eventual success. We therefore propose a universally applicable intervention that performs an initial stage of parallel exploration in order to initialize subsequent iterative optimization. Our method provides consistent gains across many harnesses and target applications, confirming the importance of initialization in LLM-driven discovery.
Figures & tables
| Task | Optimizer LLM |
| CloudCast | Gemini 3.5 Flash |
| Can’t Be Late | Gemini 3.7 Flash |
| TSP | Gemini 3.5 Flash |
| Circle Packing | GPT-OSS (120B) |
| Prompt Optim. | Llama 3.1 Instruct (8B) |
| Can’t Be Late | CloudCast | TSP | Circle Packing | Prompt Optim. | |
| Modular (n=12) | 0.02 ± 0.01 | 0.1271 ± 0.0195 | 0.095 ± 0.050 | 0.26 ± 0.06 | 0.40 ± 0.04 |
| Modular w/ Diversity Interven. (n=60) | 0.06 ± 0.02 | 0.1417 ± 0.0132 | 0.100 ± 0.022 | 0.27 ± 0.02 | 0.38 ± 0.03 |
| State-of-the-art (n=3) | 0.05 ± 0.07 | 0.1680 ± 0.3698 | 0.259 ± 0.332 | 0.25 ± 0.10 | 0.36 ± 0.23 |
| Can’t Be Late | CloudCast | TSP | Circle Packing | Prompt Optim. | |
| Modular (n=12) | -91.94 ± 2.75 | 0.0098 ± 0.0002 | -1719 ± 16 | 2.44 ± 0.08 | 0.53 ± 0.03 |
| Sim thresh. (0.95) (n=12) | -95.79 ± 2.98 | 0.0097 ± 0.0002 | -1777 ± 29 | 2.19 ± 0.38 | 0.53 ± 0.03 |
| Sim thresh. (0.8) (n=12) | -91.68 ± 2.81 | 0.0097 ± 0.0001 | -1775 ± 17 | 2.34 ± 0.26 | 0.52 ± 0.02 |
| Sim thresh. (0.95) + LLM-Judge (n=12) | -95.93 ± 2.82 | 0.0098 ± 0.0001 | -1744 ± 46 | 2.33 ± 0.23 | 0.53 ± 0.04 |
| Sim thresh. (0.8) + LLM-Judge (n=12) | -94.82 ± 2.75 | 0.0098 ± 0.0001 | -1735 ± 37 | 2.26 ± 0.30 | 0.51 ± 0.04 |
| Deduplication (n=12) | -93.93 ± 3.34 | 0.0098 ± 0.0001 | -1721 ± 12 | 2.04 ± 0.44 | 0.52 ± 0.04 |
| Can’t Be Late | CloudCast | TSP | Circle Packing | Prompt Optim. | |
| Modular (n=12) | -91.94 ± 2.75 | 0.0098 ± 0.0002 | -1719 ± 16 | 2.44 ± 0.08 | 0.53 ± 0.03 |
| State-of-the-art (n=3) | -91.32 ± 8.33 | 0.0095 ± 0.0021 | -1772 ± 250 | 2.51 ± 0.07 | 0.47 ± 0.20 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Modular ID | LLM Mutation Strategy | Parent Sampling Strategy |
| Modular 1 | DE ( 6 ) | High Score ( Alg 2 ) |
| Modular 2 | DE ( 6 ) | Tournament ( Alg 3 ) |
| Modular 3 | DE ( 6 ) | Wheel ( Alg 4 ) |
| Modular 4 | GA ( 7 ) | High Score ( Alg 2 ) |
| Modular 5 | GA ( 7 ) | Tournament ( Alg 3 ) |
| Modular 6 | GA ( 7 ) | Wheel ( Alg 4 ) |
| Task | LLMs | Token Discovery Budget | Evaluation Metric |
| Can’t Be Late Seed: 13 Instructions: 8 | Gemini 3.5 Flash Gemini 3.7 Flash | 1M | Negated cost in dollars of discovered scheduling strategy. |
| CloudCast Seed: 14 Instructions: 9 | Gemini 3.5 Flash Gemini 3.7 Flash | 2M | Score where cost of discovered algorithm in dollars is = egress costs (data_vol × edge_cost) + instance costs (runtime × cost_per_hour) |
| TSP Seed: 15 Instructions: 10 | Gemini 3.5 Flash Gemini 3.7 Flash GPT-OSS (120B) | 1M | Negated total distance of discovered path (unit less). |
| Circle Packing ( ) Seed: 16 Instructions: 11 | Gemini 2.5 Flash GPT-OSS (120B) Qwen 3.6 A3B (35B) | 200K | Sum of all circle diameters in discovered configuration. |
| Prompt Optim. Seed: 17 Instructions: 12 | Gemini 2.5 Flash Gemini 3.5 Flash Llama 3.1 Instruct (8B) | 150K | Evaluation performance on GSM8K Benchmark Cobbe et al. (2021) w/ discovered instruction prompt for downstream LLM (OLMo 2 0425 SFT (1B)) measured as a percentage |
| Task | Initial:Total Budget Ratios ( ) |
| Can’t Be Late | 0.1, 0.25, 0.5, 0.75, 0.9 |
| CloudCast | 0.1, 0.25, 0.5, 0.75, 0.9 |
| TSP | 0.1, 0.25, 0.5, 0.75, 0.9 |
| Circle Packing ( ) | 0.1, 0.25, 0.5, 0.75, 0.9 |
| Prompt Optimization (GSM8K) | 0.17, 0.33, 0.5, 0.67, 0.83 |
| Task | ||
| Can’t Be Late | 0.5 | 25% |
| CloudCast | 0.0002 | 10% |
| TSP | 30 | 10% |
| Circle Packing ( ) | 0.1 | 10% |
| Prompt Optimization (GSM8K) | 0.01 | 25% |
| Can’t Be Late | |||
| Gemini 3.7 Flash | Gemini 3.5 Flash | ||
| Modular ( ) | 0.02 ± 0.01 | 0.08 ± 0.03 | |
| Sim thresh. ( ) ( ) | 0.05 ± 0.05 | 0.13 ± 0.07 | |
| Sim thresh. ( ) ( ) | 0.13 ± 0.06 | 0.18 ± 0.06 | |
| Sim thresh. ( ) + LLM judge ( ) | 0.03 ± 0.03 | 0.08 ± 0.04 | |
| Sim thresh. ( ) + LLM judge ( ) | 0.03 ± 0.01 | 0.07 ± 0.02 | |
| Can’t Be Late | |||
| Gemini 3.7 Flash | Gemini 3.5 Flash | ||
| Modular ( ) | -91.94 ± 2.75 | -97.80 ± 1.92 | |
| Sim thresh. ( ) ( ) | -95.79 ± 2.98 | -98.14 ± 1.86 | |
| Sim thresh. ( ) ( ) | -91.68 ± 2.81 | -97.42 ± 2.34 | |
| Sim thresh. ( ) + LLM judge ( ) | -95.93 ± 2.82 | -98.64 ± 0.52 | |
| Sim thresh. ( ) + LLM judge ( ) | -94.82 ± 2.75 | -98.39 ± 1.22 | |
| Sim thresh. ( ) | Sim thresh. ( ) | Sim thresh. ( ) | Sim thresh. (Greedy, ) | ||
| Task | LLM | ||||
| Can’t Be Late | Gemini 3.5 Flash | -98.56 ± 0.43 | -98.58 ± 0.42 | -98.62 ± 0.38 | -98.63 ± 0.35 |
| Gemini 3.7 Flash Flash | -90.22 ± 0.90 | -90.52 ± 0.91 | -89.79 ± 0.85 | -89.90 ± 0.87 | |
| CloudCast | Gemini 3.5 Flash | 0.0099 ± 0.0000 | 0.0100 ± 0.0001 | 0.0099 ± 0.0000 | 0.0099 ± 0.0000 |
| Gemini 3.7 Flash | 0.0095 ± 0.0001 | 0.0095 ± 0.0001 | 0.0095 ± 0.0001 | 0.0095 ± 0.0001 | |
| TSP | GPT-OSS (120B) | -4433 ± 415 | -4674 ± 422 | -3993 ± 322 | -4071 ± 346 |