Large Language Models (LLMs) have been used for novel discovery of algorithms, theorems, drugs, and other tasks through the use of harnesses that prompt an LLM to iteratively optimize an objective. In this work, we study the relationship between the population of previous iterates and eventual discovery success. We generalize past work on harness design to develop a suite of 12 harnesses called 'Modular' and characterize their performance across 5 diverse discovery tasks, finding that discovery success is brittle and sensitive to harness design. We uncover mode collapse, characterized by a dramatic drop in the diversity of iterates, as a common failure mode. We find that popular state-of-the-art harnesses and diversity-inducing harness interventions, which aim to prolong this collapse, yield inconsistent gains. Our results instead uncover that the performance of early discoveries is predictive of eventual success. We therefore propose a universally applicable intervention that performs an initial stage of parallel exploration in order to initialize subsequent iterative optimization. Our method provides consistent gains across many harnesses and target applications, confirming the importance of initialization in LLM-driven discovery.
Figures & tables
Figure 1 : Initialization raises the floor for discovery under a fixed total token budget. We find that initializing (the Modular ) discovery harnesses with a high-performing initial population is beneficial for downstream discovery success across harnesses, tasks, and LLMs. Initialization budget is informed by the convergence criteria in Section 5 . Initial population size m=15 , similarity threshold s=0.98 .
Figure 2 : Sequential discovery can outperform parallel discovery. Comparing sequential discovery harnesses from our Modular suite (detailed in Table 5 ) with parallel discovery using the “Reflect” ( 5 ) and “K Context” prompts ( 4 ). Higher scores are better. Both parallel and sequential approaches are initialized with the same initial seed solutions. We see that generally some form of sequential compute outperforms parallel; but the types of harnesses that perform the best vary with task and LLM. Full results in Fig 9 .
Task
Optimizer LLM
CloudCast
Gemini 3.5 Flash
Can’t Be Late
Gemini 3.7 Flash
TSP
Gemini 3.5 Flash
Circle Packing
GPT-OSS (120B)
Prompt Optim.
Llama 3.1 Instruct (8B)
Table 1 : A subset of (task, LLM) pairs studied. Full set detailed in Appendix B , Table 6 .
Figure 3 : Sequential discovery is prone to mode collapse. The windowed average pairwise diversity of discoveries over the course of a trajectory across all Modular variants vs. average diversity of parallel discovery. Diversity for the parallel discovery scheme is computed over the full set of discoveries. Sequential harnesses consistently yield fewer diverse discoveries and exhibit diminished diversity over the trajectory, whereas parallel schemes discover more diverse candidates. Window size = 15 for all tasks except Circle Packing which has window size = 6. Full results in Fig 10 .
Can’t Be Late
CloudCast
TSP
Circle Packing
Prompt Optim.
Modular (n=12)
0.02 ± 0.01
0.1271 ± 0.0195
0.095 ± 0.050
0.26 ± 0.06
0.40 ± 0.04
Modular w/ Diversity Interven. (n=60)
0.06 ± 0.02
0.1417 ± 0.0132
0.100 ± 0.022
0.27 ± 0.02
0.38 ± 0.03
State-of-the-art (n=3)
0.05 ± 0.07
0.1680 ± 0.3698
0.259 ± 0.332
0.25 ± 0.10
0.36 ± 0.23
Table 2 : Discovery harnesses can discover more diverse populations. We compute the diversity of all discoveries in a trajectory, and then report the average diversity across different harnesses. Both state-of-the-art discovery harnesses ( SE , OA , OE ) and Modular harnesses with diversity interventions on average generally achieve more diverse discoveries than baseline Modular harnesses. Full results in Appendix C , Table 9 .
Can’t Be Late
CloudCast
TSP
Circle Packing
Prompt Optim.
Modular (n=12)
-91.94 ± 2.75
0.0098 ± 0.0002
-1719 ± 16
2.44 ± 0.08
0.53 ± 0.03
Sim thresh. (0.95) (n=12)
-95.79 ± 2.98
0.0097 ± 0.0002
-1777 ± 29
2.19 ± 0.38
0.53 ± 0.03
Sim thresh. (0.8) (n=12)
-91.68 ± 2.81
0.0097 ± 0.0001
-1775 ± 17
2.34 ± 0.26
0.52 ± 0.02
Sim thresh. (0.95) + LLM-Judge (n=12)
-95.93 ± 2.82
0.0098 ± 0.0001
-1744 ± 46
2.33 ± 0.23
0.53 ± 0.04
Sim thresh. (0.8) + LLM-Judge (n=12)
-94.82 ± 2.75
0.0098 ± 0.0001
-1735 ± 37
2.26 ± 0.30
0.51 ± 0.04
Deduplication (n=12)
-93.93 ± 3.34
0.0098 ± 0.0001
-1721 ± 12
2.04 ± 0.44
0.52 ± 0.04
Table 3 : Diversity interventions do not consistently improve discovery harness performance. Each cell contains the average maximal score achieved over a set of harnesses. The basic Modular harnesses generally have the highest discovery quality, with some diversity interventions matching, and rarely exceeding, performance. Sim thresh. is the cosine similarity threshold ( s ). Bold indicates highest score within a column. Full results in Appendix C , Table 10 .
Can’t Be Late
CloudCast
TSP
Circle Packing
Prompt Optim.
Modular (n=12)
-91.94 ± 2.75
0.0098 ± 0.0002
-1719 ± 16
2.44 ± 0.08
0.53 ± 0.03
State-of-the-art (n=3)
-91.32 ± 8.33
0.0095 ± 0.0021
-1772 ± 250
2.51 ± 0.07
0.47 ± 0.20
Table 4 : State-of-the-art discovery harnesses ( SE , OA , OE ) do not consistently outperform Modular harnesses. The ideal set of discovery harnesses varies across (task, LLM) settings. Bold indicates highest score within a column. Full results in Appendix C , Table 10 .
Figure 4 : Initial discovery quality can predict downstream quality. We plot the correlation between the median score of the first 15 discoveries vs. the max score achieved in a trajectory across varying discovery harnesses (72 Modular harnesses and 3 state-of-the-art harnesses).
Figure 5 : Initialized discovery can outperform purely sequential and parallel discovery. Sequential discovery harnesses initialized with a high-performing population (blue) via a fraction of the discovery budget yield the most performant discoveries. The ideal initialization to total discovery budget ratio (dotted line) can be successfully predicted via a convergence analysis Section 5 . Diversity threshold s=0.98 for all initial populations, population size m=15 . Full results in Fig 12 .
Figure 6 : High scoring trajectories have higher initial diversity but eventually experience mode collapse. We stratify the Modular harness trajectories (for CloudCast ) from Table 3 into tertiles based on the max score of a given trajectory. Windowed diversity visualized over trajectory (window size=15). Full results in Appendix C , Fig 11 .
Figure 7 : Initializing state-of-the-art harnesses improves discoveries. We find that initializing the state-of-the-art discovery harnesses ( OA , OE , SE ) with the single highest-performing seed discovery within a given initialization budget is beneficial for downstream discovery success across harnesses, tasks, and LLMs. Initialization budget is informed by the convergence criteria in Section 5 .
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Modular ID
LLM Mutation Strategy
Parent Sampling Strategy
Modular 1
DE ( 6 )
High Score ( Alg 2 )
Modular 2
DE ( 6 )
Tournament ( Alg 3 )
Modular 3
DE ( 6 )
Wheel ( Alg 4 )
Modular 4
GA ( 7 )
High Score ( Alg 2 )
Modular 5
GA ( 7 )
Tournament ( Alg 3 )
Modular 6
GA ( 7 )
Wheel ( Alg 4 )
Appendix
Table 5 : Modular variants summary.
Task
LLMs
Token Discovery Budget
Evaluation Metric
Can’t Be Late Seed: 13 Instructions: 8
Gemini 3.5 Flash Gemini 3.7 Flash
1M
Negated cost in dollars of discovered scheduling strategy.
CloudCast Seed: 14 Instructions: 9
Gemini 3.5 Flash Gemini 3.7 Flash
2M
Score =1+cost1 where cost of discovered algorithm in dollars is = egress costs (data_vol × edge_cost) + instance costs (runtime × cost_per_hour)
TSP Seed: 15 Instructions: 10
Gemini 3.5 Flash Gemini 3.7 Flash GPT-OSS (120B)
1M
Negated total distance of discovered path (unit less).
Evaluation performance on GSM8K Benchmark Cobbe et al. (2021) w/ discovered instruction prompt for downstream LLM (OLMo 2 0425 SFT (1B)) measured as a percentage
Table 7 : We vary the initialization budget ratios ( r ) across five values for each task for the Modular harness initialization experiments in Section 5 .
Figure 8 : Parallel discovery quality converges. We find that initializing (the Modular ) discovery harnesses with a high-performing initial population is beneficial for downstream discovery success across harnesses, tasks, and LLMs. The initialization budgets are informed by the convergence criteria in Section 5 . We overlay the running maximum score of discoveries from the parallel discovery scheme ( Section A.4 ) ordered by a random seed; we notice that these discoveries converge, making them amenable to convergence analysis (Appendix B.2 ). Initial population size m=15 , similarity threshold s=0.98 .
Task
δ
N
Can’t Be Late
0.5
25%
CloudCast
0.0002
10%
TSP
30
10%
Circle Packing ( n=26 )
0.1
10%
Prompt Optimization (GSM8K)
0.01
25%
Appendix
Table 8 : Task-specific δ and N values, by which convergence of a discovery trajectory can be determined.
Figure 9 : Sequential discovery can outperform parallel discovery. We present the extended version of Fig 2 here. Comparing sequential discovery harnesses from our Modular suite (detailed in Table 5 ) with parallel discovery using the “Reflect” ( 5 ) and “K Context” prompts ( 4 ). Higher scores are better. Both parallel and sequential approaches are initialized with the same initial seed solutions. We see that generally some form of sequential compute outperforms parallel; but the types of harnesses that perform the best vary with task and LLM. Different rows correspond to different discovery tasks. Note that within a column, the LLM varies across tasks.
Figure 10 : Sequential discovery is prone to mode collapse. We present the extended version of Fig 3 here. The windowed average pairwise diversity of discoveries over the course of a trajectory across all Modular variants vs. average diversity of parallel discovery. Diversity for the parallel discovery scheme is computed over the full set of discoveries. Sequential harnesses consistently yield fewer diverse discoveries and exhibit diminished diversity over the trajectory, whereas parallel schemes discover more diverse candidates. Window size = 15 for all tasks except Circle Packing which has window size = 6.
Figure 11 : High scoring trajectories eventually experience mode collapse. We present the extended version of Fig 6 here. We stratify the Modular harness trajectories from Table 3 into tertiles based on the max score of a given trajectory. Windowed diversity visualized over trajectory; window size = 15 for all tasks except Circle Packing which has window size = 6.
Can’t Be Late
Gemini 3.7 Flash
Gemini 3.5 Flash
Modular ( n=12 )
0.02 ± 0.01
0.08 ± 0.03
Sim thresh. ( s=0.95 ) ( n=12 )
0.05 ± 0.05
0.13 ± 0.07
Sim thresh. ( s=0.8 ) ( n=12 )
0.13 ± 0.06
0.18 ± 0.06
Sim thresh. ( s=0.95 ) + LLM judge ( n=12 )
0.03 ± 0.03
0.08 ± 0.04
Sim thresh. ( s=0.8 ) + LLM judge ( n=12 )
0.03 ± 0.01
0.07 ± 0.02
Appendix
Table 9 : Discovery harnesses can discovery more diverse populations. We report the comprehensive diversity results for the Section 4 experiments. Both state-of-the-art discovery harnesses ( SE , OA , OE ) and Modular harnesses with diversity interventions on average generally achieve more diverse discoveries than baseline Modular harnesses across tasks and LLMs.
Can’t Be Late
Gemini 3.7 Flash
Gemini 3.5 Flash
Modular ( n=12 )
-91.94 ± 2.75
-97.80 ± 1.92
Sim thresh. ( s=0.95 ) ( n=12 )
-95.79 ± 2.98
-98.14 ± 1.86
Sim thresh. ( s=0.8 ) ( n=12 )
-91.68 ± 2.81
-97.42 ± 2.34
Sim thresh. ( s=0.95 ) + LLM judge ( n=12 )
-95.93 ± 2.82
-98.64 ± 0.52
Sim thresh. ( s=0.8 ) + LLM judge ( n=12 )
-94.82 ± 2.75
-98.39 ± 1.22
Appendix
Table 10 : Discovery harnesses that explore more do not necessarily achieve more performant discoveries . We report the comprehensive set of results for Section 4 . Neither state-of-the-art discovery harnesses ( SE , OA , OE ) nor Modular harnesses with diversity interventions predictably achieve more performant discoveries; instead the ideal choice or harness/intervention varies across (task, LLM) settings.
Figure 12 : Initialized discovery can outperform purely sequential and parallel discovery. Sequential discovery harnesses which are initialized with a high-performing population (blue), via a fraction of the discovery budget yield the most performant discoveries. Predicted initialization budget (dotted line) is predicted via a convergence analysis Section 5 . Diversity threshold s=0.98 for all initial populations, population size m=15 . We notice in 9/13 scenarios, the predicted initialization budget coincides with the most performant initialization budget.
Sim thresh. ( s=0.8 )
Sim thresh. ( s=0.95 )
Sim thresh. ( s=0.98 )
Sim thresh. (Greedy, s=1 )
Task
LLM
Can’t Be Late
Gemini 3.5 Flash
-98.56 ± 0.43
-98.58 ± 0.42
-98.62 ± 0.38
-98.63 ± 0.35
Gemini 3.7 Flash Flash
-90.22 ± 0.90
-90.52 ± 0.91
-89.79 ± 0.85
-89.90 ± 0.87
CloudCast
Gemini 3.5 Flash
0.0099 ± 0.0000
0.0100 ± 0.0001
0.0099 ± 0.0000
0.0099 ± 0.0000
Gemini 3.7 Flash
0.0095 ± 0.0001
0.0095 ± 0.0001
0.0095 ± 0.0001
0.0095 ± 0.0001
TSP
GPT-OSS (120B)
-4433 ± 415
-4674 ± 422
-3993 ± 322
-4071 ± 346
Appendix
Table 11 : Comparing Population Initialization Methods across varying similarity thresholds. See population curation method description in Section 5 . Here we report the average score of a population initialization method across five initialization budgets r ; see task-specific initialization budgets in Table 7 . We find that greedy population curation is generally performant (highest or second-highest-performing population), but that minimal similarity filtering ( s=0.98 ) typically outperforms purely greedy populations. Therefore, we conclude that minimal filtering of exact and near duplicates from the population is beneficial ( s=0.98 ). Dark green shading in a cell indicates the highest-scoring population within a row; lighter green shading indicates the second-highest-scoring population. We do not conduct s=0.8 experiments for Circle Packing as more than half of the initialization population needed to be backfilled greedily due to a lack of sufficiently diverse candidates.
Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However, in practice these are composite systems combining several design choices about archives, parent selection, exploration, and budget allocation into a single recipe. Because discovery runs are expensive and inherently stochastic, existing harnesses are often compared using too few independent trials to distinguish key methodological improvements from run-to-run variance. We systematically decompose OpenEvolve-style evolutionary search and the TTT-Discover search harness into its constituent components and systematically evaluate 30 budget-matched harnesses across 12 model-problem pairs using more than 3.1 million LLM rollouts and repeated-trial statistical analysis. Our results show that discovery harnesses have a generalization problem: No fixed harness is reliably superior across the evaluated model-problem pairs, and variants of OpenEvolve generally underperform simpler alternatives. Thus, harness choice is better viewed as a hyperparameter rather than as a universal recipe, and should be tailored to the specific problem and underlying model. We also find that early discovery progress predicts final performance, and use this property to present a budget-matched adaptive-allocation experiment that starts multiple harnesses, prunes weak partial runs, and reallocates compute to stronger survivors, outperforming both commitment to a randomly sampled fixed harness and a non-adaptive harness ensemble. Together, these results motivate shifting from fixed harness selection to online adaptation guided by early performance. We release all run pools including baseline null distributions for every model-problem pair as reusable statistical infrastructure against for future harness proposals.
Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Large Language Models (LLMs) integrated into evolutionary search have recently produced state-of-the-art solutions on optimization tasks, including open mathematical conjectures, GPU kernel design, scientific law discovery, and combinatorial puzzles. To achieve this, prior work applied search scaffolds to one target task at a time, so every new problem is approached from scratch and the experience accumulated during search is discarded once the model finishes its attempt. This leaves the capability of iteratively evolving a solution (e.g., knowing which part to mutate and how, deciding when to backtrack) entirely in the scaffold rather than in the model itself. Whether the model itself could acquire this capability and reuse it across different tasks has been largely unexamined. To address this, we introduce Evolution Fine-Tuning (EFT), a mid-training paradigm that teaches LLMs to evolve solutions across tasks by converting evolutionary search trajectories into supervision. We construct Finch Collection, a 156K-trajectory dataset spanning 10 domains and 371 optimization tasks, and fine-tune open-source LLMs from 2B to 9B parameters. Empirically, EFT confers cross-task generalization: across 22 held-out tasks, our models surpass their base counterparts by 10.22% on average. Furthermore, when paired with test-time RL, our model matches state-of-the-art performance on two circle-packing tasks and outperforms its base-model counterpart on the Erdős minimum-overlap problem. EFT thus serves as a "practice phase" for general-purpose discovery agents that do not solve new problems from scratch.
Young-Jun Lee, Seungone Kim, Minki Kang +5
University of Minnesota · Carnegie Mellon University · KAIST +3
Evolutionary approaches to LLM-driven discovery often generate new programs from a small set of selected ancestors. This keeps contexts manageable but can omit useful evidence from other experiments, whereas including the full experimental history produces long, redundant contexts. We introduce a simple, single-agent discovery harness built around LabBook, an agent-maintained memory that serves two complementary roles: guiding retrieval of relevant evidence from a complete experimental log and informing the generation of new solutions. At each iteration, the same agent combines its memory with retrieved evidence and jointly produces the next program and an updated LabBook. This separates complete history retention from selective context construction, without requiring an explicit population or branching search structure. On 49 Frontier-CS problems, LabBook improves the observed quality-cost trade-off over the evaluated evolutionary baselines with two backbones, while remaining competitive across nine additional mathematical, systems, and heuristic-design tasks. Code will be released at https://github.com/BoYuanVisionary/LabBook.
Bo Yuan, Wenqian Ye, Zelin Zhao +4
Georgia Institute of Technology · University of Virginia, Charlottesville