Evolutionary approaches to LLM-driven discovery often generate new programs from a small set of selected ancestors. This keeps contexts manageable but can omit useful evidence from other experiments, whereas including the full experimental history produces long, redundant contexts. We introduce a simple, single-agent discovery harness built around LabBook, an agent-maintained memory that serves two complementary roles: guiding retrieval of relevant evidence from a complete experimental log and informing the generation of new solutions. At each iteration, the same agent combines its memory with retrieved evidence and jointly produces the next program and an updated LabBook. This separates complete history retention from selective context construction, without requiring an explicit population or branching search structure. On 49 Frontier-CS problems, LabBook improves the observed quality-cost trade-off over the evaluated evolutionary baselines with two backbones, while remaining competitive across nine additional mathematical, systems, and heuristic-design tasks. Code will be released at https://github.com/BoYuanVisionary/LabBook.
Figures & tables
Figure 1: Overview of the LabBook harness. At each iteration, a single agent combines the task, the latest evaluated attempt, and the compact LabBook state; it may retrieve exact evidence from the append-only experimental log before jointly producing the next program and a rewritten LabBook. Evaluation appends a new record to the log, closing the discovery loop.
Figure 2: Cost–performance on Frontier-CS with DeepSeek V4 Flash (left) and Qwen 3.6 Flash (right), averaged over three independent runs. Each point averages the best-so-far score and cumulative cost across the 49 problems. LabBook is shown after 5 iterations and every 10 iterations thereafter; the baselines are shown every 10 iterations.
Task
Method
DeepSeek V4 Flash
Qwen 3.6 Flash
Score
Cost ↓
Score
Cost ↓
Circle Packing
OpenEvolve
0.9288
0.283
0.8880
0.395
EvoX
0.8583
0.229
0.7942
0.301
AdaEvolve
0.9934
0.562
0.9510
0.403
LabBook (Ours)
1.0004
0.207
0.9347
0.816
Heilbronn Convex
OpenEvolve
0.7283
0.205
0.7519
0.548
Table 1: Comprehensive evaluation on nine cross-domain discovery tasks. Score reports the task-native objective, and cost is the cumulative measured API cost in USD at the reported checkpoint. Higher scores are better for the first six tasks. For the final three tasks, E-Graph Extraction reports total extraction cost, Operator Scheduling reports total latency, and Pedigree reports total genotype changes; lower is better for all three. Bold denotes the best score for each task and backbone, and underlining denotes the second-best distinct score. All methods use 50 iterations.
Ours
No Retrieval
No Persistent
Best-of-N
Top-score
Full History
Normalized mean
0.9982
0.9057
0.9789
0.8733
0.9411
0.9488
Mean cost (USD)
0.140
0.148
0.143
0.118
0.150
0.315
Table 2: Aggregate memory and retrieval ablation over nine tasks at 30 iterations. Normalization is performed independently per task before averaging.
Figure 3: Memory and retrieval ablation over 30 iterations with DeepSeek V4 Flash. Full LabBook is shown as a solid line; component ablations, Best-of-N, Top-score Retrieval, and Full History use distinct dashed styles. Markers denote iterations 5, 10, 20, and 30.
Figure 4: Iteration–performance trajectory for the Heilbronn Convex case study. The vertical axis begins at 0.85 to emphasize improvements beyond the initial naive-sampling regime. LabBook records and repairs an early implementation failure, preserves iteration 33 as a reliable backbone, and retrieves that program 44 iterations later to produce a new best at iteration 77.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Count
Description
AlphaEvolve ( Novikov et al., 2025 )
circle_packing
1
Place 26 non-overlapping circles in a unit square to maximize the sum of their radii.
heilbronn_convex
1
Place points in a convex region to maximize the minimum area of any triangle formed by three points.
ADRS ( Cheng et al., 2025 )
eplb
1
Replicate and place mixture-of-experts experts to balance load across GPUs.
prism
1
Place multiple models on GPUs to minimize worst-case memory pressure.
Appendix
Table 3: Algorithmic and research problem benchmarks. We evaluate 58 open-ended problems: two mathematical problems studied by AlphaEvolve, three systems problems from ADRS, four heuristic-design problems from HeuriGym, and 49 problems sampled from Frontier-CS. Frontier-CS categories follow EvoX ( Liu et al., 2026a ) .
Table 4: The 49 selected Frontier-CS problems, grouped using the category definitions adopted by EvoX ( Liu et al., 2026a ) .
Backbone
Method
Iter.
Fresh in
Cached in
Output
Total
Cost
Mean score
DeepSeek V4 Flash
OpenEvolve
50
817K
164K
292K
1,273K
$0.298
33.870
EvoX
50
879K
197K
283K
1,358K
$0.302
33.988
AdaEvolve
60
745K
476K
391K
1,612K
$0.348
37.606
LabBook
30
299K
1,426K
427K
2,152K
$0.306
39.572
Qwen 3.6 Flash
OpenEvolve
100
2,364K
0
555K
2,919K
$1.068
28.048
EvoX
100
2,296K
0
484K
2,781K
$0.976
24.992
Appendix
Table 5: Mean per-problem Frontier-CS token usage at the reported checkpoints. Token counts are rounded to the nearest thousand and cost is in USD.
Task
Experimental log (kB)
LabBook (kB)
∣H50∣/∣m50∣
Iter. 10
Iter. 30
Iter. 50
Iter. 10
Iter. 30
Iter. 50
Frontier-CS (mean)
124.9
329.4
543.3
4.20
4.84
4.89
111.2 ×
Circle Packing
79.6
288.3
581.4
1.34
1.75
2.26
257.8 ×
Heilbronn Convex
67.0
257.7
460.8
0.96
2.23
1.86
247.5 ×
EPLB
147.7
351.7
548.1
3.11
1.46
1.38
397.7 ×
PRISM
117.4
392.6
609.0
2.03
2.07
1.89
322.0 ×
Appendix
Table 6: Growth of the complete experimental log and the compact persistent memory over 50 iterations. Sizes are decimal kilobytes. The final column is the ratio between the cumulative log and LabBook at iteration 50.
Backbone
Method
Cost (USD)
Median score
DeepSeek V4 Flash
LabBook
0.463
43.065
AdaEvolve
0.348
31.792
EvoX
0.370
27.054
OpenEvolve
0.365
27.708
Qwen 3.6 Flash
LabBook
1.001
37.000
AdaEvolve
0.821
22.715
Appendix
Table 7: Median best-so-far Frontier-CS score at the final checkpoint shown for each method. We compute the median over problems within each run and then average across runs. Scores use the benchmark’s 0–100 scale.