Organizations: Center of Excellence for Generative AI, King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia. · Sakana AI, Tokyo, Japan.
LLM-based program evolution relies on evaluation feedback to guide the iterative search for high-performing programs. However, evaluation is often computationally expensive, making it essential to allocate limited resources to candidates that can most effectively advance the search. Existing LLM-based methods typically rely on fixed allocation strategies throughout the search, potentially wasting resources on low-value candidates while overlooking promising ones. We propose EvoAlloc, a self-evolving resource-allocation agent that learns from search experience to revise its strategy for allocating computational resources across candidates. EvoAlloc periodically consolidates prior search and allocation outcomes into reusable experience, which informs subsequent strategy revisions. It further uses a counterfactual exploration mechanism to occasionally evaluate candidates denied resources by the allocator, revealing their outcomes to enrich its experience for future strategy updates. Across coding and agent-harness optimization benchmarks, EvoAlloc requires 59-82% fewer full evaluations and 61-89% fewer total LLM tokens to reach baseline-level performance. Moreover, under the same full-evaluation budget, EvoAlloc achieves 8.7-12.0% higher final performance.
Figures & tables
Figure 1: EvoAlloc, a self-evolving resource allocation agent for program evolution. (a) Strategy-guided allocation decision, (b) Experience-driven Strategy evolution, and (c) counterfactual exploration jointly adapt how evaluation is allocated as search progresses.
Figure 2: Information flow between the strategy and allocation decisions in EvoAlloc. Each case reads top to bottom: the strategy and search context inform the allocation agent’s decisions, whose outcomes feed into later strategy updates. Proposals describe changes to the parent program. Matching highlights link related information across stages. Content is shortened for clarity. Strategy updates reflect feedback accumulated over multiple proposals, not the illustrated proposal alone.
Figure 3: Main results on ADAS-AIME and Circle Packing. Curves show five-seed mean best-so-far performance ( ±1 std.) against Full evaluations and cumulative LLM tokens. Bottom panels report the resources required to reach ShinkaEvolve’s final-performance and 95 %-gain targets. Token costs are the mean cumulative usage at the first target-reaching Full-evaluation index of the evaluation-aligned mean performance curve.
Figure 4: Component ablations and exploration-rate sensitivity. (a–b) Five-seed online ablations under matched Full-evaluation budgets; shading denotes ±1 standard deviation. (c) Exploration-rate sweep on Circle Packing over ρ∈{0,0.1,0.3,0.5,0.7} ; bars show the final best sum of radii, and the line shows the LLM tokens required to reach ShinkaEvolve’s final performance.
Figure 6: ADAS-AIME search trees. Both methods complete 75 full evaluations with the same seed. Shapes encode allocation decisions; colors indicate observed Full-evaluation accuracy. Gray markers denote proposals without Full outcomes; stars mark the best programs, and highlighted edges trace their lineages.
Figure 7: Allocation decisions under ablations on ADAS-AIME. Five-seed mean cumulative proposal counts by allocation decision as Full evaluations accumulate; the stacked total gives the number of proposals considered.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Evolution of allocation behavior on Circle Packing. A representative 150-full-evaluation run (seed 3) showing best-so-far performance and allocation decisions across S0–S2. Allocation panels show decisions before counterfactual exploration, with explored proposals counted by their original allocation decision. Dashed lines mark the adoption of validated Strategies.
Figure 9: Circle Packing search trees. Both methods complete 150 full evaluations with the same seed (seed 1). Shapes encode allocation decisions; colors indicate the observed sum of radii. Gray crosses denote proposals without Full outcomes; stars mark the best programs, and highlighted edges trace their lineages.
Benchmark
Method
Performance
Eval-budget AUC
Token-budget AUC
ADAS-AIME
ShinkaEvolve
35.33±2.57
31.29±1.95
30.24±2.82
One-Step Alloc
34.89±1.94
31.40±1.34
32.45±1.49
Lineage-UCB
34.00±0.89
30.88±0.53
32.32±1.07
RPM
36.22±1.13
32.49±0.69
30.31±0.94
EvoAlloc (ours)
39.56±6.80
33.69±3.33
33.15±2.89
Circle Packing
ShinkaEvolve
2.343±0.254
2.172±0.239
2.245±0.257
Appendix
Table 4: Final performance and auxiliary trajectory AUC. Five-seed mean ± one standard deviation; bold denotes the best method per benchmark and column. AUC is the normalized area under the best-so-far curve over the evaluation or token budget; higher is better.
Reference final target (100%)
95%-gain target
Method
Full evals ↓
LLM tokens (M) ↓
Full evals ↓
LLM tokens (M) ↓
ADAS-AIME
target: ≥35.33
target: ≥34.49
ShinkaEvolve
69
89.8
51
63.6
One-Step Alloc
–
–
66 ( +29% )
77.7 ( +22% )
Lineage-UCB
–
–
–
–
RPM
66 ( −4% )
140.1 ( +56% )
57 ( +12% )
117.2 ( +84% )
Appendix
Table 5: Resources to matched performance targets. First-hit costs read from the five-seed mean best-so-far curves, with mean cumulative token usage taken at the corresponding Full-evaluation index. Parentheses show changes relative to ShinkaEvolve; dashes indicate that the mean curve does not reach the target within the observed evaluation budget.
Token & proposal cost
Allocation breakdown (%)
Method
Total tokens (M)
Allocator/selector (M; % of total)
Proposal count
Direct Full
Partial → Full
Partial → Stop
Direct Discard
Explore → Full
ADAS-AIME
ShinkaEvolve
98.44
–
75.0
100.0
–
–
–
–
One-Step Alloc
92.53
0.03 ( <0.1% )
81.2
45.7
43.4
7.5
0.0
3.4
Lineage-UCB
84.93
–
95.4
78.7
–
–
21.3
–
RPM
178.27
10.98 ( 6.2 %)
1125.0
6.7
–
–
93.3
–
Appendix
Table 6: Token costs and allocation decisions. Five-seed averages at the full evaluation budget; RPM allocation statistics follow its nominal batch protocol (Section D.4 ). Allocator tokens include resource allocation, self-evolution, and shadow validation; for RPM, they are the tokens spent on pairwise selection. Dashes denote methods without an LLM-based allocator or allocation paths a method does not use. Allocation categories are mutually exclusive and sum to 100 % within each method, up to rounding, with percentages computed per seed and then averaged. Percentages describe allocator decisions rather than all executed evaluations; candidates receiving additional Full evaluations during validation are counted under their original decision categories.
Figure 10: Ablation with random allocation on Circle Packing. Random Allocation discards candidates with probability 0.5 after warm-up. Best-so-far performance is shown against Full evaluations (left) and cumulative LLM tokens (right). Lines show five-seed means, and shading denotes ±1 standard deviation. Token curves use the common observed range across methods and seeds.
Figure 11: Ablation of exploration feedback on Circle Packing. The ablated variant retains exploratory evaluations and archive updates but withholds their explicit outcome-case feedback from allocation and self-evolution. Best-so-far performance is shown against Full evaluations (left) and cumulative LLM tokens (right). Lines show five-seed means, and shading denotes ±1 standard deviation. Token curves use the common observed range across methods and seeds.
Benchmark
Strategy stage
Runs reaching the stage
Explored candidates
New-best
Improved-parent
ADAS-AIME
S0
5
10
0 ( 0.0 %)
3 ( 30.0 %)
S1
5
35
0 ( 0.0 %)
4 ( 11.4 %)
S2
4
72
5 ( 6.9 %)
18 ( 25.0 %)
S3
1
15
1 ( 6.7 %)
8 ( 53.3 %)
Total
-
132
6 ( 4.5 %)
33 ( 25.0 %)
Circle Packing
S0
5
13
2 ( 15.4 %)
7 ( 53.8 %)
Appendix
Table 7: Outcomes recovered by counterfactual exploration across Strategy stages. Counts are pooled across five seeds within each benchmark. Percentages use the number of explored candidates in each row as the denominator, rather than averaging per-seed rates. Improved-parent includes new-best.
Adoption criterion
Benchmark
Proposed challengers
Completed windows
Adopted challengers
mbest
nF
nP
ADAS-AIME
23
20
10
0
7
3
Circle Packing
44
42
11
0
11
0
Appendix
Table 8: Strategy validation outcomes aggregated over five seeds per benchmark. The final three columns count adoptions decided by fewer missed new-best candidates ( mbest ), fewer Full evaluations ( nF ), or fewer Partial evaluations ( nP ).
Context
Information
Shared by resource allocation and self-evolution
Benchmark
Evaluation protocol and available actions
Allocation guidance
Current Strategy and Experiences
Recent cases
Recently observed cases, as case cards (Appendix E.4 )
Search progress
Budget usage and recent decision counts
Additional context for resource allocation
Appendix
Table 9: Input context for resource allocation and self-evolution.
LLM-guided evolutionary search (Evolve systems) has reached state-of-the-art results on mathematical and combinatorial tasks, yet most existing systems report only the best of many runs and leave the run-to-run distribution undocumented. We ask how a fixed budget of LLM calls should be allocated, and how reliably a single run reaches the reported numbers. Sweeping the depth-breadth grid over five models and three tasks, we identify two empirical regularities: a fitness-compute envelope along which capability ordering largely collapses on effective FLOPs, and a bilinear depth-breadth fit with task-specific interaction; both are gated by model-task capability. Motivated by these regularities, we propose BaSE (Bandit-based Self-Evolving), a multi-armed bandit that allocates LLM calls across parallel trajectories. Without changing the model, prompt, or evaluator, BaSE improves mean fitness by 12.3% over the strongest island-protocol baseline across 8 (model, task) cells, with the largest gains on high-variance settings: a reliability gain from allocation alone.
Sixue Xing, Haoyu He, Kerui Wu +4
University of Notre Dame · 2Northeastern University · University of Massachusetts Amherst +4
Large language model (LLM)-driven evolution has shown promise for program search and algorithm discovery, but relying on strong models throughout long evolutionary runs is costly. A natural alternative is to combine cheap and strong models under a fixed inference budget. However, existing approaches typically allocate models at the level of individual queries or mutation steps, overlooking that evolutionary search is \textit{stateful}: each generated candidate changes the population from which subsequent mutations are produced. We empirically analyze LLM-driven evolutionary trajectories and find that search progress is strongly front-loaded, early trajectory performance is informative but noisy, and cheap models recover much of the early progress achieved by strong models at lower cost. Motivated by these findings, we propose \textbf{\model}, a training-free framework that shifts budget allocation from individual calls to evolving populations through adaptive \textit{population handoff}. A cheap model explores multiple trajectories in short blocks allocated by a bandit scheduler. Relay Gain, defined as the marginal improvement of a compact, quality-diverse candidate bank constructed for handoff, serves as the scheduler reward and determines when to hand off. The curated candidates initialize a shared strong model population for refinement. Across four benchmarks and three budgets, \model achieves the highest mean score in 11 of 12 settings, outperforming competitive baselines. Our results suggest that in stateful search, budget allocation should be organized around the population, not the individual call.
Sichun Luo, Yi Huang, Guanzhi Deng +6
1The University of Hong Kong · 2JIUTIAN Research, China Mobile · 3City University of Hong Kong +1
Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some cases may transfer across related tasks and domains. However, existing LLM-driven evolutionary frameworks largely discard such knowledge, repeatedly rediscovering similar ideas and limiting opportunities for cross-run and cross-task learning. We introduce EvoMem, a persistent memory architecture for LLM-based evolutionary program search that captures and reuses candidate mutation knowledge. EvoMem converts successful mutation events into structured, task-aware advice for future runs. It operates in two phases: after each run, it extracts and stores promising ideas with provenance, and during subsequent evolution, it retrieves a small set of relevant instructions based on the current task and program context to guide mutation. Across geometric optimization, multi-hop question answering, GPU kernel optimization, and related benchmarks, our experiments show positive average improvements in target metrics or search speed for most evaluated settings, while also revealing variability across tasks. Overall, EvoMem provides evidence that persistent memory can reduce some redundant exploration and improve the reuse and adaptation of successful strategies in LLM-driven evolutionary search.
Viktor Volkov, Valentin Khrulkov, Andrey V. Galichin +8
AXXX, Moscow, Russia · Applied AI Institute, Moscow, Russia · Lomonosov Moscow State University, Moscow, Russia