Organizations: Center of Excellence for Generative AI, King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia. · Sakana AI, Tokyo, Japan.
LLM-based program evolution relies on evaluation feedback to guide the iterative search for high-performing programs. However, evaluation is often computationally expensive, making it essential to allocate limited resources to candidates that can most effectively advance the search. Existing LLM-based methods typically rely on fixed allocation strategies throughout the search, potentially wasting resources on low-value candidates while overlooking promising ones. We propose EvoAlloc, a self-evolving resource-allocation agent that learns from search experience to revise its strategy for allocating computational resources across candidates. EvoAlloc periodically consolidates prior search and allocation outcomes into reusable experience, which informs subsequent strategy revisions. It further uses a counterfactual exploration mechanism to occasionally evaluate candidates denied resources by the allocator, revealing their outcomes to enrich its experience for future strategy updates. Across coding and agent-harness optimization benchmarks, EvoAlloc requires 59-82% fewer full evaluations and 61-89% fewer total LLM tokens to reach baseline-level performance. Moreover, under the same full-evaluation budget, EvoAlloc achieves 8.7-12.0% higher final performance.
Figures & tables
Figure 1: EvoAlloc, a self-evolving resource allocation agent for program evolution. (a) Strategy-guided allocation decision, (b) Experience-driven Strategy evolution, and (c) counterfactual exploration jointly adapt how evaluation is allocated as search progresses.
Figure 2: Information flow between the strategy and allocation decisions in EvoAlloc. Each case reads top to bottom: the strategy and search context inform the allocation agent’s decisions, whose outcomes feed into later strategy updates. Proposals describe changes to the parent program. Matching highlights link related information across stages. Content is shortened for clarity. Strategy updates reflect feedback accumulated over multiple proposals, not the illustrated proposal alone.
Figure 3: Main results on ADAS-AIME and Circle Packing. Curves show five-seed mean best-so-far performance ( ±1 std.) against Full evaluations and cumulative LLM tokens. Bottom panels report the resources required to reach ShinkaEvolve’s final-performance and 95 %-gain targets. Token costs are the mean cumulative usage at the first target-reaching Full-evaluation index of the evaluation-aligned mean performance curve.
Figure 4: Component ablations and exploration-rate sensitivity. (a–b) Five-seed online ablations under matched Full-evaluation budgets; shading denotes ±1 standard deviation. (c) Exploration-rate sweep on Circle Packing over ρ∈{0,0.1,0.3,0.5,0.7} ; bars show the final best sum of radii, and the line shows the LLM tokens required to reach ShinkaEvolve’s final performance.
Figure 6: ADAS-AIME search trees. Both methods complete 75 full evaluations with the same seed. Shapes encode allocation decisions; colors indicate observed Full-evaluation accuracy. Gray markers denote proposals without Full outcomes; stars mark the best programs, and highlighted edges trace their lineages.
Figure 7: Allocation decisions under ablations on ADAS-AIME. Five-seed mean cumulative proposal counts by allocation decision as Full evaluations accumulate; the stacked total gives the number of proposals considered.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Evolution of allocation behavior on Circle Packing. A representative 150-full-evaluation run (seed 3) showing best-so-far performance and allocation decisions across S0–S2. Allocation panels show decisions before counterfactual exploration, with explored proposals counted by their original allocation decision. Dashed lines mark the adoption of validated Strategies.
Figure 9: Circle Packing search trees. Both methods complete 150 full evaluations with the same seed (seed 1). Shapes encode allocation decisions; colors indicate the observed sum of radii. Gray crosses denote proposals without Full outcomes; stars mark the best programs, and highlighted edges trace their lineages.
Benchmark
Method
Performance
Eval-budget AUC
Token-budget AUC
ADAS-AIME
ShinkaEvolve
35.33±2.57
31.29±1.95
30.24±2.82
One-Step Alloc
34.89±1.94
31.40±1.34
32.45±1.49
Lineage-UCB
34.00±0.89
30.88±0.53
32.32±1.07
RPM
36.22±1.13
32.49±0.69
30.31±0.94
EvoAlloc (ours)
39.56±6.80
33.69±3.33
33.15±2.89
Circle Packing
ShinkaEvolve
2.343±0.254
2.172±0.239
2.245±0.257
Appendix
Table 4: Final performance and auxiliary trajectory AUC. Five-seed mean ± one standard deviation; bold denotes the best method per benchmark and column. AUC is the normalized area under the best-so-far curve over the evaluation or token budget; higher is better.
Reference final target (100%)
95%-gain target
Method
Full evals ↓
LLM tokens (M) ↓
Full evals ↓
LLM tokens (M) ↓
ADAS-AIME
target: ≥35.33
target: ≥34.49
ShinkaEvolve
69
89.8
51
63.6
One-Step Alloc
–
–
66 ( +29% )
77.7 ( +22% )
Lineage-UCB
–
–
–
–
RPM
66 ( −4% )
140.1 ( +56% )
57 ( +12% )
117.2 ( +84% )
Appendix
Table 5: Resources to matched performance targets. First-hit costs read from the five-seed mean best-so-far curves, with mean cumulative token usage taken at the corresponding Full-evaluation index. Parentheses show changes relative to ShinkaEvolve; dashes indicate that the mean curve does not reach the target within the observed evaluation budget.
Token & proposal cost
Allocation breakdown (%)
Method
Total tokens (M)
Allocator/selector (M; % of total)
Proposal count
Direct Full
Partial → Full
Partial → Stop
Direct Discard
Explore → Full
ADAS-AIME
ShinkaEvolve
98.44
–
75.0
100.0
–
–
–
–
One-Step Alloc
92.53
0.03 ( <0.1% )
81.2
45.7
43.4
7.5
0.0
3.4
Lineage-UCB
84.93
–
95.4
78.7
–
–
21.3
–
RPM
178.27
10.98 ( 6.2 %)
1125.0
6.7
–
–
93.3
–
Appendix
Table 6: Token costs and allocation decisions. Five-seed averages at the full evaluation budget; RPM allocation statistics follow its nominal batch protocol (Section D.4 ). Allocator tokens include resource allocation, self-evolution, and shadow validation; for RPM, they are the tokens spent on pairwise selection. Dashes denote methods without an LLM-based allocator or allocation paths a method does not use. Allocation categories are mutually exclusive and sum to 100 % within each method, up to rounding, with percentages computed per seed and then averaged. Percentages describe allocator decisions rather than all executed evaluations; candidates receiving additional Full evaluations during validation are counted under their original decision categories.
Figure 10: Ablation with random allocation on Circle Packing. Random Allocation discards candidates with probability 0.5 after warm-up. Best-so-far performance is shown against Full evaluations (left) and cumulative LLM tokens (right). Lines show five-seed means, and shading denotes ±1 standard deviation. Token curves use the common observed range across methods and seeds.
Figure 11: Ablation of exploration feedback on Circle Packing. The ablated variant retains exploratory evaluations and archive updates but withholds their explicit outcome-case feedback from allocation and self-evolution. Best-so-far performance is shown against Full evaluations (left) and cumulative LLM tokens (right). Lines show five-seed means, and shading denotes ±1 standard deviation. Token curves use the common observed range across methods and seeds.
Benchmark
Strategy stage
Runs reaching the stage
Explored candidates
New-best
Improved-parent
ADAS-AIME
S0
5
10
0 ( 0.0 %)
3 ( 30.0 %)
S1
5
35
0 ( 0.0 %)
4 ( 11.4 %)
S2
4
72
5 ( 6.9 %)
18 ( 25.0 %)
S3
1
15
1 ( 6.7 %)
8 ( 53.3 %)
Total
-
132
6 ( 4.5 %)
33 ( 25.0 %)
Circle Packing
S0
5
13
2 ( 15.4 %)
7 ( 53.8 %)
Appendix
Table 7: Outcomes recovered by counterfactual exploration across Strategy stages. Counts are pooled across five seeds within each benchmark. Percentages use the number of explored candidates in each row as the denominator, rather than averaging per-seed rates. Improved-parent includes new-best.
Adoption criterion
Benchmark
Proposed challengers
Completed windows
Adopted challengers
mbest
nF
nP
ADAS-AIME
23
20
10
0
7
3
Circle Packing
44
42
11
0
11
0
Appendix
Table 8: Strategy validation outcomes aggregated over five seeds per benchmark. The final three columns count adoptions decided by fewer missed new-best candidates ( mbest ), fewer Full evaluations ( nF ), or fewer Partial evaluations ( nP ).
Context
Information
Shared by resource allocation and self-evolution
Benchmark
Evaluation protocol and available actions
Allocation guidance
Current Strategy and Experiences
Recent cases
Recently observed cases, as case cards (Appendix E.4 )
Search progress
Budget usage and recent decision counts
Additional context for resource allocation
Appendix
Table 9: Input context for resource allocation and self-evolution.