SimpleEvol: An Agent-Loop Framework for LLM-Driven Automated Heuristic Design with Minimal Human Priors
Authors: Jianghan Zhu, Cong Zhang, Rongjie Zhu, Chi Zhang, Zhiguang Cao
Organizations: School of Computing and Information Systems, Singapore Management University, Singapore · Independent Researcher · School of Teacher Education, Nanjing University of Information Science and Technology, China · Department of Mathematics, National University of Singapore, Singapore
Large language models (LLMs) have emerged as powerful tools for automated heuristic design (AHD), enabling iterative generation and refinement of heuristics. However, the dominant paradigm embeds LLMs as narrow, fixed components, such as crossover or mutation, within heavily hand-engineered evolutionary frameworks. We argue this misapprehends LLMs. It treats them as specialized tools rather than general reasoners, constrains them to low-level operations, and underutilizes their autonomy. Moreover, the extensive human priors in these frameworks violate the bitter lesson principle that general methods scaling with computation surpass hand-crafted solutions. This raises a key question: which AHD framework designs best convert stronger LLM capabilities into better heuristics? To address this, we propose metrics for LLM-driven AHD framework handcraftedness (AHI) and intelligence conversion efficiency (ICE). Evaluating ten LLMs across three challenging combinatorial optimization problems, we obtain a notable finding that frameworks with fewer human priors consistently yield higher ICE. Based on this finding, we propose SimpleEvol, an agent-loop framework for AHD which removes nearly all human priors and allows the LLM to operate autonomously. SimpleEvol consistently achieves the highest ICE, often by a large margin. Our results challenge the trend toward complex AHD pipelines and point to a lighter and more model-centric alternative, suggesting that reducing human priors is a more effective strategy to scale up with model intelligence. The source code is available at https://github.com/HenryZhu1029/SimpleEvol-Master.
Figures & tables
Framework
M
K
log10(1+Q)
AHI
FunSearch
3
1
2.915
6.915
EoH
2
4
2.919
8.919
ReEvo
2
4
3.112
9.112
SimpleEvol (ours)
1
2
3.001
6.001
Table 1: AHI of different AHD frameworks.
Figure 1 : Overview of SimpleEvol.
Figure 2 : Relationship between model intelligence and AHD performance on TSP and CVRP.
Method
AHI
ICE(TSP)
ICE(CVRP)
FunSearch
6.915
1.8174
5.4381
EoH
8.919
1.5082
3.6572
ReEvo
9.112
0.8230
2.0824
SimpleEvol (ours)
6.001
2.1941
6.2128
Table 2: ICE on TSP Constructive and CVRP-ACO, together with handcraftedness index (AHI).
Figure 3 : Relative framework advantage under different model intelligence. The upper-left number in each cell denotes the performance rank on each backbone model.
Task
N=50
N=100
N=200
Fun.
EoH
ReEvo
Ours
Fun.
EoH
ReEvo
Ours
Fun.
EoH
ReEvo
Ours
TSP-SC.
5.50%
6.76%
7.98%
4.77%
7.33%
8.44%
10.06%
6.47%
10.35%
10.56%
12.10%
9.49%
CVRP-ACO
1.04%
0.71%
0.49%
0.34%
4.83%
6.70%
5.98%
4.31%
4.46%
4.28%
4.38%
4.26%
Table 3: Performance comparison of AHD frameworks using GPT-5-mini . Step-by-step construction is abbreviated as SC. FunSearch [ 53 ] is abbreviated as Fun. The best result of each setting is in bold.
Figure 7
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Template of the system generator prompt.
Figure 7 : Template of the summary user prompt.
Figure 8 : Template of the summary injection prompt.
Figure 9 : Template of the user generator prompt.
Figure 10 : Meta-information feedback of heuristic evaluation.
Figure 11 : Template of the initial user prompt.
Framework
Problem
Dtrain
Dtest
GLS
FSSP
Number of Iterations: 1000
Same as Dtrain
ACO
CVRP
Number of Ants: 30; Number of Iterations: 100
Same as Dtrain
Appendix
Table 4 : Search framework hyperparameters for different problem settings.
Benchmark
Model
MMLU-Pro
IFBench
AIME 2025
LiveCodeBench
I(m)
Score
Ratio
Score
Ratio
Score
Ratio
Score
Ratio
GPT-4o-mini
64.8
1.00
31.0
1.00
14.7
1.00
23.4
1.00
1.00
GPT-4.1-nano
65.7
1.04
32.0
1.03
24.0
1.63
32.6
1.39
1.25
Claude Sonnet 3.7
80.3
1.27
44.0
1.42
21.0
1.43
39.4
1.68
1.44
DeepSeek-v3
81.9
1.30
41.0
1.32
41.0
2.79
40.5
1.73
1.70
Appendix
Table 5 : Benchmark scores (Score) and normalized ratios (Ratio) used to compute the model intelligence score I(m) . Ratios are normalized with respect to the weakest model (GPT-4o-mini), and I(m) is computed based on aggregated normalized performance across benchmarks.
Instance
EoH
ReEvo
FunSearch
SimpleEvol (ours)
bier127.tsp
13.60%
8.33%
20.69%
7.16%
ch130.tsp
10.63%
12.24%
20.04%
10.12%
eil51.tsp
8.85%
14.20%
7.27%
11.96%
fl417.tsp
13.49%
10.59%
33.06%
10.50%
Appendix
Table 6 : Results of LLM-based AHD methods for TSP on TSPLIB instances. Following the same protocol as ReEvo [ 52 ] , we compute the reported optimality gap from the average over three runs with different starting nodes. For each instance, the best result across all methods is marked in bold .
Method
TSP Constructive
CVRP-ACO
ICE
LOO-mean
LOO-std
ICEAA
ICE
LOO-mean
LOO-std
ICEAA
FunSearch
1.8174
1.8110
0.4258
2.2103
5.4381
5.4112
0.9423
6.8866
EoH
1.5082
1.4969
0.4040
1.8132
3.6572
3.6853
0.9137
5.5700
ReEvo
0.8230
0.8201
0.2981
1.4627
2.0824
2.0223
1.1802
5.3194
Appendix
Table 7 : Robustness of AHI–ICE relationship under leave-one-out regression and an alternative intelligence metric. ICEAA is computed by replacing our benchmark-based I(m) with the Artificial Analysis Intelligence Index [ 49 ] . The highest value of ICE for each setting is marked in bold, while the second largest value is underlined.
Task
Comparison
Δ
Pr(Δ>0)
TSP Constructive
SimpleEvol − FunSearch
0.380
85.5%
SimpleEvol − EoH
0.688
95.3%
SimpleEvol − ReEvo
1.369
97.6%
SimpleEvol − MCTS-AHD
1.442
98.9%
CVRP-ACO
SimpleEvol − FunSearch
0.771
79.1%
SimpleEvol − EoH
2.562
88.7%
Appendix
Table 8 : Bayesian-bootstrap uncertainty of pairwise ICE differences. Δ denotes ICESimpleEvol−ICEA . Pr(Δ>0) is estimated from 10,000 Bayesian-bootstrap replicates.
Figure 12 : Best-so-far search trajectories under GPT-5-mini . Lower gap is better. Solid curves show the mean across independent runs, and shaded regions indicate ± 1 standard deviation. Only the CVRP-ACO panel uses a symmetric-log scale to accommodate the wide range of early-stage gaps and negative gaps relative to the reference solution.
Task
Method
50
100
200
400
600
TSP Constructive
ReEvo
0.585
0.783
0.891
0.981
0.843
EoH
1.393
0.568
1.046
1.290
1.582
FunSearch
0.482
0.862
1.328
1.315
1.776
SimpleEvol
0.936
0.823
1.403
1.835
2.436
CVRP-ACO
ReEvo
1.012
1.158
1.420
1.892
2.225
EoH
0.213
1.318
2.868
2.560
3.215
Appendix
Table 9 : ICE at different heuristic-evaluation budgets. The highest ICE at each checkpoint within each task is shown in bold.
Figure 13 : Relationship between model intelligence and AHD performance on FSSP-GLS. Left: 64-instance held-out IDD synthetic instances. Right: Taillard benchmark instances.
Task
Method
Runtime (hrs)
Input Tokens (M)
Output Tokens (M)
# Queries
TSP Constructive
FunSearch
11.13
1.111
1.479
821
EoH
5.23
0.687
1.100
828
ReEvo
3.60
3.602
2.249
1293
SimpleEvol (ours)
10.55
4.517
1.988
1001
CVRP-ACO
FunSearch
16.86
1.826
2.008
822
Appendix
Table 10 : Training time, token usage, and LLM query counts of all methods averaged across 10 models.
Figure 14 : Computational cost comparison of LLM-based AHD methods.
LLM Model
N=50
N=100
N=200
Avg Gap ↓
P(m)↑
Obj.
Gap
Obj.
Gap
Obj.
Gap
Opt (best known)
5.6750
–
7.7680
–
10.6590
–
–
1.0000
GPT-4o-mini
6.3438
11.79%
8.8025
13.32%
12.4151
16.48%
13.86%
7.2152
GPT-4.1-nano
6.2518
10.16%
8.7732
12.94%
12.2825
15.23%
12.78%
7.8256
Claude Sonnet 3.7
6.1693
8.71%
8.7684
12.88%
12.4542
16.84%
12.81%
7.8065
Appendix
Table 11 : Detailed performance of SimpleEvol across backbone models on the TSP test sets.
LLM Model
N=50
N=100
N=200
Avg Gap ↓
P(m)↑
Obj.
Gap
Obj.
Gap
Obj.
Gap
Opt (best known)
5.6750
–
7.7680
–
10.6590
–
–
1.0000
GPT-4o-mini
6.3942
12.67%
8.8689
14.17%
12.4607
16.90%
14.58%
6.8574
GPT-4.1-nano
6.3294
11.53%
8.8852
14.38%
12.3365
15.74%
13.88%
7.2029
Claude Sonnet 3.7
6.1857
9.00%
8.7873
13.12%
12.4140
16.47%
12.86%
7.7749
Appendix
Table 12 : Detailed performance of FunSearch across backbone models on the TSP test sets.
LLM Model
N=50
N=100
N=200
Avg Gap ↓
P(m)↑
Obj.
Gap
Obj.
Gap
Obj.
Gap
Opt (best known)
5.6750
–
7.7680
–
10.6590
–
–
1.0000
GPT-4o-mini
6.5120
14.75%
9.1810
18.19%
12.7930
20.02%
17.65%
5.6647
GPT-4.1-nano
6.2992
11.00%
8.7966
13.24%
12.2683
15.10%
13.11%
7.6261
Claude Sonnet 3.7
6.2576
10.27%
8.7952
13.22%
12.5031
17.30%
13.60%
7.3546
Appendix
Table 13 : Detailed performance of EoH across backbone models on the TSP test sets.
LLM Model
N=50
N=100
N=200
Avg Gap ↓
P(m)↑
Obj.
Gap
Obj.
Gap
Obj.
Gap
Opt (best known)
5.6750
–
7.7680
–
10.6590
–
–
1.0000
GPT-4o-mini
6.3526
11.94%
8.8555
14.00%
12.4488
16.79%
14.24%
7.0207
GPT-4.1-nano
6.3886
12.58%
9.0631
16.67%
12.6207
18.40%
15.88%
6.2957
Claude Sonnet 3.7
6.2048
9.34%
8.6615
11.50%
12.2265
14.71%
11.85%
8.4402
Appendix
Table 14 : Detailed performance of ReEvo across backbone models on the TSP test sets.
LLM Model
N=50
N=100
N=200
Avg Gap ↓
P(m)↑
Obj.
Gap
Obj.
Gap
Obj.
Gap
Opt (best known)
5.6750
–
7.7680
–
10.6590
–
–
1.0000
GPT-4o-mini
6.2326
9.83%
8.7412
12.53%
12.2513
14.94%
12.43%
8.0446
GPT-4.1-nano
6.2886
10.81%
9.0136
16.04%
12.5707
17.94%
14.93%
6.6989
Claude Sonnet 3.7
6.2341
9.85%
8.6815
11.76%
12.2565
14.99%
12.20%
8.1968
Appendix
Table 15 : Detailed performance of MCTS-AHD across backbone models on the TSP test sets.
LLM Model
N=50
N=100
N=200
Avg Gap ↓
P(m)↑
Obj.
Gap
Obj.
Gap
Obj.
Gap
Opt (best known)
8.8880
–
14.9320
–
27.1590
–
–
1.0000
GPT-4o-mini
9.2702
4.30%
16.1382
8.08%
28.6203
5.38%
5.92%
16.8932
GPT-4.1-nano
9.1717
3.19%
16.1514
8.17%
28.9365
6.54%
5.97%
16.7569
Claude Sonnet 3.7
9.2351
3.90%
15.9603
6.89%
28.3293
4.31%
5.03%
19.8669
Appendix
Table 16 : Detailed performance of SimpleEvol across backbone models on the CVRP test sets.
LLM Model
N=50
N=100
N=200
Avg Gap ↓
P(m)↑
Obj.
Gap
Obj.
Gap
Obj.
Gap
Opt (best known)
8.8880
–
14.9320
–
27.1590
–
–
1.0000
GPT-4o-mini
9.2106
3.63%
16.5272
10.68%
29.5915
8.96%
7.76%
12.8926
GPT-4.1-nano
9.3651
5.37%
16.5791
11.03%
29.1303
7.26%
7.89%
12.6813
Claude Sonnet 3.7
9.0786
2.14%
16.1344
8.05%
28.8580
6.26%
5.48%
18.2346
Appendix
Table 17 : Detailed performance of FunSearch across backbone models on the CVRP test sets.
LLM Model
N=50
N=100
N=200
Avg Gap ↓
P(m)↑
Obj.
Gap
Obj.
Gap
Obj.
Gap
Opt (best known)
8.8880
–
14.9320
–
27.1590
–
–
1.0000
GPT-4o-mini
9.1090
2.49%
16.1046
7.85%
28.4425
4.73%
5.02%
19.9131
GPT-4.1-nano
9.2774
4.38%
16.5862
11.08%
29.2250
7.61%
7.69%
13.0055
Claude Sonnet 3.7
8.9577
0.78%
15.8456
6.12%
28.6107
5.35%
4.08%
24.4939
Appendix
Table 18 : Detailed performance of EoH across backbone models on the CVRP test sets.
LLM Model
N=50
N=100
N=200
Avg Gap ↓
P(m)↑
Obj.
Gap
Obj.
Gap
Obj.
Gap
Opt (best known)
8.8880
–
14.9320
–
27.1590
–
–
1.0000
GPT-4o-mini
9.3572
5.28%
16.1092
7.88%
29.2078
7.54%
6.90%
14.4882
GPT-4.1-nano
9.1683
3.15%
16.1404
8.09%
28.9847
6.72%
5.99%
16.6958
Claude Sonnet 3.7
9.0870
2.24%
15.9228
6.64%
28.2793
4.12%
4.33%
23.0788
Appendix
Table 19 : Detailed performance of ReEvo across backbone models on the CVRP test sets.
LLM Model
N=50
N=100
N=200
Avg Gap ↓
P(m)↑
Obj.
Gap
Obj.
Gap
Obj.
Gap
Opt (best known)
8.8880
–
14.9320
–
27.1590
–
–
1.0000
GPT-4o-mini
9.2480
4.05%
15.7320
5.36%
28.3610
4.43%
4.61%
21.6855
GPT-4.1-nano
9.1572
3.03%
16.0282
7.34%
29.1074
7.17%
5.85%
17.0997
Claude Sonnet 3.7
9.1691
3.16%
15.7503
5.48%
28.2169
3.90%
4.18%
23.9271
Appendix
Table 20 : Detailed performance of MCTS-AHD across backbone models on the CVRP test sets.
LLM Model
Avg Gap (IDD) ↓
P(m)idd↑
Avg Gap (Taillard) ↓
P(m)OOD↑
LLM-based AHD: FunSearch
GPT-4o-mini
3.07%
32.5862
0.33%
307.5031
GPT-4.1-nano
3.02%
33.1680
0.27%
366.1662
Appendix
Table 21 : Performance comparison across AHD frameworks on IDD and Taillard (OOD) instances.
Resources
Type
License / Availability
URL
LKH3
Code
Available for academic research use
http://webhotel4.ruc.dk/ keld/research/LKH-3/
POMO
Code
Available online
https://github.com/yd-kwon/POMO/tree/master
DeepACO
Code
MIT License
https://github.com/henry-yeh/DeepACO
FunSearch
Code
Apache License
https://github.com/google-deepmind/funsearch
EoH
Code
MIT License
https://github.com/FeiLiu36/EoH
ReEvo
Code
MIT License
https://github.com/ai4co/reevo
Appendix
Table 22 : Licenses for codebases, datasets, and public leaderboard for benchmark scores.
School of Computer Science, University of Nottingham Ningbo China, Ningbo, China · School of Computer Science, University of Nottingham, Nottingham, UK
Guangdong Provincial Key Laboratory of Brain-Inspired Intelligent Computation, Department of Computer Science and Engineering, Southern University of Science and Technology · Zhongguancun Academy · The Hong Kong University of Science and Technology