Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments still rely on hand-written rules that operators can read, audit, and execute within tight latency budgets. LLM-based Automatic Heuristic Design (AHD) promises to automate writing such rules. However, existing AHD frameworks were developed for combinatorial problems fully specified to the LLM, and they learn only from a scalar fitness score. In real systems, the behaviour that determines a good heuristic, such as processor speeds or power consumption, is unknown a priori: the score reveals which heuristic performs better, but not why. This missing information is recorded in the system logs that every evaluation produces. Exploiting it is non-trivial: logs are massive and noisy, the relevant signals depend on the objective, and their content and format vary across hardware and software stacks, so they can neither be fed to an LLM as is nor processed by a fixed parser. We propose TRACE, which couples an evolutionary AHD loop with an agentic knowledge-extraction workflow. A Reasoner agent analyzes the log schema in light of the objective and formulates hypotheses about the system dynamics; a Coder agent writes and executes schema-specific code to test them, producing insights or executable tools for the evolved heuristics. We evaluate TRACE on a synthetic cloud benchmark and a 5G vRAN scenario built from industrial testbed measurements and operational traffic traces. TRACE consistently outperforms state-of-the-art AHD methods in resource assignment problems and yields more auditable heuristics at under 2% overhead.
Figures & tables
Figure 1: Scalar-feedback AHD (left) learns what works from a single fitness value. TRACE (right) distills system logs into knowledge about why it works (e.g., per-processor speed and energy profiles) and feeds it back into heuristic generation.
Figure 2: Overview of TRACE. Right: the generation loop evolves heuristics; each evaluation returns a fitness J(g) , stored in the algorithm database, and an execution log Λ(g) , stored in a log buffer. Left: at the start of each epoch e , the Agentic Workflow distills the log buffer into Diagnostic Knowledge Ke , merged into the System Information of the generation prompt.
Figure 3: Two runs of the Agentic Workflow. Top: the Reasoner hypothesizes that processing time can be predicted from task features; the Coder writes code that trains a regressor on the logs, yielding a callable surrogate predict() . Bottom: characterizing the energy profile of each processor yields a semantic insight (per-processor average power).
Figure 4: Cost J (log scale, lower is better) over 10 independent runs for Scenario 1 ( N=10 , left) and Scenario 2 ( N=50 , right).
Strategy
Scenario 1 ( N=10 )
Scenario 2 ( N=50 )
Loss
Delay (ms)
Cost ( J )
Loss
Delay (ms)
Cost ( J )
Random
0.70
32.15
103.71
0.10
65.22
117.91
Oracle-Time
0.03
1.30
4.92
0.07
4.81
46.74
Oracle-Queue
0.62
15.76
90.23
0.00
17.32
17.32
FunSearch
0.05
2.43
8.50
0.08
49.39
90.98
EoH
0.05
2.43
8.50
0.01
11.51
15.88
Table 1: Cloud-computing use case: loss rate, mean delay (ms), and cost J (lower is better) of the fixed baselines and of the best of 10 runs of each AHD method, over low, medium, and high loads.
Strategy
Loss
Delay (ms)
Energy ( μ J)
Cost
Random
0.10
0.30
670.01
1.58
Oracle-Time
0.00
0.35
884.22
0.89
Oracle-Queue
0.08
0.27
700.76
1.47
CPU-Only
0.27
0.59
170.46
2.79
HA-Only
0.00
0.21
1132.39
1.13
EoH
0.00
0.31
533.33
0.58
Table 2: 5G vRAN use case: TB loss rate, mean processing delay, mean energy per TB, and cost (equation 2 ; lower is better) for fixed policies (top), industry standards (middle), and the best of 10 runs of each AHD method (bottom).
Strategy
Semantic Grounding
Auditability
Simulatability
Mean
FunSearch
1.00
1.47
2.00
1.49
EoH
1.00
1.33
4.47
2.27
ReEvo
1.00
2.00
5.00
2.67
HSEvo
1.27
1.33
2.27
1.62
TRACE
3.33
4.67
1.00
3.00
Table 3: LLM-as-a-Judge auditability scores for the best heuristics in Scenario 1, averaged over a three-judge panel and five queries per judge. Per-judge scores and rubric definitions are in Appendix F .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Hardware / Service
Tokens (in)
Tokens (out)
Time (s)
Reasoner (Hypothesis gen.)
Google API
454–580
21–25
14–31
Coder (Artifact generation)
Local GPU A100
980–995
1050–1350
7.3–7.8
Sandbox execution
Local CPU Intel i7
–
–
0.1–3.85
Heuristic generation
Local GPU A100
479–495
262–738
1.7–3.5
Heuristic evaluation
Local CPU Intel i7
–
–
0.82–1.3
Appendix
Table 4: Per-component computational profile. Token consumption and execution times are reported as the 5th and 95th percentiles over 10 independent runs.
Block
Component
Avg. call (s)
# calls
Total (s)
% total
Reasoner (Hypothesis gen.)
22.5
1
22.5
4.86
Agentic
Coder (Artifact generation)
7.41
1
7.41
1.60
Workflow
Sandbox execution
2.24
1
2.24
0.48
Sub-total
32.15
6.95
Standard AHD
Heuristic generation
1.92
150
288.0
62.25
steps
Heuristic evaluation
0.95
150
142.5
30.80
Appendix
Table 5: Average wall-clock breakdown for a single TRACE epoch in which the agentic workflow is executed ( I=150 generation steps), over 10 independent runs. TRACE’s Agentic Workflow incurs only 6.95% of the total epoch duration; the remaining 93.05% is attributed to the standard AHD steps shared with baselines such as EoH.
Figure 5: LLM choice for TRACE’s Coder/Evolution loop: distribution of the best per-run cost across 10 independent runs on Scenario 1.
Figure 6: EoH cost distribution across 10 independent runs on Scenario 1 under four LLM backbones, including gemini-3-pro-preview (the model used by TRACE’s Reasoner). Even under the strongest backbone, EoH remains 1 – 2 orders of magnitude above TRACE (cf. Fig. 4 , left, where TRACE’s median cost is below 1 ).
Strategy
Semantic Grounding
Auditability
Simulatability
Mean
Gemini
Claude
GPT
Gemini
Claude
GPT
Gemini
Claude
GPT
FunSearch
1.00
1.00
1.00
1.40
1.00
2.00
1.00
4.00
1.00
1.49
EoH
1.00
1.00
1.00
1.00
1.00
2.00
3.40
5.00
5.00
2.27
ReEvo
1.00
1.00
1.00
2.00
2.00
2.00
5.00
5.00
5.00
2.67
HSEvo
1.00
1.00
1.80
1.00
1.00
2.00
1.00
3.80
2.00
1.62
TRACE
4.00
3.00
3.00
5.00
4.00
5.00
1.00
1.00
1.00
3.00
Appendix
Table 6: Multi-judge auditability scores for the best-found heuristics in Scenario 1. Each cell reports the mean over five queries at temperature 0.7 ; the Mean column averages the nine (rubric, judge) means per method.
Figure 7: Latency-energy trade-off in heterogeneous 5G edge clouds.
Systems resource management tasks rely primarily on hand-designed heuristics. However, growing hardware heterogeneity and workload diversity require heuristics specialized to particular deployment instances, making manual design expensive and difficult to scale. In this paper, we explore how to synthesize systems heuristics using LLMs. The main challenge is ensuring that generated heuristics execute safely, integrate correctly with the surrounding system, and still achieve strong performance. We propose Vulcan, a framework that identifies LLM-friendly interfaces that isolate core decision logic from the rest of the implementation. With Vulcan, LLM-generated code is restricted to simple stateless decision functions, while trusted runtime abstractions provide rich derived statistics for meaningful policy exploration without system-integration bugs. To ensure execution safety, LLMs synthesize heuristics in a restricted language, Anvil, that guarantees important properties by construction. We evaluate Vulcan across three well-studied domains and demonstrate up to 4.9× higher savings for spot-VM scheduling, up to 2× lower miss ratios for cache eviction, and up to 14% higher application performance for tiered-memory systems, while ensuring execution safety throughout.
Automated Heuristic Design (AHD) with Large Language Models (LLMs) has shown remarkable progress in discovering high-quality heuristics. However, existing LLM-based AHD methods optimize heuristics for a fixed training instance set and may fail catastrophically when deployed under real-world distributional shifts. We propose Robust Adversary Instance Search (RAISE), a framework that integrates constrained worst-case instance search within a principled neighborhood of the training distribution into the LLM-based evolutionary search loop. RAISE treats robust AHD as a constrained adversarial instance search problem: the outer loop evolves heuristics via LLM operators, while an LLM-free inner loop efficiently identifies hard instances within an epsilon-ball around the training instance set using a basis distribution parameterization with boundary projection. Comprehensive experiments on Online Bin Packing (OBP), Online Job Shop Scheduling (OJSP), and Online Vehicle Routing (OVRP) across five distribution families demonstrate that existing LLM-based AHD methods degrade by up to 19 times under distribution shift, while RAISE consistently maintains strong performance across all tested distributions and problem scales
Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most methods treat the simulator as a black box: aggregate scores rank candidates without revealing why they fail or which policy logic should change. We present an LLM-guided heuristic design framework that uses repeated simulation for selection and event-level traces for diagnosis. Each incumbent is assessed through multiple replications, while replaying its lowest-scoring one produces a queryable trace. A manager agent formulates bottleneck hypotheses from this evidence, and editing agents implement parallel code-level revisions. After execution checks and repeated evaluation, best-so-far selection retains only improvements. LLM revision occurs between evaluation batches, while a fixed policy controls each simulation run. We evaluate the framework in a discrete-event simulation of dynamic production and automated guided vehicle (AGV) scheduling. Across five independent optimization runs with Gemini-3.1-Pro, final mean scores averaged 77.51 on the simulator's 0-100 scale. In the highest-scoring run, trace-based diagnoses motivated proactive charging, distance-aware AGV assignment, and rebalanced dispatch priorities, raising the best-so-far mean score from 62.49 to 78.61. On 100 matched seeds, the best final policy outscored representative rolling-MILP, rule-based, and metaheuristic policies on every seed and retained its advantage under random faults without re-optimization. After separate re-optimization for a longer horizon and variable order interarrival times, the resulting policies again outscored all baselines. Ablations with two LLM backbones showed that removing either parallel candidate generation or trace-database access reduced final mean scores. These results show that simulation traces can guide targeted code-level policy improvement in complex simulation-based scheduling.
Jinbo Li, Chuanhao Li
Department of Industrial Engineering, Tsinghua University