OptiCom : A Unified Framework for State-Conditioned Composition in LLM-Driven Optimization
Authors: Chenxing Wei, Sichen Liu, Lizhao Liu, Ningyuan Sun, Chen Bingzhou, Ying He, Bo Jiang, Fei Yu, +1 more
Organizations: School of Computing and Data Science, The University of Hong Kong · Shenzhen Loop Area Institute · ByteDance · Huazhong University of Science and Technology · School of Information Technology, Carleton University · Hong Kong University of Science and Technology (Guangzhou)
Large language models (LLMs) are increasingly deployed to solve complex scientific and practical problems via iterative optimization. However, dynamically coordinating diverse search mechanisms as candidate quality, failure modes, and resource budgets evolve remains a critical open challenge. Targeted empirical diagnostics reveal that mechanism effectiveness is highly state-dependent. Motivated by this, we analyze how individual decisions drive final outcomes, decomposing the expected terminal improvement under a shared budget into cumulative decision opportunities minus cumulative selection losses. Guided by this opportunity-loss theoretical foundation, we propose OptiCom, a unified framework that represents LLM-driven optimizers within a shared configuration space: C=(A,Q,O,E,M,S), corresponding to artifact, query, operator, evaluation, memory, and strategy. Operating within this space, a fast LLM-based Optimization Controller dynamically composes immediate mechanisms through structured Action Packages, while a slower Strategy Adapter refines long-term selection preferences, operator weights, and templates based on accumulated trajectory feedback. Comprehensive evaluations across 32 benchmark groups demonstrate the superiority of framework: OptiCom achieves an average Max-score rank of 1.72 among 14 evaluated configurations, securing the top score in 23 groups. Ultimately, these results highlight the broad applicability and high extensibility of OptiCom as a general-purpose paradigm for robust LLM test-time scaling.
Figures & tables
Feedback-guided revision and experience reuse
Evolutionary search and adaptive control
Shared abstractions and meta-optimization
TextGrad Yuksekgonul et al. (2025)
AdaEvolve Cemri et al. (2026)
DSPy Khattab et al. (2024)
A
Optimizable graph variables
Candidate programs
LM program instructions and demonstrations
Q
Graph context and propagated feedback
Selected programs and search evidence
Training examples and execution traces
O
Textual-feedback-guided updates
LLM-generated program variations
Instruction and demonstration updates
E
Objective feedback
Program fitness
User-specified task metric
M
Graph and update state
Populations and improvement statistics
Optimizer-dependent candidates and search state
Table 1: Representative AQOEMS mappings across three research directions. Each method has a framework instantiation; mapping and alignment details appear in Appendix A .
Figure 1: State-conditioned diagnostics on Heilbronn Triangle: three distinct situations per state and ten trials per mechanism–situation pair (30 trials per mechanism–state rate). (a) Averages over Q/O/E/M/S alternatives; (b) query and (c) operator comparisons. R marks the reference setting. Memory comparisons enable retrieval; other dimensions change one mechanism from the reference. Appendix B details the protocol and complete six-panel results.
Figure 2: Overview of OptiCom. The Optimization Controller selects and composes Q/O/E/M mechanisms through an Action Package, while the slower Strategy Adapter updates the guidance used by subsequent Controller decisions. Solid arrows show execution and state flow; gray dashed arrows indicate configuration, and purple dashed arrows indicate slower strategy adaptation. Annotations involving Δt and ℓt highlight how specific architectural components are explicitly designed to expand search opportunities ( Δt ) and reduce selection losses ( ℓt ).
Domain
Math
Systems
GPU
Algorithms
Reasoning
Prompts
Quantum
Benchmark
Heilbronn
LLM-SQL
TriMul
Frontier-CS
ARC
HotpotQA
QNN
Method
Max
Mean
Max
Mean
Max
Mean
Max
Mean
Max
Mean
Max
Mean
Max
Mean
ProTeGi
0.828
0.813
0.630
0.622
2.541
2.531
0.632
0.619
0.473
0.461
0.441
0.427
0.8833
0.8367
LATS
0.887
0.867
0.686
0.682
2.576
2.542
0.627
0.609
0.457
0.432
0.413
0.409
0.8500
0.8300
EvoX
0.943
0.931
0.684
0.677
2.593
2.583
0.626
0.617
0.461
0.443
0.386
0.357
0.8167
0.8067
EvoX (official)
0.956
0.941
0.696
0.681
2.592
2.579
0.613
0.597
0.473
0.451
0.397
0.363
0.8333
0.8133
Table 2: Main results across seven representative benchmarks over five runs. Unqualified method names denote method-inspired framework profiles; “official” denotes the official SkyDiscover implementation. Scores follow task-specific scales. Higher scores are better. Bold and underlined values indicate the best and second-best distinct results.
Figure 3: (a) Average within-group ranks of OptiCom and 13 method-inspired profiles across 32 benchmark groups, computed separately from five-run Max and Mean scores (lower is better). (b) Archived best-so-far scores over optimization iterations on Heilbronn Triangle.
Config.
Heilbronn
HotpotQA
Full OptiCom
0.953±0.008
0.501±0.017
Fixed composition
0.704±0.034
0.409±0.127
Random composition
0.672±0.171
0.381±0.153
w/o Strat. Adapter
0.713±0.027
0.451±0.058
Operator-only adapt.
0.697±0.056
0.428±0.076
w/o Exp. Memory
0.893±0.024
0.483±0.039
Table 3: Core ablations with Doubao-Seed-2.0-pro. Scores report mean ± std over 5 runs.
Figure 4: Heilbronn Triangle case study: best-so-far combined score (blue) and recorded cumulative API tokens (orange). The run reaches 0.9608 at iteration 13 using ≈ 293k tokens, then spends another 572k through iteration 30 without improvement. Selected execution events are annotated; their causal effects and the impossibility of further improvement are not established.
Heilbronn Triangle
HotpotQA
Backbone
w/o Adapter
Full OptiCom
Gain
w/o Adapter
Full OptiCom
Gain
Doubao-Seed-2.0-pro
0.713±0.027
0.953±0.008
+0.240
0.451±0.058
0.501±0.017
+0.050
GPT-5.5
0.796±0.011
0.987±0.002
+0.191
0.471±0.037
0.553±0.012
+0.082
Claude Opus 4.6
0.787±0.012
0.973±0.007
+0.186
0.469±0.041
0.554±0.010
+0.085
GLM-5.3
0.783±0.014
0.969±0.008
+0.186
0.461±0.059
0.523±0.019
+0.062
Kimi-K3
0.791±0.011
0.974±0.004
+0.183
0.465±0.053
0.547±0.016
+0.082
Table 4: Cross-backbone evaluation: mean ± SD over five runs (at most 30 iterations). Gain denotes improvement from enabling the Adapter. Bold and underlining mark the best and second-best distinct values per column, separately for means ( ↑ ), SDs ( ↓ ), and gains ( ↑ ).
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Dimension
Optimization concept
Functional responsibility
Illustrative realizations
A : Artifact
Solution representation
What is optimized, and how candidates are represented.
Source code, numerical parameters, mathematical constructions, prompts, or images.
Q : Query
Information acquisition
What evidence and context are acquired or assembled to inform an update.
Inspect selected failures, request profiling evidence, retrieve past attempts, or use only the current candidate and feedback.
O : Operator
Search and update operations
How candidates are generated or modified.
Generate a new candidate, repair an error, apply a local revision, or recombine existing candidates.
E : Evaluation
Objective and constraint assessment
How candidate quality and feasibility are assessed, and what feedback is returned.
Correctness checks, runtime measurements, diagnostic tests, or higher-fidelity verification.
M : Memory
Search history and experience
What information persists across iterations, and how it is maintained.
Retain an incumbent, a candidate archive, evaluation records, or reusable lessons; omit persistent experience storage when unused.
S : Strategy
Search control
How the other functions are coordinated and resources are allocated as optimization proceeds.
Follow a prescribed schedule, apply state-dependent rules, or adapt selection preferences using accumulated outcomes.
Appendix
Table 5: Functional dimensions of AQOEMS. The dimensions describe responsibilities rather than mandatory modules or sequential execution steps. The examples are possible realizations, not required components; their use may depend on the optimization state through S .
Method / status
Artifact, query, and operator
Evaluation, memory, and strategy
OPRO Yang et al. (2024) B
A: Candidate solutions, including prompts. Q: Task description and selected scored solutions. O: Generate proposals conditioned on that history.
E: Task objective values. M: Evaluated candidate–score records. S: Repeated generation and evaluation with history selection and stopping rules.
TextGrad Yuksekgonul et al. (2025) B
A: Optimizable variables in a computational graph. Q: Relevant graph context and propagated feedback. O: Variable updates guided by textual feedback.
E: Objective evaluation and associated criticism. M: Graph state, current variables, and optimizer-dependent update history. S: Feedback propagation and variable-update procedure.
ProTeGi Pryzant et al. (2023) B
A: Task prompts. Q: Minibatch examples and observed errors. O: Generate textual critiques and corresponding prompt revisions.
E: Predictive performance on evaluation examples. M: Beam candidates and evaluation statistics. S: Beam expansion and bandit-based candidate selection.
Reflexion Shinn et al. (2023) B
A: Task attempts, including programs or action trajectories. Q: Current feedback and previous verbal reflections. O: Produce a subsequent attempt informed by reflection.
E: Task-dependent external or internal feedback. M: Episodic verbal reflections. S: Attempt, evaluate, reflect, and retry.
GEPA Agrawal et al. (2025) B
A: Prompts in an AI system. Q: Execution trajectories, diagnostics, and selected candidate context. O: Reflective mutation and combination of complementary candidates.
E: Task evaluations, including performance across examples. M: Candidate pool, scores, and ancestry. S: Pareto-aware selection and evolutionary updates under an evaluation budget.
FunSearch Romera-Paredes et al. (2023) B
A: Functions within a program specification. Q: Selected prior functions assembled into a prompt. O: LLM-generated function variants.
E: Automated execution and scoring. M: An island-structured database of evaluated programs. S: Program sampling, insertion, and island reset rules.
Appendix
Table 6: Complete functional mappings. B denotes a corresponding baseline profile in the supplied registry; C denotes conceptual coverage only. B does not certify equivalence to the original implementation. The two right columns jointly specify all six AQOEMS dimensions.
Aspect
Mechanism to preserve
Alignment check
Candidate generation
The information and update rule used to produce new candidates.
Check that generation receives the intended parents, examples, feedback, or revision instructions.
Selection and search
The method’s candidate selection, branching, population, or tree-search rule.
Check that these rules affect executed actions, rather than appearing only in configuration descriptions.
Feedback use
The feedback representation and its role in subsequent optimization.
Check that evaluation records, reflections, or textual critiques reach the decisions they are intended to guide.
Memory and reuse
The information retained across iterations and the rule for retrieving it.
Check that retained candidates or experiences remain available and are retrieved according to the configured mechanism.
Strategy adaptation
The trigger, evidence, and scope of changes to the optimization procedure.
Check that strategy updates affect later decisions; distinguish adaptive rules within a configuration from changes to the configuration itself.
Evaluation and budget
Task validity requirements, scoring semantics, and resource constraints.
Check that candidates use the shared benchmark evaluator and that reported limits include the relevant optimization operations.
Appendix
Table 7: Criteria for checking framework instantiations. These checks distinguish preservation of a method’s defining mechanism from changes introduced by shared interfaces. The table specifies verification criteria rather than certifying equivalence to official implementations.
State
Initial condition
Success criterion
s1
Execution failure or hard-constraint violation
Recover a candidate that executes and satisfies the required hard constraints.
s2
Failed correctness tests
Preserve feasibility and pass all required correctness checks.
s3
Objective stagnation
Preserve feasibility and correctness and exceed the predefined objective-improvement criterion.
s4
Limited remaining budget
Reach the s3 objective criterion within the declared remaining budget.
s5
High performance variability
Reduce variability beyond the predefined criterion without exceeding the allowed degradation in mean quality, feasibility, or correctness.
Appendix
Table 8: Optimization states and state-resolution criteria in the Heilbronn diagnostic study.
ID
Situation
Starting condition
s1-1
Runtime exception
Point generation or updating raises an exception, such as an out-of-range index or incompatible array dimensions, and terminates without returning a point set.
s1-2
Construction timeout
Excessive search, too many internal iterations, or an ineffective termination condition prevents the construction from finishing within the declared execution limit.
s1-3
Invalid output interface
The program terminates but returns the wrong number of points, an array other than shape (11,2) , or coordinates containing NaN or infinity.
s2-1
Incorrect boundary handling
The program returns a well-formed point set, but rectangular clipping or an incorrect projection leaves points outside the sloping boundaries of the equilateral triangle.
s2-2
Incorrect area objective
The search uses an incorrect internal area calculation, such as omitting the absolute determinant or using an incorrect normalization factor, so its internal objective disagrees with the specified geometric objective.
s2-3
Incomplete triple enumeration
The internal objective evaluates only a subset of triples, such as consecutive or locally selected points, and can miss the triple determining the true minimum area.
Appendix
Table 9: Starting situations for the expanded Heilbronn Triangle diagnostic study. Each state contains three situations with ten trials per mechanism and situation.
Figure 5: Complete mechanism-level view of the Heilbronn Triangle motivation study. Each displayed mechanism–state value averages the three situations in that state, with ten trials per situation. (a) Dimension-level averages. (b)–(f) Query, operator, evaluation, memory, and strategy mechanisms. R denotes the reference setting. The memory panel uses the gate-open comparison described in this appendix so that stored episodic, structured, and adaptive memory can be read by the memory-lookup query mechanism.
ID
Mechanism
s1-1
s1-2
s1-3
s2-1
s2-2
s2-3
s3-1
s3-2
s3-3
s4-1
s4-2
s4-3
s5-1
s5-2
s5-3
Q1
none
3
3
4
3
3
3
3
3
3
4
3
4
2
2
2
Q2
self-only (R)
5
4
6
5
5
4
4
3
4
4
3
3
3
3
3
Q3
env-probe
9
7
8
8
8
7
4
4
4
2
3
3
3
3
3
Q4
memory-lookup
3
2
4
3
3
3
3
3
3
4
3
4
2
2
2
O1
local-revision (R)
5
4
6
5
5
4
4
3
4
4
3
3
3
3
3
O2
repair
8
6
8
8
8
7
2
2
2
5
4
6
2
2
2
Appendix
Table 10: Successful trials out of ten for all 15 situations in the strict base intervention matrix. R marks the reference mechanism. Memory rows M2 and M3 are inert here because the reference query does not read persistent memory; the gate-open memory comparison is reported separately in Table 11 .
ID
Mechanism
s1-1
s1-2
s1-3
s2-1
s2-2
s2-3
s3-1
s3-2
s3-3
s4-1
s4-2
s4-3
s5-1
s5-2
s5-3
M2
episodic + Q4
6
5
7
6
6
5
5
3
5
5
4
4
3
3
4
M3
structured + Q4
6
5
7
7
6
6
6
4
6
7
7
7
5
5
6
M4
adaptive + Q4
7
6
8
7
6
7
6
5
7
7
7
7
5
5
7
Appendix
Table 11: Gate-open memory comparison. Episodic, structured, and adaptive memory are paired with Q4 memory lookup so that stored experience can affect later decisions. Entries are successful trials out of ten.
Figure 6: Separate and joint query/operator changes across the three s2 situations. (a) Success rate over 30 trials per configuration; higher is better. (b) Mean iterations to first success among successful trials only; lower is better. The reference, Q-only, O-only, and Q+O settings obtain 46.7%, 76.7%, 76.7%, and 83.3% success, respectively, with conditional mean iterations of 5.2, 2.9, 3.6, and 2.1.
Configuration
Query
Operator
Successes
Avg. iterations
R
self-only
local-revision
14/30
5.2
Q-only
env-probe
local-revision
23/30
2.9
O-only
self-only
repair
23/30
3.6
Q+O
env-probe
repair
25/30
2.1
Appendix
Table 12: Joint Q/O results across the three s2 situations. Average iterations are conditional on success.
Field
Decision
Consumer and constraints
qt
Query plan
Query executor: permitted tools, evidence requests, and retrieval operations; may be empty.
ct
Context scope
Context builder: candidate identifiers, diagnostics, and experience to expose within the context limit.
ot
Update operator
Operator executor: a registered mechanism with compatible artifact inputs and required parent candidates.
wt
Branch width
Scheduler: a positive bounded number of candidate proposals, subject to available resources.
et
Evaluation plan
Evaluator: registered stages, effort limits, and feedback form; final required checks remain unchanged.
ρt
Retention policy
Archive and experience stores: which evaluated candidates, records, and lessons to retain or expose.
Appendix
Table 13: Action Package fields and their execution responsibilities. Fields select registered functionality; they do not redefine the task contract.
Domain
Benchmark family
Evaluation target
Math
Mathematical optimization
Numerical objectives subject to task-specific feasibility constraints.
Systems
ADRS
Workload performance, cost, and composite system objectives.
GPU
GPU Mode; KernelBench
Numerical correctness and kernel execution performance; scaled inverse runtime or eager-baseline speedup, respectively.
Algorithms
Frontier-CS; ALE-Bench-Lite
Bounded algorithmic-problem scores or private contest-performance scores over the stated problem sets.
Reasoning
ARC
Correctness of candidate transformations on benchmark inputs.
Creative
Sky Festival
Satisfaction of semantic and compositional image requirements.
Appendix
Table 14: Benchmark domains and evaluation targets. Evaluation uses task-specific benchmark harnesses and the stated aggregation rules.
Figure 7: Archived best-so-far scores on (a) Erdos Min Overlap and (b) Second Autocorr Ineq. OptiCom establishes an early advantage on Erdos Min Overlap, whereas EvoX obtains a higher final score on Second Autocorr Ineq. Highlighted annotations mark the final recorded best score of OptiCom and its corresponding iteration.
Figure 8: Archived best-so-far scores on (a) Circle Packing Rect and (b) First Autocorr Ineq. Insets reveal differences among high-scoring candidates that are difficult to distinguish on the full vertical scale. OptiCom achieves the highest final score on Circle Packing Rect, while EvoX finishes slightly higher on First Autocorr Ineq.
Figure 9: Archived best-so-far scores on (a) Heilbronn Triangle and (b) Heilbronn Convex 13. On Heilbronn Triangle, OptiCom resumes improvement after an initial plateau and reaches 0.961 at iteration 13 . On Heilbronn Convex 13, successive improvements produce a final recorded score of 0.939 at iteration 14 .
Figure 10: Archived best-so-far scores on (a) Minimizing Max Min Dist 2 and (b) Minimizing Max Min Dist 3. Large early improvements are followed by smaller refinements, with the final recorded improvements of OptiCom occurring at iterations 26 and 13 , respectively. The displayed score of 1.000 is rounded and does not constitute an optimality certificate.
Figure 11: Archived best-so-far scores on (a) Circle Packing and (b) Sums Diffs Finite Sets. OptiCom makes a large early improvement on Circle Packing, but AlphaEvolve later achieves a higher score. On Sums Diffs Finite Sets, smaller improvements accumulate until iteration 20 . The Circle Packing panel marks the shorter recorded budget of LATS.
Figure 12: Archived best-so-far scores on (a) Matmul and (b) Hexagon Packing 12. On Matmul, OptiCom overcomes an initial plateau and reaches 0.842 at iteration 11 , matching the strongest final score shown. On Hexagon Packing 12, improvements to 0.764 remain insufficient to match several baselines, illustrating a limitation of the observed optimization trajectory.
Heilbronn Triangle
HotpotQA
Backbone
Full OptiCom
w/o Adapter
Full OptiCom
w/o Adapter
Doubao-Seed-2.0-pro
0.953±0.008
0.713±0.027
0.501±0.017
0.451±0.058
GPT-5.5
0.987±0.002
0.796±0.011
0.553±0.012
0.471±0.037
Claude Opus 4.6
0.973±0.007
0.787±0.012
0.554±0.010
0.469±0.041
GLM-5.3
0.969±0.008
0.783±0.014
0.523±0.019
0.461±0.059
Kimi-K3
0.974±0.004
0.791±0.011
0.547±0.016
0.465±0.053
Appendix
Table 16: Cross-backbone evaluation over five runs per configuration, with a maximum of 30 iterations. All optimization-related LLM calls use the listed backbone. Bold indicates the higher mean within each backbone–task pair.
Backbone
Heilbronn Triangle
HotpotQA
Doubao-Seed-2.0-pro
+0.240
+0.050
GPT-5.5
+0.191
+0.082
Claude Opus 4.6
+0.186
+0.085
GLM-5.3
+0.186
+0.062
Kimi-K3
+0.183
+0.082
Appendix
Table 17: Absolute mean-score gains from enabling the Strategy Adapter, computed as Full OptiCom minus w/o Adapter using Table 4 . These are within-backbone comparisons under the same 30-iteration cap.
Configuration
Final score
Adapter calls
Total tokens
No adaptation
0.713±0.027
0
853K
Event-triggered adaptation
0.953±0.008
10
865K
Periodic adaptation ( k=5 )
0.939±0.010
5
861K
Every-iteration adaptation
0.954±0.003
29
877K
Appendix
Table 18: Strategy adaptation frequency on Heilbronn Triangle. Scores report mean ± standard deviation over five runs. Adapter calls and total API tokens are run averages; K denotes one thousand tokens. Updates are not invoked after the final iteration.
Final score
Total tokens
Feedback
Full OptiCom
w/o Adapter
Full OptiCom
w/o Adapter
Basic
0.897±0.057
0.679±0.126
809K
799K
Rich
0.953±0.008
0.713±0.027
865K
853K
Appendix
Table 19: Feedback access on Heilbronn Triangle over five runs. Both conditions use the same evaluator and final scoring rule. Scores are mean ± standard deviation; token counts are mean total API consumption per run.
Initial configuration
Fixed
Adaptive OptiCom
Local-revision-first
0.813±0.017
0.954±0.008
Diagnosis-and-repair-first
0.826±0.012
0.956±0.008
Exploration-and-recombination-first
0.857±0.011
0.951±0.008
Appendix
Table 20: Initialization sensitivity on Heilbronn Triangle. Each entry reports the mean and standard deviation over five runs with a maximum of 30 iterations. Adaptive OptiCom begins with the corresponding initial composition and can adjust it after the first iteration.
Recent work has demonstrated the promise of orchestrating large language models (LLMs) within evolutionary and agentic optimization systems. However, the mechanisms driving these optimization gains remain poorly understood. In this work, we present a large-scale study of LLM-guided evolutionary search, collecting optimization trajectories for 15 LLMs across 8 tasks. Although zero-shot problem-solving ability correlates with final optimization outcomes, it explains only part of the variance: models with similar initial capability often induce dramatically different search trajectories and outcomes. By analyzing these trajectories, we find that strong LLM optimizers behave as local refiners, producing frequent incremental improvements while progressively localizing the search in semantic space. Conversely, weaker optimizers exhibit large semantic drift, with sporadic breakthroughs followed by stagnation. Notably, various measures of solution novelty do not predict final performance; novelty is beneficial only when the search remains sufficiently localized around high-performing regions of the solution space. Our results highlight the importance of trajectory analysis for understanding and improving LLM-based optimization systems and provide actionable insights for their design and training.
Xinhao Zhang, Xi Chen, François Portet +1
Univ. Grenoble Alpes, CNRS, Grenoble INP, LIG, 38000 Grenoble, France
Agentic AI systems built on large language models (LLMs) offer significant potential for automating complex workflows, from software development to customer support. However, LLM agents often underperform due to suboptimal configurations; poorly tuned prompts, tool descriptions, and parameters that typically require weeks of manual refinement. Existing optimization methods either are too complex for general use or treat components in isolation, missing critical interdependencies. We present ARTEMIS, a no-code evolutionary optimization platform that jointly optimizes agent configurations through semantically-aware genetic operators. Given only a benchmark script and natural language goals, ARTEMIS automatically discovers configurable components, extracts performance signals from execution logs, and evolves configurations without requiring architectural modifications. We evaluate ARTEMIS on four representative agent systems: the \emph{ALE Agent} for competitive programming on AtCoder Heuristic Contest, achieving a \textbf{13.6% improvement} in acceptance rate; the \emph{Mini-SWE Agent} for code optimization on SWE-Perf, with a statistically significant \textbf{10.1% performance gain}; and the \emph{CrewAI Agent} for cost and mathematical reasoning on Math Odyssey, achieving a statistically significant \textbf{36.9% reduction} in the number of tokens required for evaluation. We also evaluate the \emph{MathTales-Teacher Agent} powered by a smaller open-source model (Qwen2.5-7B) on GSM8K primary-level mathematics problems, achieving a \textbf{22% accuracy improvement} and demonstrating that ARTEMIS can optimize agents based on both commercial and local models.
Paul Brookes, Vardan Voskanyan, Rafail Giavrimis +18
TurinTech AI, London, UK · University of Surrey, Guildford, United Kingdom · University of Leeds, Leeds, UK +4
The high cost and data scarcity in scientific exploration have motivated the use of large language models (LLMs) as knowledge-driven components in Bayesian optimization (BO). However, existing approaches typically embed LLMs directly into the sampling or surrogate modeling pipeline, without fully leveraging their significantly lower evaluation cost compared to real-world experiments. To address this limitation, we propose LLM-Accelerated Bayesian Optimization (LABO), a framework that combines LLM predictions with experimental observations within a single BO loop. LABO employs a gating criterion to dynamically balance the reliance on LLM predictions versus actual experiments. By leveraging inexpensive LLM evaluations to broadly explore the search space and reserving costly real experiments only for regions with high uncertainty, LABO achieves more sample-efficient optimization. We provide a theoretical analysis with a cumulative regret bound that formalizes this efficiency gain. Empirical results across diverse scientific tasks demonstrate that LABO consistently outperforms existing methods under identical experimental budgets. Our results suggest that LABO offers a practical and theoretically grounded approach for integrating LLMs into scientific discovery workflows.
Zhuo Chen, Xinzhe Yuan, Jianshu Zhang +8
equal contribution · Shanghai Artificial Intelligence Laboratory, Shanghai, China · School of Mechanical Engineering, Shanghai Jiao Tong University, Shanghai, China +7