Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making, showing great potential due to the integration with mathematically rigorous tools. However, existing agentic BBO studies use different task domains and system configurations, making their results difficult to compare and the effects of individual design choices hard to isolate. We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol. In our experiments, agentic BBO achieves higher family-averaged scores than direct LLM-based methods in all five domains and outperforms the best numerical optimizers in four. We further study three factors shaping agent performance: optimization tools, task information and prior knowledge, and the role of the LLM during search. Our results show that additional numerical tools do not consistently improve performance, task semantics are broadly useful while more specific priors are less reliable, and numerical optimizers can effectively absorb gains from search trajectories established by the agent. Finally, we introduce a five-task frontier challenge within AgenticBBO-Bench and evaluate seven LLMs under the Codex agent harness, where GPT-6 Astra and DeepSeek-V4.1-Flash lie on the Pareto frontier of performance and cost among the evaluated models. Our code is available at https://github.com/lamda-bbo/agentic-bbo.
Figures & tables
Figure 1: Frontier-model comparison on AgenticBBO-Bench. Left: overall benchmark scores. Right: score versus official list-price token cost per task (log scale). Models with similar optimization performance can incur substantially different inference costs.
Figure 2: Comparison between traditional Bayesian optimization and Agentic BBO. Traditional BO follows a fixed optimizer-driven procedure, whereas Agentic BBO lets an LLM agent organize the intermediate decision process before each expensive evaluation. The actions shown in the agent loop are representative examples rather than a fixed or exhaustive workflow: the agent may dynamically choose, skip, repeat, combine, or introduce other actions before proposing a candidate.
Family
Search Space
Semantics
Feasibility Constraints
Evaluator
Synthetic
Continuous
–
–
Analytic function
HPO
Mixed
✓
–
Model training
Database Tuning
Mixed
✓
–
Learned surrogate
Chip Design
Continuous
✓
✓
Placement evaluator
Molecular Design
Structured discrete
✓
✓
Molecular scoring
Table 1: Overview of the five benchmark task families. Semantics indicates whether task or variable meanings are exposed to the optimizer, while Feasibility Constraints indicates whether the search space contains candidates that may be infeasible or invalid. Evaluator describes how objective values are obtained in each task family.
Table 2: Results of the broad benchmark and numerical-tool study. Scores are normalized such that higher values indicate better optimization performance, and the best result in each comparison is shown in bold . A dash indicates that the method is not applicable to that task family.
Figure 3: Tool use in Agentic BBO. (a) Tool-use patterns across the diagnostic tasks. (b) A representative Bigblue1 trajectory in which the agent fits a proxy model and uses it to refine the placement.
Figure 4: Effect of task information and prior knowledge. (a) Performance on six real-world tasks with anonymous inputs, task semantics, or additional domain priors. (b) Performance on controlled objectives under different forms of prior information.
Table 3: Results of the LLM-role study and the five-task frontier challenge. Higher scores indicate better optimization performance, and the best result in each comparison is shown in bold .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Experimental panel
Tasks
Initial + new evaluations
Broad BBOB
24
20+100
Broad HPO
25
5+25
Broad DBTune
6
50+200
Broad BBOPlace
12
50+200
Broad GuacaMol
10
50+200
Tool / policy diagnostics: BBOB
2
20+50
Appendix
Table 4: Evaluation budgets used across the experimental panels. “Initial + new” denotes the number of shared initial observations followed by the number of additional objective evaluations available to each method.
Session
v0
v1
v2
v3
v4
v5
v6
Pick
Held-out S
BBOB 1
0.472
0.654
0.660
0.651
0.636
0.634
0.612
v5
0.548
BBOB 2
0.472
0.470
0.417
0.442
0.488
0.460
0.393
v5
0.484
BBOB 3
0.472
0.670
0.683
0.759
0.735
0.741
0.700
v4
0.630
HPO 1
0.507
0.574
0.577
0.607
0.594
0.594
0.594
v4
0.532
HPO 2
0.507
0.285
0.603
0.625
0.589
0.622
0.615
v5
0.531
HPO 3
0.507
0.577
0.492
0.455
0.290
0.597
0.607
v6
0.279
Appendix
Table 5: Development trajectories and held-out performance. Development entries average four tasks and two seeds, while held-out S averages two disjoint tasks and four seeds. Bold marks the historically selected version.
Session
Pick
Implemented search logic
BBOB 1
v5
GP-LogEI with an adaptive incumbent-centered candidate region.
BBOB 2
v5
UCB for the first 12 new evaluations, followed by EI on the same GP.
BBOB 3
v4
Matérn-5/2 GP with LogEI, conditional inverse-hyperbolic-sine response compression, and a trust region.
HPO 1
v4
Warped-coordinate GP with global EI candidate pools and dimension-dependent greedy local search.
HPO 2
v5
Warped-coordinate GP-EI with an adaptive trust region and separation from observed candidates.
HPO 3
v6
Reference GP-EI augmented with scheduled Halton coverage probes and a structured-pool duplicate fallback.
Appendix
Table 6: Search logic implemented by the six selected programs.
Bayesian optimization (BO) has become the standard tool for sample-efficient optimization and owes its efficiency to uncertainty-aware search driven by generic statistical priors. Richer domain priors can improve BO in principle, but encoding them through tailored kernels or problem structure is difficult and rarely done in practice. LLMs can help sidestep this difficulty by making informal priors from natural language, code, and documentation directly available to the optimizer. However, existing LLM-based BO methods either insert the LLM into a fixed role (surrogate, acquisition proxy, or configuration interface) or hand it broad control, sacrificing the systematic exploration that makes BO reliable. We introduce agentic Bayesian optimization: a paradigm in which an LLM agent is the central decision maker in the BO loop while a Bayesian backend provides the uncertainty-aware optimization substrate. The agent configures the problem, queries the backend, selects and commits evaluations, and can revise the optimization strategy during the run by tightening bounds, switching acquisition functions, proposing targeted evaluations, or even reframing the problem following new instructions or observed evidence. We instantiate this idea in Sara, a surrogate-augmented autoresearch agent, and lenz, a modular BoTorch-based backend that the agent can inspect and modify through a structured interface. Across synthetic and real-world benchmarks, Sara preserves the reliability of state-of-the-art BO without prior knowledge, outperforms LLM-based baselines, and uses natural-language priors to improve beyond standard BO. We further demonstrate the practical value of agentic BO in dynamic settings, where Sara reconfigures the full optimization problem on the fly as requirements change, a capability not previously available in standard BO.
Optimization of LLM training and inference configurations, such as hyperparameters, data mixtures, and prompts, is critical to performance, but it is often approached heuristically in practice, leading to potentially suboptimal outcomes. By framing them as noisy, expensive, and derivative-free optimization problems, Bayesian optimization (BO) and other black-box optimization (BBO) methods offer a promising yet underexplored direction for principled, sample-efficient methods. However, LLM training and inference costs are prohibitively high for most of the BBO research community, and new methods are often only evaluated on synthetic test functions and small-scale datasets that fail to capture the challenges of modern LLM optimization problems. This impedes the development of BBO methods and makes it difficult to assess their effectiveness on modern LLM tasks. We introduce BoLT, the first LLM-centric benchmark that democratizes LLM research for the BBO community. BoLT is released at https://github.com/chewwt/bolt. BoLT covers broad and well-motivated LLM optimization problems, involving multi-fidelity, multi-objective, heteroscedastic noise, and high-dimensional search spaces. Each problem in BoLT is grounded in real experimental data and made fully reproducible and accessible through lightweight surrogate models fitted to the results of thousands of real LLM experiments. We benchmark BoLT against an extensive range of BO and BBO methods, showing that selected BO methods consistently outperform others across tasks and highlighting gaps in existing BBO methods on LLM tasks, underscoring the need to modernize benchmarks for the BBO community.
Ruth Wan Theng Chew, Zhiliang Chen, Apivich Hemachandra +1
Bayesian Optimization (BO) is widely used for optimizing expensive black-box functions, yet many real-world optimization problems contain substantially richer information than function evaluations alone. Examples include training curves in hyperparameter optimization, expert notes and images in scientific experimentation, and prior knowledge about where optima may lie. We show that large language models (LLMs) can effectively leverage such rich auxiliary information to guide optimization. Motivated by these findings, we develop three methods for incorporating auxiliary information into BO using LLMs. Across hyperparameter optimization benchmarks and a real-world nuclear fusion optimization task, our methods consistently outperform both standard BO and existing LLM-based optimization approaches. Our results demonstrate the effectiveness of LLMs for leveraging rich auxiliary information in BO.
Tejus Gupta, Efe Mert Karagözlü, Rohit Sonker +2
School of Computer Science Carnegie Mellon University Pittsburgh, PA 15213