Agentic High-Dimensional Bayesian Optimization with Hypothesis- and Evidence-Guided Search
Organizations: State Key Laboratory of Novel Software Technology, Nanjing University · School of Artificial Intelligence, Nanjing University
Abstract
High-dimensional Bayesian optimization (HDBO) seeks sample-efficient optimization when the number of variables is large relative to the evaluation budget. Recent LLM-based and agentic BO methods incorporate task knowledge and adapt search decisions during a run, but have primarily been evaluated on low- and moderate-dimensional problems. We ask whether this paradigm can transfer to the higher-dimensional regime. Our experiments show that these methods do not remain reliable in the high-dimensional regime, where the challenge is not only where to evaluate, but also which modeling assumption and search geometry to use when the objective's useful structure is unknown. We therefore introduce HERA, a Hypothesis- and Evidence-guided Research Agent that uses task context, optimization feedback, and structural diagnostics to revise search hypotheses, select and configure HDBO strategies, and determine their execution length. PRISM, its numerical optimization engine, generates and evaluates candidates sequentially within each search block, updating numerical models after each observation. HERA remains competitive with strong numerical HDBO baselines and outperforms the evaluated LLM-based and agentic methods on four metadata-free synthetic functions. Across eight real-world tasks, HERA achieves the best mean final objective among all evaluated systems on most benchmarks. Further analyses show that structural diagnostics change strategy use, metadata effects vary across tasks, and adaptive search blocks reduce inference cost.
Figures & tables
| Aspect | Sara + lenz | HERA + PRISM |
|---|---|---|
| Agent role | Experiment designer | Structure strategist |
| Decision target | Next evaluation point | Structural hypothesis |
| Candidate generation | LLM or lenz | PRISM only |
| Control level | Point-level | Strategy-level |
| Control frequency | Every evaluation | Every search block |
| Adaptation | Point and BO policy | HDBO strategy |
| Tool | Returned information or effect |
|---|---|
| get_state | Budget and progress, active strategy and scoreboard, geometry, surrogate credibility, structural evidence, and optional metadata/catalog. No state change. |
| propose_initial_solution | Validate and queue one complete initial design. The next search block evaluates it first. |
| set_search_strategy | Atomically switch, reconfigure, or restart the numerical strategy without an objective evaluation. |
| run_search_block | Return every observed value and best-so-far, block improvement, and remaining budget. |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Representation or construction | |
|---|---|---|
| Hartmann-100 | 100 | Six latent inputs formed by a seeded block-average embedding. |
| Ackley-200 | 200 | Dense orthogonal mixing over all coordinates; shifted coordinates. |
| Lévy-300 | 300 | Fifteen block-average latent inputs; shifted coordinates. |
| Rosenbrock-300 | 300 | Full 300-coordinate chained Rosenbrock; shifted coordinates. |
| Rover | 60 | Trajectory-design parameters. |
| HalfCheetah | 102 | Continuous-control policy parameters. |
| Tool | Contract | Budget |
|---|---|---|
| get_state | Inspect a bounded state view. Summary mode reuses compact evidence; full mode refreshes surrogate credibility and structure diagnostics. | None |
| propose_initial_solution | Propose one prior-informed design after the shared opening design. | One when evaluated |
| set_search_strategy | Atomically switch, reconfigure, or restart a strategy. Arguments are checked against the selected strategy’s semantic schema. | None |
| run_search_block | Execute the requested number of sequential proposals, validate every candidate, call the evaluator, and update the shared observation history. | Evaluations returned |
| Task | HERA | Full-dimensional BO | TuRBO |
|---|---|---|---|
| Rover | |||
| HalfCheetah | |||
| MOPTA | |||
| Mazda | |||
| Ant | |||
| MM1 |
| Task | Sobol | LLAMBO | Centaur | Sara |
|---|---|---|---|---|
| Rover | ||||
| HC | ||||
| MOPTA | ||||
| Mazda | ||||
| Ant | ||||
| MM1 |
| Channel | Principal quantities |
|---|---|
| Exploration geometry | Normalized OTSD, recent OTSD gain, observation entropy, and their trajectories. |
| Surrogate credibility | Posterior standard-deviation summaries over Sobol probes; held-out RMSE, standardized residuals, coverage, and . |
| Variable sparsity | LassoBO active-set size and fraction, selected coordinates, and a thresholded provisional verdict. |
| Low-dimensional subspaces | Surrogate-gradient energy dimension and its ratio to over probe points. |
| Additive structure | Interaction-graph fragmentation, number and size of candidate groups, linked pairs, and strongest normalized mixed differences. |
| Model | HalfCheetah | MOPTA | Mazda |
|---|---|---|---|
| DeepSeek-V4.1-Flash | |||
| Kimi-K3 | |||
| GPT-6-astra |