Language models are increasingly used as proposal models for black-box search, from program optimization to molecular design. Existing approaches typically improve proposals through iterative prompting or parameter updates, offering limited control over how completely and efficiently the model's search space is explored. Continuous optimization methods, such as Bayesian optimization, provide a principled way to search but require a suitable domain to operate over. To address this, we introduce BOReFT, which learns a compact, low-dimensional space of hidden-state interventions in a frozen language model, and uses this space as the search domain for Bayesian optimization with an external scoring function. Empirically, we find that the learned domain spans semantic regions and exhibits smoothness properties that support search. Theoretically, we show that semantic coverage and interpolation control the best score available in the learned space, and that decoding from this space yields a standard stochastic-bandit observation model for adaptive search. We evaluate BOReFT on the interpretable word search task "Semantle" and on three more real-world discovery tasks in de novo molecule property optimization. Compared to strong LLM baselines, BOReFT finds in Semantle a higher number of hidden targets and, on two out of three molecular objectives, achieves higher property scores. Consequently, our method provides a principled new bridge between discrete proposal spaces of LLM-based search and continuous black-box optimization.
Figures & tables
Figure 1: BOReFT overview. Bayesian optimization searches a learned intervention manifold B , where each code b steers an LLM activation h′ to generate a discrete candidate scored by objective f .
Figure 2
Figure 3: Best-so-far performance on held-out Semantle search. Best semantic similarity found at each verification step, averaged over 5 targets and 3 warmstart runs, with a band of ±1 SD. We plot BOReFT only in the after-SFT block to group together methods that all use the training data.
Before SFT
After SFT
Method
EM ( ↑ )
Sim. ( ↑ )
Novelty ( ↑ )
EM ( ↑ )
Sim. ( ↑ )
Novelty ( ↑ )
Discrete BO
0/15
0.802
0.000
0/15
0.802
0.000
Random
0/15
0.710
0.561
1/15
0.797
0.119
SDPO-TTT
0/15
0.709
0.679
0/15
0.719
0.676
AutoDiscovery
0/15
0.757
0.736
1/15
0.787
0.154
MiGrATe
1/15
0.798
0.630
1/15
0.785
0.621
Table 1: Held-out Semantle search. Exact-match counts and mean best similarity over 15 runs (5 held-out targets and 3 warmstarts) with the rate of novelty w.r.t. the train set. Bold and underline mark the best and second-best. BOReFT scores are repeated under before-SFT for clarity.
Before SFT
After SFT
Method
DRD2 ( ↑ )
GSK3 β ( ↑ )
JNK3 ( ↑ )
DRD2 ( ↑ )
GSK3 β ( ↑ )
JNK3 ( ↑ )
Random
0.346/0.717
0.310/0.590
0.106/0.110
0.235/0.403
0.220/0.260
0.094/0.120
SDPO-TTT †
– / 0.981
– / 0.280
– / 0.180
– / 0.780
– / 0.340
– / 0.190
AutoDiscovery
0.534/0.929
0.316/0.380
0.158/0.160
0.264/0.801
0.226/0.290
0.094/0.160
MiGrATe
0.880/1.000
0.452/0.580
0.216/0.260
0.844/0.958
0.526/0.710
0.218/0.260
BOPRO
0.253/1.000
0.470/0.630
0.284/0.420
0.086/0.152
0.322/0.520
0.100/0.130
Table 2: Molecular property optimization. Mean and best ( mean/max ) activation probability over 5 seeds at a budget of 500 evaluations. Bold and underline mark the best and second-best. BOReFT scores are repeated under before-SFT for clarity. † SDPO-TTT did not finish all five seeds.
Figure 4: Training-set size and method ablations. Left: Mean best similarity and EM on Semantle, along with median coverage error over 30 runs across train and test. Right: Mean best similarity and median interpolation distance Δinterp on the same 30 runs, with EM counts in orange.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Search against baselines on molecular property optimization. Best-so-far property at each verification, mean ± 1 SD across seeds. The top row is before SFT and the bottom row is after SFT. The first ten evaluations are shared warm starts. We plot BOReFT only after SFT; its scores are repeated in the before-SFT columns of Table 2 . SDPO-TTT, AutoDiscovery, and MiGrATe are the no-instruction runs. A curve with no band is a single finished seed.
Method
Update
Proposal rule
Scored / round
Random (base)
none
task prompt only
1
Random (post-SFT)
none
same, after LoRA-SFT on the training set
1
OPRO
none
scored-history in-context prompt
1
BOPRO
GP over embeddings
OPRO prompt over 5 nearest neighbors
1
AutoDiscovery
UCB1 MCTS
OPRO prompt on the selected branch
1 of 8
MiGrATe
LoRA GRPO
mixed on-policy / greedy / neighborhood group
4
Appendix
Table 3: Search-method settings. All methods use the same base model within a task. AutoDiscovery expands eight completions per node and scores one; the rest remain as untried children. MiGrATe scores four new completions per round (two on-policy and two neighborhood); greedy group members are reused from history. BOPRO’s GP is fit in the Qwen3 embedding space of observed solutions, not in the BOReFT code space.
Before SFT
After SFT
Method
Exact ↑
Sim. ↑
Exact ↑
Sim. ↑
Discrete BO
15/15
1.000
–
–
Random
0/15
0.758
3/15
0.825
SDPO-TTT
0/15
0.763
0/15
0.766
AutoDiscovery
0/15
0.772
2/15
0.811
MiGrATe
0/15
0.788
0/15
0.782
Appendix
Table 4: Semantle search on training targets, before and after supervised fine-tuning. The table reports exact-match counts and mean best semantic similarity on five targets from the representation-training set, over fifteen runs. The left block uses the frozen generator. The right block uses the LoRA adapter of Section B.6 . BOReFT does not use that adapter. Discrete BO searches the training vocabulary, so the right block has no entry for it. In each column, bold marks the best value and underline marks the second best. The italic row is a reference, and it is not included in that ranking.
Semantle
Molecules
Method
Before SFT
After SFT
Before SFT
After SFT
Discrete BO
0.000
0.000
–
–
Random
0.791
0.201
0.212
0.068
SDPO-TTT
0.750
0.747
0.176
0.183
AutoDiscovery
0.364
0.218
0.075
0.066
MiGrATe
0.678
0.650
0.112
0.102
Appendix
Table 5: Proposal repetition before and after supervised fine-tuning. Each entry is the fraction of post-warm-start proposals that repeat an earlier solution in the same run. Lower is better. Semantle pools the training-target and held-out runs. Molecular rates pool finished seeds on DRD2, GSK3 β , and JNK3. SDPO-TTT before SFT is the one finished JNK3 seed. After SFT, SDPO-TTT is one finished seed on each property, and AutoDiscovery omits four unfinished DRD2 seeds. Bold and underline mark the best and second-best rate in each column. The italic row is a reference, and it is not ranked. BOReFT does not use the fine-tuning adapter, so its entries repeat.
Train
Held out
Search domain
V/V0
Exact ↑
Sim. ↑
Rep. ↓
Exact ↑
Sim. ↑
Rep. ↓
AABB
at k=0
1
6/15
0.863
0.505
8/15
0.892
0.452
at k=0.1
4.71×105
3/15
0.819
0.627
4/15
0.836
0.577
at k=0.5
1.08×1021
0/15
0.758
0.851
0/15
0.746
0.844
at k=1
7.31×1032
0/15
0.726
0.943
0/15
0.707
0.913
Appendix
Table 6: Semantle search under alternative domains. Exact matches and mean best semantic similarity, using the protocol of Table 1 on the same checkpoint. Repetition ( ↓ ) uses the definition in Table 5 , pooled within each split: the fraction of post-warm-start proposals whose solution repeats an earlier proposal in the same run. V/V0 is the Lebesgue volume of the search domain relative to the mean box ( k=0 ). The axis-aligned bounding boxes follow Eq. 30 , and the covering ellipsoid is as described in Section B.8 over the posterior means.
Figure 6: Geometric prior via semantic representations. First two principal components of the frozen semantic representations that supply the geometric prior. The top row is molecular optimization and the bottom row is Semantle. The left column is Qwen3-Embedding-0.6B. The right column is the MiST chemistry-pretrained checkpoint on the top row and Llama-3.2-1B-Instruct on the bottom row. Points are colored by semantic category, with one palette per row. These are the representations passed to gψ . The learned Semantle codes are shown separately in Fig. 2 .
Train
Held out
Temperature
Exact ↑
Sim. ↑
Rep. ↓
Exact ↑
Sim. ↑
Rep. ↓
T=0
6/15
0.863
0.505
8/15
0.892
0.452
T=0.25
3/15
0.830
0.455
5/15
0.860
0.412
T=0.50
5/15
0.861
0.333
4/15
0.839
0.345
T=0.75
1/15
0.807
0.234
3/15
0.839
0.231
Appendix
Table 7: Semantle search under decoding temperature. Exact matches, mean best semantic similarity, and repetition rate, using the protocol of Table 1 on the same checkpoint and the mean-box domain of Eq. 8 . Repetition is defined as in Table 6 . T=0 is greedy decoding as in the primary runs.
Train
Held out
Kernel
Exact ↑
Sim. ↑
Rep. ↓
Exact ↑
Sim. ↑
Rep. ↓
Matérn- 2.5
6/15
0.863
0.505
8/15
0.892
0.452
Matérn- 1.5
3/15
0.834
0.533
6/15
0.865
0.454
Matérn- 0.5
5/15
0.857
0.321
4/15
0.857
0.365
RBF
4/15
0.837
0.493
4/15
0.832
0.498
Appendix
Table 8: Semantle search under GP kernel. Exact matches, mean best semantic similarity, and repetition rate, using the protocol of Table 1 on the same checkpoint and the mean-box domain of Eq. 8 . Repetition is defined as in Table 6 . All kernels use ARD lengthscales. Matérn- 2.5 is the primary setting.
Train
Held out
Acquisition
Exact ↑
Sim. ↑
Rep. ↓
Exact ↑
Sim. ↑
Rep. ↓
LogEI
6/15
0.863
0.505
8/15
0.892
0.452
UCB
4/15
0.845
0.461
6/15
0.865
0.442
Thompson sampling
2/15
0.813
0.201
3/15
0.826
0.205
Appendix
Table 9: Semantle search under acquisition function. Exact matches, mean best semantic similarity, and repetition rate, using the protocol of Table 1 on the same checkpoint and the mean-box domain of Eq. 8 . Repetition is defined as in Table 6 . LogEI is the primary setting. UCB uses coefficient β=0.2 .
Figure 7: Search under training ablations. Best-so-far semantic similarity (mean ± 1 SD across 3 seeds) on five training and five held-out Semantle targets. The first ten objective evaluations are shared warm starts.
Train
Held out
Method
Exact ↑
Sim. ↑
Exact ↑
Sim. ↑
Δinterp↓
BOReFT
6/15
0.863
8/15
0.892
0.338
w/o self-distillation
6/15
0.876
3/15
0.829
0.355
w/o reconstruction
2/15
0.821
5/15
0.852
0.396
w/o shared encoder
1/15
0.792
4/15
0.833
0.382
w/o variational training
0/15
0.759
3/15
0.799
0.356
Appendix
Table 10: Search and interpolation in ablated spaces. Exact matches and mean best semantic similarity on Semantle, split by whether the hidden target was included in representation training (15 runs per split: 5 targets × 3 seeds). Δinterp is the median distance from the average embedding of temperature- 1 decodes to the semantic interpolant xα , over 40 target pairs at t∈{0.4,0.5,0.6} . An endpoint-only reference that always emits the nearer training word has Δinterp=0.365 .
Figure 8: Search vs. intervention rank. Mean best semantic similarity (solid) and exact-match rate (dashed) at N=3072 , pooled over 10 targets × 3 seeds.
Optimization of LLM training and inference configurations, such as hyperparameters, data mixtures, and prompts, is critical to performance, but it is often approached heuristically in practice, leading to potentially suboptimal outcomes. By framing them as noisy, expensive, and derivative-free optimization problems, Bayesian optimization (BO) and other black-box optimization (BBO) methods offer a promising yet underexplored direction for principled, sample-efficient methods. However, LLM training and inference costs are prohibitively high for most of the BBO research community, and new methods are often only evaluated on synthetic test functions and small-scale datasets that fail to capture the challenges of modern LLM optimization problems. This impedes the development of BBO methods and makes it difficult to assess their effectiveness on modern LLM tasks. We introduce BoLT, the first LLM-centric benchmark that democratizes LLM research for the BBO community. BoLT is released at https://github.com/chewwt/bolt. BoLT covers broad and well-motivated LLM optimization problems, involving multi-fidelity, multi-objective, heteroscedastic noise, and high-dimensional search spaces. Each problem in BoLT is grounded in real experimental data and made fully reproducible and accessible through lightweight surrogate models fitted to the results of thousands of real LLM experiments. We benchmark BoLT against an extensive range of BO and BBO methods, showing that selected BO methods consistently outperform others across tasks and highlighting gaps in existing BBO methods on LLM tasks, underscoring the need to modernize benchmarks for the BBO community.
Ruth Wan Theng Chew, Zhiliang Chen, Apivich Hemachandra +1
Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization procedure: the algorithm determines which candidates to explore and how the search progresses, while the language model generates or refines prompt proposals. We introduce RLMOpt, a prompt optimizer that makes the search policy itself language-model-driven through a recursive language model (RLM). The RLM agent operates over a tool-based environment, inspecting task information, analyzing failures, generating candidates, allocating evaluation budget, and deciding when to stop. A deterministic harness complements the agent by enforcing objective scoring, Pareto-based selection, and regression constraints. We evaluate RLMOpt across four benchmarks spanning structured clinical information extraction (Chia), multi-hop question answering (HotpotQA), verifiable instruction following (IFBench-2025), and multi-turn tool-calling agents (BFCL). In a matched comparison at a single seed, RLMOpt obtains the best held-out score on all four benchmarks and leads the four-task mean (0.610 against 0.589 for GEPA). Repeating each benchmark across seeds yields 11 matched benchmark-seed comparisons, in which RLMOpt outperforms GEPA in 9 cases. Across all 11 runs, it never produced a prompt that underperformed its seed, whereas GEPA fell below its starting point twice. It is also more efficient, achieving these results with fewer search rollouts while producing prompts that are 27-79% the size of those produced by GEPA. Our results further show that optimization gains are determined primarily by the headroom available in the seed prompt, rather than by the search budget. Efficient optimization therefore depends on reaching the available headroom reliably and with minimal search
Gradient-free post-training has emerged as a compelling alternative to gradient-based optimization for large language models (LLMs), but existing approaches remain costly. We ask whether structured search can identify a strong single expert under a modest evaluation budget. Motivated by evidence that useful weight updates lie in low-dimensional subspaces, we apply Bayesian optimization within a random linear embedding of weight space. Our method requires no backpropagation and uses a Gaussian process surrogate to guide candidate evaluations efficiently. Across several reasoning benchmarks with Qwen2.5-Instruct models from 0.5B to 3B parameters, Bayesian optimization using five times less candidate evaluations matches or exceeds RandOpt. These results show that surrogate-guided search can substantially reduce the evaluation cost of gradient-free post-training while producing stronger deployable single experts.
Nigel Bastian Cendra, Abdelhamid Ezzerg, Fernando Julio Cendra +2
University College London · Institut Polytechnique de Paris · University of Oxford