cs.LGSep 27, 2026

BOReFT: Manifold Steering of Language Models for Black-box Optimization

Authors: Dhruv Agarwal, Rico Angell, Kavitha Srinivas, Tahira Naseem, Horst Samulowitz, Willie Neiswanger, Andrew McCallum

Organizations: University of Massachusetts Amherst · New York University · IBM Research · University of Southern California

Abstract

Language models are increasingly used as proposal models for black-box search, from program optimization to molecular design. Existing approaches typically improve proposals through iterative prompting or parameter updates, offering limited control over how completely and efficiently the model's search space is explored. Continuous optimization methods, such as Bayesian optimization, provide a principled way to search but require a suitable domain to operate over. To address this, we introduce BOReFT, which learns a compact, low-dimensional space of hidden-state interventions in a frozen language model, and uses this space as the search domain for Bayesian optimization with an external scoring function. Empirically, we find that the learned domain spans semantic regions and exhibits smoothness properties that support search. Theoretically, we show that semantic coverage and interpolation control the best score available in the learned space, and that decoding from this space yields a standard stochastic-bandit observation model for adaptive search. We evaluate BOReFT on the interpretable word search task "Semantle" and on three more real-world discovery tasks in de novo molecule property optimization. Compared to strong LLM baselines, BOReFT finds in Semantle a higher number of hidden targets and, on two out of three molecular objectives, achieves higher property scores. Consequently, our method provides a principled new bridge between discrete proposal spaces of LLM-based search and continuous black-box optimization.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. BoLT: A Benchmark to Democratize Black-box Optimization Research for Expensive LLM Tasks

    May 16, 2026Ruth Wan Theng Chew, Zhiliang Chen, Apivich Hemachandra +1Large Language Model BenchmarksBlack-Box Optimization

  2. RLMOpt: Adaptive Prompt Optimization via Recursive Language Models

    Aug 11, 2026Subhash Bangalore Satheesha, Nirvik Pande, Deepthi Duddempudi +1Automatic Prompt OptimizationPrompt Engineering

  3. Can Bayesian Optimization Efficiently Find a Strong Single Expert in Neural Thickets?

    Aug 11, 2026Nigel Bastian Cendra, Abdelhamid Ezzerg, Fernando Julio Cendra +2Large Language Model TrainingGradient-Based Optimization