Bayesian optimization (BO) is a powerful paradigm for optimizing expensive black-box functions. Traditional BO methods typically rely on separate hand-crafted acquisition functions and surrogate models for the underlying function, and often operate in a myopic manner. In this paper, we propose a novel direct regret optimization approach that jointly learns the optimal model and non-myopic acquisition by distilling from a set of candidate models and acquisitions, and explicitly targets minimizing the multi-step regret. Our framework leverages an ensemble of Gaussian Processes (GPs) with varying hyperparameters to generate simulated BO trajectories, each guided by an acquisition function drawn from a pool of conventional choices and terminated by a Bayesian early stop criterion. These trajectories train an end-to-end Decision Transformer that selects the next query so as to improve the ultimate objective, following a dense training sparse learning paradigm: the transformer is trained on abundant simulated data, while a limited number of real evaluations refine the GPs online. On synthetic and real-world benchmarks, our method attains the best or near-best final simple regret against standard, lookahead, trust-region and amortized BO baselines, with the largest gains in high-dimensional settings. Ablations attribute the gains jointly to region-of-interest filtering and the learned policy, and matched-budget comparisons against explicit two-step lookahead acquisitions show that the advantage is not shared by lookahead alone.
Figures & tables
Figure 1 : The DRO framework. Gray arrows indicate processes involving simulated data, while green arrows correspond to real data. The dense training ( gray ) and sparse learning ( green ) steps are depicted in hollow arrows.
Figure 2 : Upper: Performance on Ackley Function across dimensions (2D, 5D, 10D, 20D - Best Objective Value Found). Higher values are better. PFNs4BO is omitted from Ackley 20D due to scalability. Lower: Performance on HPO Tasks (Simple Regret). Lower values are better.
Figure 3 : Performance on LunarLander and ablation studies for DRO on Ackley 10D. Higher values are better (Best Objective Value Found). All DRO variants in ablations run for 10 trials.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Performance comparison of different acquisition function strategies used within DRO’s simulation rollouts on the Ackley 10D function (Best Objective Value Found). ‘DRO ROTATE’ refers to the strategy of cycling through EI, UCB, PI, and MES for generating simulated trajectories. Higher values are better.
Figure 5 : Performance on the simplified Ackley benchmark (2D, 8D, 16D) reproducing the logEI study setting ( Ament et al., 2023 ) . Differences from the main Ackley experiments: (i) noiseless observations; (ii) no output shift; (iii) reduced asymmetric domain [(−32.768/4,32.768/2)]D ; (iv) Sobol initialization with Ninit=2(d+1) . DRO performs best or near-best across all shown dimensions. Higher is better (best objective value found).
Method
Ackley 8D
Ackley 16D
BO
0.154
0.145
DRO
3.680
3.548
SCoreBO
1.411
1.344
TuRBO
0.283
0.484
PFNs4BO
0.585
0.590
Appendix
Table 1 : Comparison of wall time (seconds/iteration) on Ackley benchmarks.
Method
ackley_2d
ackley_5d
ackley_10d
ackley_20d
xgboost
lunarlander
adam_wine_acc
adam_breast_acc
adam_iris_acc
DRO-PPO
-0.04 ± 0.01
-0.87 ± 0.22
-2.69 ± 0.16
-17.09 ± 0.00
0.940 ± 0.004
291.69 ± 33.09
0.873 ± 0.000
0.919 ± 0.000
0.992 ± 0.000
DRO
-0.12 ± 0.02
-1.69 ± 0.17
-3.12 ± 0.10
-4.39 ± 0.19
0.938 ± 0.004
344.12 ± 66.55
0.861 ± 0.014
0.922 ± 0.002
0.982 ± 0.001
Appendix
Table 2 : Ultimate Results (Best Objective Value Found): DRO vs. DRO-PPO. Best results are in bold.
Method
ackley_2d
ackley_5d
ackley_10d
ackley_20d
lunarlander
DRO-PPO
-0.68 ± 0.17
-2.89 ± 0.37
-4.49 ± 0.21
-17.09 ± 0.00
265.42 ± 33.45
DRO
-0.57 ± 0.23
-2.70 ± 0.27
-5.19 ± 0.78
-10.37 ± 0.67
317.22 ± 66.76
Appendix
Table 3 : Results at Iteration 100 (Intermediate Performance): DRO vs. DRO-PPO. Best results are in bold.
Method
Ackley 5D
Ackley 10D
BO (logEI)
−16.56±0.42
−19.29±0.10
Two-step EI
−16.87±0.42
−19.69±0.22
KG (one-step)
−5.58±0.28
−5.74±0.31
Two-step KG
−5.19±0.44
−7.61±0.25
TuRBO
−5.12±1.37
−7.13±0.88
DRO
−2.63±0.22
−5.05±0.71
Appendix
Table 4 : Matched-budget comparison with lookahead acquisitions on Ackley (best observed objective at iteration 100 ; mean ± standard error over 10 trials, higher is better).
Variant
it =100
it =200
it =500
logEI, global (paper run)
−19.29±0.11
−18.89±0.27
−18.49±0.27
DRO-GLOBAL (transformer, no ROI)
–
–
−15.12±0.33
ROI-logEI (one acquisition, no transformer)
−18.29±1.01
−16.17±1.51
−12.11±2.49
DRO-NoDT (ensemble + ROI + pool, no transformer)
–
–
−10.01±2.56
TuRBO (paper run)
−7.13±0.89
−3.51±1.04
–
DRO (paper run)
−5.05±0.75
−3.61±0.10
−3.12±0.10
Appendix
Table 5 : Component attribution on Ackley 10D (best observed objective, mean ± standard error over 10 trials, higher is better).
Rollout pool
it =100
it =200
it =500
DRO control (EI, UCB, PI, MES; fresh)
−4.59±0.30
−3.42±0.17
−2.77±0.13
+ two-step EI (uniform rotation)
−4.40±0.37
−3.32±0.14
−2.58±0.28
+ posterior variance ( 10% share)
−5.07±0.50
−3.60±0.21
−2.86±0.13
+ posterior variance (equal share) †
–
–
−8.70±1.32
Appendix
Table 6 : Composition of the rollout acquisition pool on Ackley 10D (best observed objective, mean ± standard error over 10 paired trials, higher is better). † Run under the stored-run protocol; compare with the paper run in Table 5 .
Ensemble
it =100
it =200
it =500
RBF only, uniform allocation (control)
−4.59±0.30
−3.42±0.17
−2.77±0.13
Mixed kernels, marginal-likelihood weights
−4.13±0.39
−3.09±0.18
−2.50±0.16
Appendix
Table 7 : Mixed-kernel ensemble with evidence-weighted rollout allocation on Ackley 10D (best observed objective, mean ± standard error over 10 paired trials, higher is better).