cs.LGSep 30, 2026

Prequential E-Values for Selected-GP Near-Optimality Certificates

Authors: Ami Tavory, Noa Cohen

Organizations: Meta Platforms

Abstract

When optimizing an expensive black-box function sequentially, as in hyperparameter optimization, we may want to stop once the best evaluated value is certified within ε\varepsilon of the global optimum. Such a certificate needs two ingredients: a lower confidence bound for the selected value and an upper confidence envelope over the domain, typically supplied by a Gaussian process (GP). GP-UCB-style stopping rules are valid when the kernel and constants defining this envelope are fixed before the run, but the practical temptation is to tune the envelope from the same adaptive evaluations and then certify as if it had been fixed. We use prequential e-values to make this selection auditable: starting from a predeclared set of fully specified GP/RKHS envelopes, each candidate is tested by its own one-step-ahead e-process, contradicted candidates are deleted, and certification uses the largest upper bound among the survivors. With a valid selected-point lower bound and one declared candidate having valid latent coverage and noise calibration, the rule is anytime-valid. On a 512-seed noisy RBF stress sweep, it roughly halves false-certification risk at comparable power versus fit-then-certify. Relative to random fixed GP precommitment on smooth d=3,4d=3,4 objectives, each additional false certificate is accompanied by 3.0 and 13.5 additional correct certificates, respectively.

Explore similar work

May 21, 2026cs.LG

Regret-Based (ε,δ)(ε,δ)-optimal Stopping Criteria for Bayesian Optimization

Bayesian optimization (BO) is a widely used iterative black-box optimization method that utilizes Gaussian process (GP) surrogate models. In practice, BO is typically terminated after a fixed evaluation budget is exhausted, which can incur unnecessary cost and provides no optimality guarantee on solution quality. Recent research in developing a practical stopping criterion has made empirical progress, yet a theoretically sound stopping criterion remains a work in progress. In this work, we present provably tighter instantaneous regret bounds for GP upper confidence bound (GP-UCB) at any given iteration. Then, we propose stopping criteria for GP-UCB based on this tighter bound that ensures an εε-optimal solution with high probability 1−δ1-δ upon termination. Numerical experiments are performed to validate and demonstrate the effectiveness and efficiency of our stopping criteria.
May 19, 2026stat.ML

Goal-Oriented Lower-Tail Calibration of Gaussian Processes for Bayesian Optimization

Gaussian process (GP) predictive distributions are commonly used in Bayesian optimization (BO) to guide the selection of evaluation points for expensive objective functions. The choice of kernel and hyperparameters has a strong influence on the exploration--exploitation trade-off. For minimization, sampling criteria such as expected improvement (EI) depend on both the probability mass below the current best value and the shape of the predictive distribution in this region. This article studies goal-oriented calibration of GP predictive distributions below a low threshold tt in the noiseless setting, for standard GP models with hyperparameters selected by maximum likelihood. We consider two complementary forms of calibration below tt for inputs distributed according to a reference measure μμ: occurrence calibration over the design space and thresholded μμ-calibration on sublevel sets of the form {x∈X,f(x)≤t}\{x\in\mathbb{X}, f(x)\le t\}. We propose tcGP, a post-hoc method that combines these two forms of calibration for GP predictive distributions below tt. With fixed GP hyperparameters, the exact EI sampling criterion based on tcGP generates a sequence of evaluation points that is dense in the design space. Experiments on standard benchmarks show improved lower-tail calibration and BO performance relative to standard GP models and globally calibrated GP models.
Jun 7, 2026cs.LG

A Joint Finite-Sample Certificate for Adaptive Selective Conformal Risk Control

Selective predictors answer on confident inputs and abstain elsewhere; deploying one safely needs a single finite-sample certificate that simultaneously upper-bounds the selected risk, lower-bounds the acceptance probability \pacc\pacc above a floor \pmin\pmin, and lower-bounds the deployment utility. This certificate must be valid under adaptive threshold selection from a finite grid of mm pairs on \ncert\ncert samples. We give such a certificate for bounded, possibly non-monotone losses by treating the selected risk directly as a ratio rather than through a Hoeffding-style range bound. The construction couples three confidence bounds: a variance-adaptive empirical-Bernstein bound on the ratio risk, a Clopper--Pearson bound on acceptance, and a two-sided closeness bound on utility. Together they lower-bound the certified policy's utility absolutely and to within 2\gammau2\gammau of the best over the \emph{certified set}, both non-vacuous whenever feasible; a regime-scoped third leg matches an external oracle, informative only where the risk margin \gammar<α\gammar < α and vacuous at the headline operating points. Relative to the range-only Hoeffding-ratio construction this sharpens the acceptance-floor dependence from 1/\pmin1/\pmin to 1/\pmin1/\sqrt{\pmin}, and a closed-form corollary identifies a per-pair regime in which our risk bound dominates a Hoeffding conformal risk control (Hoeffding--CRC) selective bound. Empirically, on ImageNet (three ResNets) and COCO val 2017 panoptic, the certificate opens a +22+22 pp certified-acceptance frontier over Hoeffding--CRC and is ≈10×{\approx}10{\times} tighter than a non-vacuous matched-valid baseline; these gains are regime-scoped, not universal, and absent on ADE20K. The certifier runs in O(\ncertm)O(\ncert m) time.