Abstract
We study the problem of estimating the guidance that steers the distribution learned by a diffusion generative model toward a tilted target q0∝wp0 at inference time. Relying on the stochastic optimal control approach, we observe that the exact drift correction is the gradient of the logarithm of Doob's h-function, and we study the problem of estimating it from a sample. In the present paper, we assume that the score of the pretrained model is available, that the tilting weight is bounded and positive, and that the reference distribution has a bounded support, no smoothness of the weight is required. Introducing a penalized least-squares risk in which the penalty is the residual of the space-time harmonicity equation satisfied by the h-function, measured in a dual Sobolev norm, we derive high-probability bounds on the squared error of the resulting guidance estimate. Since the penalty vanishes at the target, the estimator is free of regularization bias, and in favourable scenarios its rate of convergence is faster than the minimax rate of estimating first-order derivatives of a smooth regression function. Assuming that w is bounded and positive with Ep0[w−s]<∞ for some s∈(0,∞], and that the reference data are compactly supported, we prove that the guidance is estimable in squared L2 at rate εns/(s+4), where εn=n−2(β−1)/(2(β−1)+d). We also transfer the obtained bounds to the total variation distance between the marginals of the estimated and the exactly guided samplers, and illustrate the performance of the suggested approach with numerical experiments.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Explore similar work
Sep 25, 2026cs.LG
Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to estimate. In this work, we propose Diffusion Online
h-guidance Fine-tuning (DOHF), which turns Doob's
h-transform into a practical online training algorithm. DOHF assigns optimality weights to generated samples, estimates the normalized local correction
∇logh under the current rollout policy, and distills it directly into the generative model. Theoretically, we characterize the population-optimal DiffusionNFT update as well as the various classfier free guidance methods through a unified
h-transform perspective. Methodologically, our framework accommodates black-box and non-differentiable rewards without additional network evaluations. We further show improved alignments under three empirical scenarios. Our work demonstrates how adapting probabilistic conditioning through inexpensive estimation and iterative distillation can improve generative learning across statistical sampling and visual generation.
Zhengyi Guo, Jiayuan Sheng, Wenpin Tang +1
Sep 30, 2026stat.ML
Inference-time alignment of flow and diffusion-based models is critical for achieving flexible generative modeling. Theoretically, Doob's
h-transform provides an elegant solution to this problem, and most existing methods are based on this principle. However, in practice, estimating the optimal guidance derived from Doob's
h-transform at inference time is challenging. To deal with this issue, we regard inference-time alignment as a sequential optimization problem in the space of probability measures and propose a novel framework called
Steepest Guidance, based on the principle of maximizing local improvement in the objective. We provide a theoretical analysis of the proposed method and demonstrate its effectiveness through extensive experiments.
Shokichi Takakura, Akifumi Wachi, Rei Higuchi +2
LY Corporation · The University of Tokyo · RIKEN AIP
Jun 1, 2026cs.LG
Reward guidance algorithms steer a learned generative process toward the reward-tilted measure at inference time. While empirically powerful, these methods are prone to reward hacking: the guided model over-optimizes the reward at the cost of fidelity to the learned distribution. Prior work has attributed this to the complexity of neural reward functions or implicit biases in diffusion training, but its fundamental origins remain poorly understood. We show that reward hacking arises from an approximation made in most practical implementations of reward-guided diffusion -- finite-particle plug-in estimation of the Doob h-function -- even in the simplest non-trivial settings of Gaussian and Gaussian mixture targets with quadratic rewards. In closed form, we isolate two distinct failure modes of the plug-in estimator: it leads to reward hacking within each mode and it cannot select high-reward modes. We propose a closed-form reward damping schedule that corrects the within-mode bias with no additional compute, and clarify the role of best-of-n sampling in compensating for the mode selection failure. Experiments on Gaussian mixture targets, a 2D checkerboard, and FLUX.1 text-to-image generation confirm that our theoretical insights carry over to practical settings.
Sanjit Dandapanthula, Nicholas M. Boffi
Carnegie Mellon University