cs.LGSep 25, 2026

DOHF: Online Diffusion Fine-tuning with Doob's hh-transform Guidance

Authors: Zhengyi Guo, Jiayuan Sheng, Wenpin Tang, David D. Yao

Abstract

Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to estimate. In this work, we propose Diffusion Online hh-guidance Fine-tuning (DOHF), which turns Doob's hh-transform into a practical online training algorithm. DOHF assigns optimality weights to generated samples, estimates the normalized local correction ∇log⁡h\nabla\log h under the current rollout policy, and distills it directly into the generative model. Theoretically, we characterize the population-optimal DiffusionNFT update as well as the various classfier free guidance methods through a unified hh-transform perspective. Methodologically, our framework accommodates black-box and non-differentiable rewards without additional network evaluations. We further show improved alignments under three empirical scenarios. Our work demonstrates how adapting probabilistic conditioning through inexpensive estimation and iterative distillation can improve generative learning across statistical sampling and visual generation.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Are we really tilting? The mechanics of reward guidance in flow and diffusion models

    Jun 1, 2026Sanjit Dandapanthula, Nicholas M. BoffiReward GradientsReward Functions

  2. Explicit Critic Guidance for Aligning Diffusion Models

    May 26, 2026Zhengyang Liang, Qihang Zhang, Ceyuan YangDiffusion AlignmentDiffusion-Based Reinforcement Learning Methods