cs.LGSep 29, 2026

Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows

Authors: Junhyun Ha, Juho Lee, Byoungwoo Park

Organizations: KAIST

Abstract

Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or deterministic editors limit expressiveness to represent multiple separated modes for the same proposal. In this work, we introduce Proposal-Conditioned Refinement Flows (PReFlow), a policy extraction method combining critic-based proposal selection with a conditional refinement flow. To optimize proposal selection and refinement together, we formulate a KL-regularized objective whose optimum induces a Gibbs policy over final actions under a Gaussian-smoothed behavior prior. The refinement flow can represent multiple high value modes, while a proposal-centered Gaussian reference regulates large action changes. This Gaussian reference further enables us to make use of simulation-free, closed form adjoint matching targets from sampled endpoints and critic gradients, yielding a single velocity regression loss without a backward adjoint solve. On 50 OGBench tasks, PReFlow achieves competitive offline performance and the highest aggregate score among the compared methods after online fine-tuning, reaching 91% after 500K environment steps.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Refinement-based Flow Policy Optimization

    Sep 14, 2026Bumgeun Park, Hyukjun Yang, Donghwan LeeFlow PoliciesDistributional Reinforcement Learning

  2. ReFPO: Reflow Regularization for Flow Matching Policy Gradients

    Jun 19, 2026Ge Wang, Yibo Peng, Fan Feng +10Flow PoliciesOffline Reinforcement Learning

  3. Discrete Flow Matching for Offline-to-Online Reinforcement Learning

    May 12, 2026Fairoz Nower Khan, Nabuat Zaman Nahim, Peizhong JuFlow PoliciesOffline Reinforcement Learning