cs.AISep 28, 2026

Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning

Authors: Hengrui Zhang, Yuhu Cheng, C. L. Philip Chen, Xuesong Wang

Organizations: School of Information and Control Engineering, China University of Mining and Technology · School of Computer Science and Engineering, South China University of Technology

Abstract

Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoal decomposition; however, high-level decision-making still relies on noise-sensitive value estimates, leading to unstable behavior in complex environments. We address this limitation by proposing \textbf{D}iffusion \textbf{S}ubgoal \textbf{P}lanning (\textbf{DSP}), a diffusion-based framework for high-level subgoal generation. DSP casts high-level planning as guided generative inference over goal-conditioned subgoals and learns both conditional and unconditional flows, enabling classifier-free guidance to introduce a goal-directed bias at inference time. By removing explicit value-based guidance from high-level planning, DSP generates reachable and goal-directed subgoals through a generative model while retaining hierarchical execution. Experiments on offline GCRL benchmarks demonstrate that DSP outperforms prior methods on a range of navigation and manipulation tasks, with particularly strong performance in maze environments that require multi-step subgoal planning.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL

    Feb 3, 2026Jinwoo Choi, Sang-Hyun Lee, Seung-Woo SeoGoal-Conditioned Reinforcement LearningHigh-Level Subgoal Generation

  2. Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning

    May 25, 2026Hyungkyu Kang, Byeongchan Kim, Min-hwan OhGoal-Conditioned Reinforcement LearningGoal-Conditioned Value Function