Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning
Organizations: School of Information and Control Engineering, China University of Mining and Technology · School of Computer Science and Engineering, South China University of Technology
Abstract
Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoal decomposition; however, high-level decision-making still relies on noise-sensitive value estimates, leading to unstable behavior in complex environments. We address this limitation by proposing \textbf{D}iffusion \textbf{S}ubgoal \textbf{P}lanning (\textbf{DSP}), a diffusion-based framework for high-level subgoal generation. DSP casts high-level planning as guided generative inference over goal-conditioned subgoals and learns both conditional and unconditional flows, enabling classifier-free guidance to introduce a goal-directed bias at inference time. By removing explicit value-based guidance from high-level planning, DSP generates reachable and goal-directed subgoals through a generative model while retaining hierarchical execution. Experiments on offline GCRL benchmarks demonstrate that DSP outperforms prior methods on a range of navigation and manipulation tasks, with particularly strong performance in maze environments that require multi-step subgoal planning.
Figures & tables
| Datasets | Flat Policies | Hierarchical Policies | ||||||
|---|---|---|---|---|---|---|---|---|
| GCBC | CFGRL | GCIVL | OTA | Pi-HIQL | HIQL | HIQL w/o | DSP | |
| pointmaze-medium-navigate-v0 | ||||||||
| pointmaze-large-navigate-v0 | ||||||||
| pointmaze-giant-navigate-v0 | ||||||||
| pointmaze-teleport-navigate-v0 | ||||||||
| antmaze-medium-navigate-v0 | ||||||||
| Datasets | Succ. | ||||
|---|---|---|---|---|---|
| AntMaze-Giant | 1 | 1.6 | 100.0 | 68.8 | 46.0 |
| 3 | 1.7 | 100.0 | 62.1 | 69.2 | |
| 5 | 1.8 | 98.8 | 51.6 | 56.4 | |
| 10 | 2.1 | 96.9 | 40.6 | 27.0 | |
| Scene-Play | 1 | 1.8 | 98.0 | 92.4 | 35.6 |
| 3 | 2.0 | 96.5 | 90.7 | 50.4 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Hyperparameter | Value |
|---|---|
| Learning rate | |
| Optimizer | Adam |
| Batch size | |
| Total gradient steps | |
| MLP dimensions | |
| Activation function | GELU |
| Environment | Datasets | |||
| pointmaze | pointmaze-medium-navigate-v0 | |||
| pointmaze-large-navigate-v0 | ||||
| pointmaze-giant-navigate-v0 | ||||
| pointmaze-teleport-navigate-v0 | ||||
| pointmaze-medium-stitch-v0 | ||||
| pointmaze-large-stitch-v0 |
| Datasets | MLE | Diffusion | HIQL | DSP |
|---|---|---|---|---|
| antmaze-medium-navigate-v0 | ||||
| antmaze-large-navigate-v0 | ||||
| antmaze-giant-navigate-v0 | ||||
| cube-single-play-v0 | ||||
| cube-double-play-v0 | ||||
| scene-play-v0 |
| Datasets | Representation | Merlin | HDMI | HD | DSP |
|---|---|---|---|---|---|
| PointMaze-Large | Full XY | ||||
| PointMaze-Giant | Full XY | ||||
| AntMaze-Large | Full | ||||
| XY | |||||
| AntMaze-Giant | Full | ||||
| XY |
| Method | GCIVL | OTA | Pi-HIQL | HIQL | DSP |
|---|---|---|---|---|---|
| Training Time (min) | 24 | 56 | 53 | 42 | 57 |
| Datasets | Flat Policies | Hierarchical Policies | |||||||
|---|---|---|---|---|---|---|---|---|---|
| GCBC | CFGRL | GCIVL | HIQL | HIQL w/o | DSP | DSP 1 | DSP mix | DSP | |
| pointmaze-medium-stitch-v0 | |||||||||
| pointmaze-large-stitch-v0 | |||||||||
| pointmaze-giant-stitch-v0 | |||||||||
| pointmaze-teleport-stitch-v0 | |||||||||
| antmaze-medium-stitch-v0 | |||||||||
| Datasets | Fraction | HIQL | HIQL w/o | DSP | |
|---|---|---|---|---|---|
| AntMaze-Large | 1.00 | 1.00 | |||
| 0.50 | 0.96 | ||||
| 0.25 | 0.91 | ||||
| 0.10 | 0.85 | ||||
| HumanoidMaze-Giant | 1.00 | 1.00 | |||
| 0.10 | 0.94 |
| Datasets | HIQL | HIQL w/o | DSP |
|---|---|---|---|
| antmaze-medium-explore-v0 | |||
| antmaze-large-explore-v0 | |||
| antmaze-teleport-explore-v0 |
| Datasets | HIQL | DSP |
|---|---|---|
| visual-antmaze-medium-navigate | ||
| visual-antmaze-large-navigate | ||
| visual-antmaze-giant-navigate | ||
| visual-scene-play |
| Datasets | HDMI | DTAMP | HD | SIHD | DSP |
|---|---|---|---|---|---|
| antmaze-umaze-v2 | |||||
| antmaze-umaze-diverse-v2 | |||||
| antmaze-medium-play-v2 | |||||
| antmaze-medium-diverse-v2 | |||||
| antmaze-large-play-v2 | |||||
| antmaze-large-diverse-v2 |
| Datasets | PPO | HIQL | DSP | |
|---|---|---|---|---|
| ShadowHandOver | 46.5 | |||
| ShadowHandCatchUnderarm | 37.2 | |||
| ShadowHandOver | 46.5 | |||
| ShadowHandCatchUnderarm | 37.2 |