Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning horizon before generating plan content, even though the appropriate horizon depends on the route itself. A horizon that is too short can force infeasible transitions, whereas one that is too long can introduce redundant motion. We introduce HorizonFlow, a hierarchical planner that treats plan length as an output of generation rather than a prescribed input. Its subgoal route planner guides its action-prefix controller through a sequence of latent subgoals. Both components combine insertion-based generation with flow matching to jointly generate continuous plan content and length, using the partially generated plan to guide token insertion. HorizonFlow reuses the resulting length information to select candidates and steer generation toward shorter plans without a separate learned value model. Across Maze2D, Multi2D, and OGBench navigation and visual manipulation benchmarks, HorizonFlow achieves the highest average performance among the compared methods.
Figures & tables
Figure 1: Schematic of horizon mismatch and count-based selection. (a, b) Prescribed horizons that are too short or too long yield an infeasible jump or detour, respectively. (c) HorizonFlow generates variable-length candidates and selects the one with the fewest tokens.
Figure 2: HorizonFlow overview. (a) Training predicts missing-token counts and denoising velocities from deleted and noised offline plans. (b) Sampling inserts and refines tokens before σ=1 , then only refines; optional Feynman–Kac (FK) steering guides generation, followed by minimum-count selection. (c) The route planner (RP) selects a route; the prefix controller (PC) selects an action prefix toward its first subgoal, with receding-horizon execution.
Figure 3: Route-planner samples on Maze2D-Large. Rows show start–goal pairs; columns show generation time σ and generated count n . Latent subgoals are mapped to maze coordinates, connected and colored by order coordinate rk . White circles/orange stars mark starts/goals. Beyond the dashed cutoff at σ=1 , counts remain fixed while refinement continues. These are generated plans, not executed trajectories.
Environment
Task
Diffuser
VHD
HDMI
HD
DF
SSD
Ours
Maze2D
U-Maze
113.9 ± 3.1
118.5 ± 6.7
120.1 ± 5.6
128.4 ± 36.0
116.7 ± 2.0
144.6 ± 7.6
137.5 ± 0.8
Medium
121.5 ± 2.7
130.5 ± 3.6
121.8 ± 3.6
135.6 ± 30.0
149.4 ± 7.5
134.4 ± 13.6
154.8 ± 1.5
Large
123.0 ± 6.4
142.9 ± 7.1
128.6 ± 6.5
155.8 ± 25.0
159.0 ± 2.7
183.5 ± 19.2
205.2 ± 2.4
Single-task Average
119.5
130.6
123.5
139.9
141.7
154.2
165.8
Multi2D
U-Maze
128.9 ± 1.8
137.6 ± 3.9
131.3 ± 4.0
144.1 ± 12.0
119.1 ± 4.0
158.2 ± 10.1
145.1 ± 1.2
Medium
127.2 ± 3.4
146.3 ± 2.0
131.6 ± 4.2
140.2 ± 16.0
152.3 ± 9.9
155.2 ± 17.7
170.5 ± 1.0
Table 1: Benchmark performance. Values are mean ± standard deviation; highlighted entries mark row-wise highest means. HorizonFlow uses 5, 8, and 4 training seeds for Maze2D/Multi2D, OGBench navigation, and visual manipulation, respectively, with 100 episodes per environment and protocol for Maze2D/Multi2D and 50 per task for OGBench.
Figure 4: Planning-horizon and candidate-count diagnostics on Maze2D-Large. Both models generate XY routes tracked by a shared proportional–derivative (PD) controller. (a–c) Fixed-length (no-insertion) model with prescribed horizon H ; outlines mark H=N∗ . (d–f) Joint content–length model with FK steering and minimum-count selection over K candidates. Panels show normalized score, reach rate, and median steps to goal among successful rollouts; columns are nominal-step bins N∗ (100 pairs each). Darker is better; colors are normalized per column across both models.
Figure 5: Sampling, replanning, and inference cost on OGBench navigation. Columns vary RP candidate count (a, e), PC candidate count (b, f), RP update period relative to training stride (c, g), and executed prefix fraction (d, h), with one-step execution shown separately; other settings stay at their defaults (stars). Rows report mean success (a–d) and estimated amortized model-inference time (e–h; ms/step, log scale), averaged equally over the nine navigation environments on a matched evaluation subset. Solid/dashed curves denote FK steering on/off; horizontal lines in (e–h) mark baseline timing references. Error bars combine per-environment standard deviations across seeds (details in Appendix F ).
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Pool
Success (%)
Available
Miss (%)
Median ρ
Δ steps
FK-off, K=16
99.86
600/600
0.14
0.82
−104
FK-on, K=16
100.00
600/600
0.00
0.87
−42
FK-on, K=64
99.94
600/600
0.06
0.84
−54
FK-on, K=128
100.00
600/600
0.00
0.83
−58
FK-on, K=256
99.92
600/600
0.08
0.82
−63
Appendix
Table 2: Within-pair execution diagnostics for count-based selection. Success is the expectation under uniform minimum-count tie-breaking; available counts pools containing a successful candidate, and miss reports expected selection failure conditional on that availability. Median ρ summarizes within-pool Spearman correlations between count and steps to goal among successful candidates. Δ steps compares minimum-count selection with the mean steps to goal of successful candidates in the same pool; negative values denote earlier arrival.
Figure 6: Diffuser diagnostics on Maze2D-Large. (a–c) Normalized score, reach rate, and median steps to goal among successful rollouts across planning horizon H and nominal steps N∗ . Outlines mark H=N∗ . (d) Normalized score under different horizon-assignment and selection protocols: LP denotes learned length prediction; single, value, and exec denote single-candidate, value-based, and execution-based selection. Darker is better; colors are normalized within columns in (a–c) and across all cells in (d).
Figure 7: Diffuser plan samples on Maze2D-Large. Rows show start–goal pairs labeled by nominal steps N∗∈{64,192,320} ; columns vary the planning horizon H∈{64,128,256,384} . Each panel overlays five generated position sequences with the same endpoints. Circles mark starts, stars mark goals, and gray regions are walls. These are generated reference plans, not executed trajectories.
Figure 8: HorizonFlow XY plan samples on Maze2D-Large. Rows use the same start–goal pairs as Figure 7 ; columns compare prescribed planning horizons with adaptive generation (variable). Colors distinguish generated reference plans, dots mark tokens, circles mark starts, stars mark goals, and gray regions are walls. Count labels in this figure include both anchors; ranges give the minimum and maximum counts among the displayed samples.
Figure 9: Length signals on matched Maze2D-Large start–goal pairs. (a–c) HIQL value-derived length, VHD predicted length, and HorizonFlow generated non-anchor count versus nominal steps N∗ ; ρ denotes Spearman correlation. (d) Correlation recomputed on pairs with N∗≥x as the minimum nominal-step threshold x increases.
Figure 10: Full-system route-planner length controls. Success rates on AntMaze-Giant (left) and HumanoidMaze-Giant (right). The bars compare minimum-count RP selection with FK enabled or disabled, a single joint RP sample, maximum-count selection without FK, and a VHD-style predicted-length control with one content sample. The RP uses 16 candidates unless K=1 is indicated. Error bars are standard deviations over three training seeds, with 125 evaluation problems per seed. The pretrained encoder and PC are shared by the matched controls.
Figure 11: Action-prefix controller controls with the route planner fixed. Success rates on AntMaze-Giant (left) and HumanoidMaze-Giant (right) for the default minimum-count PC, a single generated candidate, a full-horizon fixed-length Transformer, maximum-count selection without FK, and a behavior-cloning policy. Each domain uses 125 problems for each of seeds 43 – 45 .
Figure 12: Sampler implementation checks and resolution sensitivity. Top: gap-target assignment and conservation checks, missing-count head calibration, and rank correlation between intermediate FK scores and final generated counts. Bottom: rollout success, average generated non-anchor count, and fraction of RP candidates at the length capacity as the RP integration resolution varies. Solid blue curves enable FK and dashed gray curves disable it; circles denote AntMaze-Giant and triangles denote HumanoidMaze-Giant.
Setting
Route planner (RP)
Prefix controller (PC)
Generated token
(r,z) , dz=16
(r,a) ; state-free interiors
Capacity (anchors included)
64
32
Clean count distribution
Log-uniform [2,64]
Log-uniform [3,32]
Bucket capacities
8, 16, 32, 64
8, 16, 32
Transformer width / blocks
640 / 10
512 / 4
Attention heads
8
8
Appendix
Table 3: Default OGBench architecture, training, and inference settings.
Environment
T
ΔRP
HRP
Kdiv
NPC
Batch
Maze2D-UMaze
300
5
2
1
10
1024
Maze2D-Medium
600
10
5
1
20
1024
Maze2D-Large
800
13
6
1
26
1024
PointMaze / AntMaze
1000
16
8
1
32
1024
HumanoidMaze-Medium / Large
2000
32
16
2
32
1024
HumanoidMaze-Giant
4000
64
32
4
32
1024
Appendix
Table 4: Environment-dependent settings of the common recipe. T is the episode limit, ΔRP the training stride, HRP the RP update period, and NPC the PC capacity including anchors. Batch size applies to both stages.
Environment
Diffuser
HD
DF
PointMaze Medium
256
500
250
PointMaze Large
500
500
500
PointMaze Giant
500
1000
500
AntMaze Medium
500
500
500
AntMaze Large
500
500
1000
AntMaze Giant
500
500
250
Appendix
Table 5: Selected training horizon parameters for reproduced OGBench planning baselines, in environment steps. HD uses a high-level span parameter; DF uses a training window length.
Dataset
GPU
HorizonFlow
Diffuser
HD
DF
PointMaze-Medium
RTX 5090
12.1 (4.8 + 7.2)
4.9
4.2
0.8
PointMaze-Large
RTX 5090
12.1 (4.8 + 7.2)
6.8
4.1
1.4
PointMaze-Giant
RTX 5090
12.2 (4.8 + 7.3)
6.8
4.4
1.4
AntMaze-Medium
RTX 5090
12.1 (4.8 + 7.2)
6.8
4.1
1.5
AntMaze-Large
RTX 5090
12.1 (4.8 + 7.2)
6.8
4.1
3.6
AntMaze-Giant
RTX 5090
12.1 (4.8 + 7.2)
6.9
4.1
0.8
Appendix
Table 6: Wall-clock training time per seed in hours. HorizonFlow entries report total time (RP + PC). Dashes indicate unreported measurements.
Figure 13: Architecture ablations on OGBench navigation. HorizonFlow is compared with replacing the prefix controller by a behavior-cloning policy (BC for PC) and removing the hierarchy (flat). Bars report mean success across navigation environments and training seeds; error bars show standard deviations across seed-level averages.
Environment
Control period
Diffuser
HD
DF
HorizonFlow
PointMaze-Medium
100
0.27–236.57
0.26–285.88
0.26–217.11
4.54–41.54
PointMaze-Large
100
0.27–248.96
0.27–285.73
0.28–293.53
4.76–42.13
PointMaze-Giant
100
0.27–249.04
0.27–297.68
0.27–296.08
4.65–42.14
AntMaze-Medium
100
0.27–250.56
0.27–289.47
0.28–290.66
4.81–42.75
AntMaze-Large
100
0.26–250.26
0.28–287.69
0.28–422.73
4.76–42.71
AntMaze-Giant
100
0.28–249.87
0.27–289.06
0.28–217.24
4.61–42.58
Appendix
Table 7: Step-level device-inference latency estimates on OGBench navigation, in milliseconds. Each entry gives the controller-only minimum and the constructed planning-update upper estimate. The reference control period is included for comparison.
Figure 14: Sampling, replanning, and inference cost on PointMaze. Blocks show Medium, Large, and Giant, top to bottom. Each block reports success rate (%) (a–d) and estimated inference time (e–h, ms/step, logarithmic scale) for RP and PC candidate counts, RP update ratio, and executed prefix fraction, respectively. Solid/dashed curves indicate FK steering on/off; stars mark defaults.
Figure 15: Sampling, replanning, and inference cost on AntMaze. Blocks show Medium, Large, and Giant, top to bottom. Each block reports success rate (%) (a–d) and estimated inference time (e–h, ms/step, logarithmic scale) for RP and PC candidate counts, RP update ratio, and executed prefix fraction, respectively. Solid/dashed curves indicate FK steering on/off; stars mark defaults.
Figure 16: Sampling, replanning, and inference cost on HumanoidMaze. Blocks show Medium, Large, and Giant, top to bottom. Each block reports success rate (%) (a–d) and estimated inference time (e–h, ms/step, logarithmic scale) for RP and PC candidate counts, RP update ratio, and executed prefix fraction, respectively. Solid/dashed curves indicate FK steering on/off; stars mark defaults.
Figure 17: Sampling, replanning, and inference cost on visual cube manipulation. Blocks show Cube-Single, Cube-Double, and Cube-Triple, top to bottom. Each block reports success rate (%) (a–d) and estimated inference time (e–h, ms/step, logarithmic scale) for RP and PC candidate counts, RP update ratio, and executed prefix fraction, respectively. Solid/dashed curves indicate FK steering on/off; stars mark defaults.
Figure 18: Sampling, replanning, and inference cost on visual Scene. Panels (a–d) report success rate (%), and (e–h) report estimated inference time (ms/step, logarithmic scale), for RP and PC candidate counts, RP update ratio, and executed prefix fraction, respectively. Solid/dashed curves indicate FK steering on/off; stars mark defaults.
Figure 19: Sampling, replanning, and inference cost on Maze2D. Blocks show U-Maze, Medium, and Large, top to bottom. Each block reports D4RL normalized score (a–d) and estimated inference time (e–h, ms/step, logarithmic scale) for RP and PC candidate counts, RP update ratio, and executed prefix fraction, respectively. Solid/dashed curves indicate FK steering on/off; stars mark defaults.
School of Information and Control Engineering, China University of Mining and Technology · School of Computer Science and Engineering, South China University of Technology
Department of Electrical and Computer Engineering, Seoul National University, Seoul, South Korea · Department of Automotive Engineering, Ajou University, Gyeonggi-do, South Korea