The difficulty of learning goal-reaching policies is often attributed to a "curse of horizon" that manifests as bias accumulation in temporal-difference backups and noisy advantage estimates. In this work, we identify an additional informational curse of horizon in goal-conditioned policy learning, where increasing the goal relabeling horizon can significantly reduce policy generalization and performance. Through a series of controlled experiments with oracle planners, we decouple the goal horizons sampled during training from those that the policy is asked to reach at test time. Even when evaluated only on a sequence of nearby subgoals, goal-conditioned behavioral cloning (BC) policies suffer from severe, training horizon-dependent performance degradation that is mitigated by reinforcement learning (RL) objectives. We explain this phenomenon as a horizon-dependent decrease in the conditional mutual information between actions and hindsight-relabeled goals, and find empirically that both BC and RL policies trained on longer-horizon goals exhibit a shift in sensitivity from goal to state information, as measured by the policy's input Jacobians. Motivated by this observation, we find that distilling the input Jacobians of short-horizon policies into long-horizon policies yields significant performance gains, especially in combinatorial manipulation tasks. Taken together, our results highlight goal relabeling horizon as an important consideration when learning generalist policies from offline data.
Figures & tables
Figure 1: Training on more distant relabeled goals harms nearby goal-reaching capabilities . We record the success rates across policies trained on a geometric distribution of relabeling goal offsets with discount γ , where γ=1 corresponds to a uniform distribution over future states. At evaluation, policies are given optimal Hplan -step subgoals from oracle planners . Success rates are averaged over the last three checkpoints, three seeds, and four environments: cube-double , cube-triple , cube-quadruple , and scene .
Figure 2: Increasing training goal horizon globally decreases policy goal sensitivity . We measure the proportion of the input Jacobian norm corresponding to the goal at fixed state-goal offsets, plotted with log2 scaling in cube-triple . Results are averaged over the last three training checkpoints, three seeds, and 1024 randomly sampled state-goal pairs per horizon, with additional environments in Figure 8 .
Figure 3
Environment
Dataset
GCBC
GCIVL
GCIQL
QRL
CRL
HIQL
SAW
S 2 AW
pointmaze
pointmaze-medium-navigate-v0
9±6
63±6
53±8
82±5
29±7
79±5
97±2
96±3
pointmaze-large-navigate-v0
29±6
45±5
34±3
86±9
39±7
58±5
85±10
83±8
pointmaze-giant-navigate-v0
1±2
0±0
0±0
68±7
27±10
46±9
68±8
73±9
antmaze
antmaze-medium-navigate-v0
29±4
72±8
71±4
88±3
95±1
96±1
97±1
97±1
antmaze-large-navigate-v0
24±2
16±5
34±4
75±6
83±4
91±2
90±3
93±2
antmaze-giant-navigate-v0
0±0
0±0
0±0
14±3
16±3
65±5
73±4
87±3
Table 1: Evaluating Sobolev bootstrapping on offline goal-conditioned RL tasks. We compare our method’s average (binary) success rate ( % ) against the numbers reported in Park et al. (2024) and Zhou and Kao (2025) across the five test-time goals for each environment, averaged over the last three checkpoints and eight seeds ± standard deviations. Numbers within 5% of the best value in the row are in bold .
Figure 3: Sobolev bootstrapping exhibits more interpretable sensitivity structure . We plot the proportion of the total input Jacobian norm corresponding to each cube during a triple pick-and-place on-policy rollout for SAW without (top) and with (bottom) Sobolev bootstrapping. For fairness, we also analyze a scripted optimal trajectory in Figure 11 .
Figure 4: Sobolev bootstrapping recovers policy goal sensitivity . We compare the absolute goal Jacobian norm between S 2 AW and SAW, averaging over batches of goals sampled at fixed intervals from the current state up to the maximum horizon. Results are averaged across four seeds and 1024 state-goal pairs for each horizon, with shaded standard deviations.
Environment
GCIQL
GCIQL ℓ1
SAW
SAW ℓ1
SAW ℓ2
S 2 AW
cube-single
68±6
74±4
76±8
85±6
85±4
86±5
cube-double-task1
74±8
76±8
64±9
81±9
65±14
78±12
cube-triple-task1
13±3
16±10
18±11
24±9
17±12
74±9
cube-quadruple-task6
0±1
0±1
2±2
2±2
1±1
35±14
Table 2: Sobolev bootstrapping improves single pick-and-place performance in the presence of distractor objects. We evaluate policies trained in multi-cube environments on single pick-and-place tasks. While naive ℓ1 or ℓ2 Jacobian regularization [Appendix H.4 ] yields modest improvements on simpler tasks, Sobolev bootstrapping matches or outperforms all baselines. Other settings are identical to Table 1 .
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Training on distant relabeled goals harms nearby goal-reaching capabilities . We record the success rates across learned policies trained on a uniform distribution of relabeling goal offsets [1,Htrain] , where Htrain=∞ corresponds to a uniform distribution over all future in-trajectory states. At evaluation, policies are given optimal Hplan -step subgoals from oracle planners. Success rates are averaged over the last three checkpoints and three seeds.
Figure 6: Training on a larger proportion of distant relabeled goals harms nearby goal-reaching capabilities . We record the success rates across learned policies trained on a geometric distribution of relabeling goal offsets with discount γ , where γ=1 corresponds to a uniform distribution over future states. At evaluation, policies are given optimal Hplan -step subgoals from oracle planners. Success rates are averaged over the last three checkpoints and three seeds.
Figure 7: Training on more random relabeled goals harms nearby goal-reaching capabilities . We record the success rates across learned policies trained on a mixture of goals either sampled from a uniform distribution over future states or from an arbitrary dataset state with probability prand∈{0.25,0.50,0.75} . At evaluation, policies are given optimal Hplan -step subgoals from oracle planners. Success rates are averaged over the last three checkpoints and three seeds.
Figure 8: Increasing training goal horizon globally decreases policy goal sensitivity . We measure the proportion of the input Jacobian norm corresponding to the goal at fixed state-goal offsets, plotted with log2 scaling. Results are averaged over the last three training checkpoints, three seeds, and 1024 randomly sampled state-goal pairs per horizon.
Figure 9: Training on distant relabeled goals globally decreases policy goal sensitivity . We measure the proportion of the input Jacobian norm corresponding to the goal at fixed state-goal offsets, plotted with log2 scaling. Results are averaged over the last three training checkpoints, three seeds, and 1024 randomly sampled state-goal pairs per horizon.
Figure 10: Random goal sampling globally decreases policy goal sensitivity. We directly reduce the action-goal MI by replacing in-trajectory goals with randomly sampled dataset states with probability prandom∈{0.25,0.5,0.75} , and measure the goal Jacobian norm proportion at fixed state-goal offsets.
Figure 11: Sobolev bootstrapping pays attention to the correct object during optimal trajectories. To enable fair comparisons across later task stages, we plot the input Jacobian norms with respect to each object over an independent scripted optimal trajectory Ahn et al. (2025) for the cube-triple triple pick-and-place task. S 2 AW (top) allocates a higher proportion of the Jacobian norm to the focal cube both during approach and grasp, whereas SAW’s (bottom) Jacobian norm only changes after the cube is already grasped. The shaded area denotes when the cube is actively grasped.
Dataset
Expectile τ
AWR α
KLD β
Subgoal steps k
Sobolev weight λ
pointmaze-medium-navigate-v0
0.7
3.0
3.0
25
0.03
pointmaze-large-navigate-v0
0.7
3.0
3.0
25
1.00
pointmaze-giant-navigate-v0
0.7
3.0
3.0
25
1.00
antmaze-medium-navigate-v0
0.7
3.0
3.0
25
0.03
antmaze-large-navigate-v0
0.7
3.0
3.0
25
0.03
antmaze-giant-navigate-v0
0.7
3.0
3.0
25
0.30
Appendix
Table 3: Hyperparameters for SAW with Sobolev bootstrapping.
Dataset
λ=0
0.001
0.01
0.1
cube-double-play-v0
40±7
40±11
55±7
35±9
cube-triple-play-v0
4±2
9±2
30±7
23±10
scene-play-v0
63±6
54±10
71±5
62±5
antmaze-giant-navigate-v0
73±4
76±4
82±4
85±4
Appendix
Table 4: Sobolev loss weight sensitivity analysis. Sobolev bootstrapping is robust to different weighting factors within a reasonable scale (around 0.01 for cube-double and scene ). Notably, improvements on antmaze-giant and cube-triple are robust across various weighting scales. Experimental results are averaged across four seeds.
hours / 1M iters
GCIQL
SAW
S 2 AW
cube-quadruple-play
0.26
0.46
0.75
scene-play
0.25
0.45
0.74
humanoidmaze-giant-navigate
0.26
0.46
1.75
Appendix
Table 5: Training time comparison. We select the highest-dimensional environments and report the hours required to complete a standard 1M-step training run on an NVIDIA 5090 GPU.
Environment
Dataset
GCBC
GCIVL
GCIQL
QRL
CRL
HIQL
SAW
S 2 AW
S 2 AW ℓ
maze-giant
antmaze-giant-navigate-v0
0±0
0±0
0±0
14±3
16±3
65±5
73±4
87±3
80±5
humanoidmaze-giant-navigate-v0
0±0
0±0
0±0
1±0
3±2
12±4
35±4
38±6
35±5
cube
cube-single-play-v0
6±2
53±4
68±6
5±1
19±2
44±9
72±5
86±5
79±8
cube-double-play-v0
1±1
36±3
40±5
1±0
10±2
6±2
40±7
50±6
55±8
cube-triple-play-v0
1±1
1±0
3±1
0±0
4±1
3±1
4±2
30±5
30±4
cube-quadruple-play-v0
0±0
0±0
0±0
0±0
0±0
0±0
0±0
3±2
2±1
Appendix
Table 6: Architectural ablations to the Sobolev SAW target actor. We report the performance of Sobolev SAW where the target actor has the same architecture as the full policy, averaged over 4 seeds with standard deviations after the ± sign. Numbers within 5% of the best value in the row are in bold .
Dataset
RIS
Sobolev RIS
pointmaze-medium-navigate-v0
88±6
96±1
pointmaze-large-navigate-v0
63±13
75±3
pointmaze-giant-navigate-v0
57±12
67±5
antmaze-medium-navigate-v0
96±1
97±1
antmaze-large-navigate-v0
89±3
90±1
antmaze-giant-navigate-v0
65±4
69±0
Appendix
Table 7: Evaluating RIS with Sobolev bootstrapping. We evaluate the effectiveness of Sobolev bootstrapping on RIS by comparing it against the numbers reported in Zhou and Kao (2025) . Results are averaged over 4 seeds, and other settings are identical to Table 1 .
Environment
OTA n=5
OTA n=10
OTA n=20
SHARSA
DQC
S 2 AW
antmaze-giant-navigate
77 ±4
-
-
72 ±1
25 ±6
87 ±3
cube-double-play
2 ±0
3 ±1
2 ±1
5 ±1
24 ±4
57 ±6
cube-triple-play
1 ±0
0 ±0
1 ±0
2 ±1
11 ±4
30 ±5
Appendix
Table 8: Additional horizon reduction baselines . The additional baselines do not readily generalize to the standard datasets and perform poorly. All results are averaged across 4 seeds.
SAW
SAW ℓ1
SAW ℓ2
S 2 AW
antmaze-giant-navigate
73 ±4
72 ±3
74 ±2
87 ±3
cube-double-play
40 ±7
47 ±10
37 ±7
57 ±6
cube-triple-play
4 ±2
5 ±4
4 ±1
30 ±5
scene-play
63 ±6
69 ±8
62 ±2
74 ±6
Appendix
Table 9: Uniform Jacobian regularization. We apply ℓ1 and ℓ2 norm penalization with coefficient 0.3 to the SAW baseline. We observe that ℓ1 penalization is more effective, leading to notable performance gains in cube-double and scene environments. Results are averaged across four seeds and five evaluation tasks.
Department of Electrical and Computer Engineering, Seoul National University, Seoul, South Korea · Department of Automotive Engineering, Ajou University, Gyeonggi-do, South Korea