Do Better Goal Representations Improve Goal-Conditioned Reinforcement Learning?
Authors: Syed Nazmus Sakib, Abdul Monaf Chowdhury, Nafiul Haque, Shifat E Arman, Md Mehedi Hasan
Organizations: Department of Robotics and Mechatronics Engineering, University of Dhaka, Bangladesh · Department of Computer Science, University of Oxford, UK
Goal-conditioned reinforcement learning (GCRL) relies heavily on how target goals are represented to the policy. While recent methods encode goals via temporal distance, occupancy, or controllability, it remains unclear how much downstream performance actually depends on representation quality. We study this in offline GCRL by constructing an exact temporal-distance goal representation in deterministic mazes. We then systematically corrupt its geometric quality while keeping the downstream learner fixed. Across OGBench navigation tasks and two algorithms, large changes in goal-representation quality produce almost no change in performance. However, applying the same interventions to the agent's current state more than doubles success, revealing the state pathway as the true bottleneck. Building on this insight, we show that simple random Fourier positional encodings substantially improve performance on the hardest navigation tasks without map information or objective modifications. Overall, our findings suggest that in state-based offline navigation, improving how the agent's current state is represented matters far more than refining the goal representation. Code will be released soon.
Figures & tables
Figure 1: Two sides of the same interface. Left, top: goal-side representations, the axis this literature optimises. A goal is described by its temporal distances to other states. Left, bottom: the property that motivates them, namely that such distances compose, so that a policy accurate within one radius extends to two and four ( Myers et al., 2025a ) . Right: our proposal. We encode the agent’s own position instead, lifting observations that are close together in raw coordinates into a representation in which they are far apart, and from which the policy acts.
Figure 2: Goal representations are saturated; the state pathway is not. GCIVL on antmaze-large. Grey: published goal methods. Yellow: exact BFS goal representation. Violet: adding unprivileged state position features.
Environment
Algo
Dual (published)
Learned (ours)
Ideal (ours)
Difference
pointmaze-med
GCIVL
76±7
70.9±6.1 (5)
62.7±3.0 (5)
−8.2[−15.6,−0.7]
pointmaze-large
GCIVL
46±6
46.4±6.5 (5)
46.3±7.9 (5)
−0.1[−10.7,+10.5]
antmaze-med
GCIVL
75±4
76.5±7.4 (5)
68.4±4.0 (5)
−8.0[−17.2,+1.1]
antmaze-large
GCIVL
28±11
32.2±8.9 (8)
30.0±8.0 (8)
−2.2[−11.3,+6.9]
humanoid-med
GCIVL
29±3
27.8±3.6 (5)
34.2±3.1 (5)
+6.4[+1.5,+11.4]
pointmaze-med
CRL
33±1
37.8±3.3 (5)
57.4±11.6 (5)
+19.5[+5.4,+33.7]
Table 1: An exactly optimal goal representation does not reliably improve control. GCIVL and CRL, concatenation interface, success rate ×100 , mean ± standard deviation over n seeds given in parentheses. The published column is taken from Park et al. (2026a) , their Table 1 for GCIVL and Table 6 for CRL, over 8 seeds. The difference is ideal minus learned with a 95% Welch interval; bold marks intervals excluding zero.
Figure 3: Exact goal representations yield no consistent training advantage. Evaluation curves (mean ± 1 SE over 2–8 seeds per panel; violet is learned, gold exact). Exact and learned representations exhibit indistinguishable learning rates and asymptotic success.
Figure 4: Control is insensitive to goal quality. Results on antmaze-large (mean ± 1 SD, 5–8 seeds). Across corruption ladders spanning exact geodesics to zero distance information, performance differences are statistically indistinguishable from zero.
Figure 5: Position features. Success with and without the features of equation equation 6 , for both goal representations, at frequency scales corresponding to wavelengths 8.5 , 2.1 and 0.5 maze cells.
Figure 6: State-side interventions on antmaze-large (GCIVL; mean ± 1 SD, 5–8 seeds). (a) Applying an identical fixed random N(0,1) code to the state more than doubles success, while doing nothing on the goal side. (b) The state code does not require distance geometry: corrupted and purely random codes match or outperform exact geodesics.
Figure 7: Controls for the method. antmaze-large, GCIVL. Position features work without any goal representation at all (second bar). Standardising the coordinates in place, which changes their scale without adding a basis, is severely harmful (fifth bar), and widening the actor changes nothing (sixth bar).
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Optimiser
Adam
Learning rate
3×10−4
Batch size
1024
Gradient steps
106
Hidden dimensions (value, actor, representation)
(512,512,512)
Nonlinearity
GELU
Appendix
Table 2: Hyperparameters. Shared values follow the benchmark. Algorithm-specific and method-specific values are grouped below them.
Figure 8: Exact and learned goal representations across the full navigation suite. Mean ± 1 SD over 5 to 8 seeds. CRL was not run on humanoidmaze-medium.
Figure 9: Effect of the two corruption procedures on temporal-distance structure. Spearman correlation between the decoded and true distance, within quartiles of true distance. Gaussian corruption degrades all ranges; far-field corruption keeps the nearest quartile exact and reverses the furthest.
Figure 10: Success against feature wavelength. The runs of Figure 5 , plotted against wavelength in maze cells. Dashed lines in the matching colour mark the same agent without position features.
Figure 11: Full interface sweep. Cross-attention minus concatenation, with everything else held fixed within each row. Whiskers are 95% Welch intervals. Violet marks intervals excluding zero in favour of cross-attention, gold in favour of concatenation, grey a tie.
Figure 12: Interfaces on manipulation tasks. GCIVL with the learned representation. Cross-attention is worse than concatenation on cube-single-play, −3.6[−6.5,−0.7] , and no better on scene-play, −6.5[−13.2,+0.2] .
This paper investigates robust representation learning in offline goal-conditioned reinforcement learning (GCRL). Particularly in sparse reward scenarios, learning representations that align state and goal latents is a challenge that frequently culminates in representation divergence where the encoder drifts toward a low-dimensional, goal-agnostic subspace that destabilizes policy learning. We address this issue by showing that an agent must acquire a fundamental understanding of its environment across multiple scales, from local physical dynamics to long-horizon goal-directed structure. Building on this insight, we propose Ms.PR, a framework that leverages multi-scale predictive supervision to enforce goal-directed alignment within the latent space. We demonstrate that Ms.PR leads to improved representation quality and strong performance on both vision and state-based tasks. Furthermore, we show that our approach is exceptionally resilient under realistic, challenging data regimes, maintaining state-of-the-art performance across a wide variety of tasks, trajectory stitching scenarios, and extreme noise conditions.
Valliappan Chidambaram Adaikkappan, David Meger, Sai Rajeswar +1
Goal-conditioned reinforcement learning hinges on how the goal is encoded. Contrastive, metric, temporal-distance, and information-theoretic encoders differ in objective. They still share one trait. None of them sees the current state. Such a state-independent embedding cannot mark which part of the goal still needs action. The policy must then recover that cue by inverting both encoders. We propose DAGR. It refines the static embedding of any late-fusion encoder into a state-conditioned one through multi-scale gated cross-attention. A near-identity gated residual preserves the base representation. Difference-aware Goal Cross-Attention then biases the attention scores using a per-token state-goal discrepancy map. On OGBench, DAGR improves navigation. Our ablations trace the gain to the gated residual, not to the difference bias that names the method. On manipulation and puzzle tasks it matches or falls below the base. DAGR is a structured refinement, not a universal improvement.
Xing Lei, Wenyan Yang, Xuetao Zhang +1
Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University · Department of Electrical Engineering and Automation, Aalto University · School of Engineering, Westlake University
Offline goal-conditioned reinforcement learning (GCRL) provides a practical framework for obtaining goal-reaching policies from fixed datasets. However, learning a reliable goal-conditioned value function in long-horizon tasks remains challenging. In this paper, we identify erroneous generalization in goal-conditioned value functions as a fundamental bottleneck, and demonstrate that appropriate inductive bias in the value function is crucial for addressing the bottleneck. Building on these findings, we propose Latent-Aligned Value Learning (LAVL), an offline GCRL algorithm that integrates latent-representation-based value generalization with hierarchical planning in a unified framework. Extensive experiments on OGBench demonstrate that LAVL consistently outperforms existing offline GCRL methods, achieving the highest performance on 20 out of 22 datasets. Notably, LAVL exhibits strong performance in long-horizon tasks and trajectory stitching datasets, where prior methods suffer significant performance degradation. Our code is available at https://github.com/oh-lab/LAVL.git.