cs.LGFeb 5, 2026

Beyond Success Rates: Trainability and Extractability for Offline GCRL

Authors: Jan Malte TöpperwienAditya MohanMarius Lindauer

Organizations: Institute of Artificial Intelligence (LUH|AI), Leibniz University Hannover · L3S Research Center

Abstract

Offline goal-conditioned reinforcement learning (GCRL) is typically benchmarked by the best tuned success rate of each method. This score measures attainable performance, but it does not reveal how reliably a learned goal-conditioned signal can be extracted into a policy: a method could succeed across many value-learning and extraction settings, or only at a narrow, hard-to-find configuration. We study this gap across four methods, GCIQL, GCIVL, QRL, and CRL, under a shared advantage-weighted regression (AWR) extractor. For each method, we construct trainability landscapes over the optimizer learning rate, which affects value learning and actor optimization, and AWR temperature, which controls how selectively the actor imitates high-advantage transitions. Across AntMaze, Cube, and Scene, we observe distinct regimes: high-scoring methods may be broadly accessible or brittle, while broad relative basins may still sit below low absolute ceilings. To interpret these differences, we pair landscapes with post-hoc diagnostics of future-vs-random goal discrimination and AWR weight concentration. Their relationship to downstream success is task-dependent. On AntMaze, where future goals align with path-like progress, these diagnostics explain landscape regimes. On Cube and Scene, goal ranking and manipulation control decouple: methods can rank goals well while failing downstream, or succeed through action-conditioned advantages despite weak future-vs-random separation. These results show that peak tuned success alone does not establish broadly extractable goal-conditioned behavior. Trainability landscapes expose this gap, while extraction diagnostics offer a lower-cost lens on how learned signals become policies.

Explore similar work

May 25, 2026cs.LG

Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning

Offline goal-conditioned reinforcement learning (GCRL) provides a practical framework for obtaining goal-reaching policies from fixed datasets. However, learning a reliable goal-conditioned value function in long-horizon tasks remains challenging. In this paper, we identify erroneous generalization in goal-conditioned value functions as a fundamental bottleneck, and demonstrate that appropriate inductive bias in the value function is crucial for addressing the bottleneck. Building on these findings, we propose Latent-Aligned Value Learning (LAVL), an offline GCRL algorithm that integrates latent-representation-based value generalization with hierarchical planning in a unified framework. Extensive experiments on OGBench demonstrate that LAVL consistently outperforms existing offline GCRL methods, achieving the highest performance on 20 out of 22 datasets. Notably, LAVL exhibits strong performance in long-horizon tasks and trajectory stitching datasets, where prior methods suffer significant performance degradation. Our code is available at https://github.com/oh-lab/LAVL.git.
Hyungkyu Kang, Byeongchan Kim, Min-hwan Oh
May 10, 2026cs.LG

Functional Graphs for Predicting and Explaining Goal Failure in Sparse Goal-Conditioned RL

Sparse goal-conditioned reinforcement learning can produce policies whose failures are hidden by aggregate success rates. We analyze trained goal-conditioned value policies through the deterministic functional graphs induced by greedy evaluation: for each goal, every state maps to a single successor, decomposing behavior into attractors and basins. This reveals a local-to-global structure in learned policies. We define local goal support (LGS), a one-step statistic measuring the fraction of valid neighboring states whose greedy successor is the goal. In deterministic sparse GridWorlds, zero LGS exactly precludes goal entry from non-goal starts. Empirically, weak LGS is a strong diagnostic of goal-level failure across update rules, curricula, larger grids, and bottleneck geometries: the fixed rule LGS <= 0.5 identifies low-success goals with precision 0.921, recall 0.929, and F1 0.925 in the main 8x8 TD setting, with similar performance across variants. However, local support is not sufficient for global success: some supported goals still fail because distant states are captured by competing attractors or fragmented basin structure. We therefore introduce a compact post-hoc taxonomy of policy-induced graphs -- goal-dominant, competitor-dominated, partial/contested, and fragmented -- to characterize residual failure modes beyond local support. These results show that sparse GCRL failures can be understood as structured policy-induced dynamics, and that local one-step policy structure provides a cheap post-training diagnostic for goal-level failure.
Shalley Dash
May 7, 2026cs.LG

Unifying Goal-Conditioned RL and Unsupervised Skill Learning via Control-Maximization

Unsupervised pretraining has driven empirical advances in goal-conditioned reinforcement learning (GCRL), but its theoretical foundations remain poorly understood. In particular, an influential class of methods, mutual information skill learning (MISL), discovers behaviorally diverse skills that can later be used for downstream goal-reaching. However, it remains a theoretical mystery why skills learned through MISL should support goal-reaching. A subtle challenge is that both GCRL and MISL are umbrella terms: different GCRL tasks use distinct criteria for measuring goal-reaching performance, while different MISL methods optimize distinct notions of behavioral diversity. We address this challenge and unify GCRL and MISL as instances of control maximization. We identify three canonical GCRL formulations and prove that they are fundamentally inequivalent: they can induce incompatible optimal policies even in the same environment. Nevertheless, they all share a common interpretation: a well-performing goal-conditioned policy is one whose future trajectory is highly sensitive to the commanded goal, with the precise notion of sensitivity determined by the GCRL formulation. Noting that MISL objectives can be understood as measures of skill-sensitivity akin to goal-sensitivity, we show that MISL objectives are bounded by formulation-specific downstream goal-sensitivities. These bounds establish a precise correspondence between MISL methods and downstream GCRL tasks: for every GCRL formulation, there exists a matching MISL objective for which more diverse skills afford greater downstream goal sensitivity. Our results thus lay a theoretical foundation for RL pretraining and have important practical implications, such as suggesting which pretraining objectives to use when a user cares about a specific class of downstream tasks.
Alireza Modirshanechi, Benjamin Eysenbach, Peter Dayan +1