cs.AIOct 1, 2026

Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning

Authors: Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira, Marlos C. Machado

Organizations: Universidade Federal de Minas Gerais · University of Alberta · Alberta Machine Intelligence Institute (Amii) · Canada CIFAR AI Chair

Abstract

Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstraction, which treats k environment steps as a single transition, restores this signal at long range, but no single fixed k suits all state-goal distances: large k preserves value differences across long temporal distances while collapsing distinctions between nearby states, and small k does the reverse. We make this trade-off explicit and introduce Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on k. GITA trains one policy by aggregating advantage-weighted supervision across multiple k values, so scales assigning larger positive advantages to a state-goal pair contribute more strongly to its update. GITA does not need to choose between local resolution and long-range signal; it retains both without committing to a single k. On OGBench, GITA outperforms a broad range of offline GCRL baselines, raising average success rate across all tasks by 25 percentage points (73% relative improvement) over HIQL. It also improves over the strongest fixed-k method, OTA, by 7 percentage points (14% relative).

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GTRL: Grounding Divide-and-Conquer Value Learning with Temporal Differences

    Sep 27, 2026Abdul Monaf Chowdhury, MD Sameer Iqbal Chowdhury, Shifat E Arman +1Goal-Conditioned Reinforcement LearningTemporal Difference

  2. Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning

    May 25, 2026Hyungkyu Kang, Byeongchan Kim, Min-hwan OhGoal-Conditioned Reinforcement LearningGoal-Conditioned Value Function