Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning
Organizations: Universidade Federal de Minas Gerais · University of Alberta · Alberta Machine Intelligence Institute (Amii) · Canada CIFAR AI Chair
Abstract
Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstraction, which treats k environment steps as a single transition, restores this signal at long range, but no single fixed k suits all state-goal distances: large k preserves value differences across long temporal distances while collapsing distinctions between nearby states, and small k does the reverse. We make this trade-off explicit and introduce Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on k. GITA trains one policy by aggregating advantage-weighted supervision across multiple k values, so scales assigning larger positive advantages to a state-goal pair contribute more strongly to its update. GITA does not need to choose between local resolution and long-range signal; it retains both without committing to a single k. On OGBench, GITA outperforms a broad range of offline GCRL baselines, raising average success rate across all tasks by 25 percentage points (73% relative improvement) over HIQL. It also improves over the strongest fixed-k method, OTA, by 7 percentage points (14% relative).
Figures & tables
| Non-Hierarchical | Hierarchical | |||||||||
| Environment | Type | Size | GCBC | GCIVL | GCIQL | QRL | CRL | HIQL | OTA | GITA |
| PointMaze | navigate | large | ||||||||
| navigate | giant | |||||||||
| stitch | large | |||||||||
| stitch | giant | |||||||||
| AntMaze | navigate | large | ||||||||
| Environment | Type | Size | Unconditioned Value | Highest Advantage Factor | Mean Advantage Before Exponentiation | GITA with Geometric | GITA |
|---|---|---|---|---|---|---|---|
| PointMaze | navigate | large | |||||
| navigate | giant | ||||||
| stitch | large | ||||||
| stitch | giant | ||||||
| AntMaze | navigate | large | |||||
| navigate | giant |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Environment | Variant | Observation | Action Dim. | Max. Episode Length |
| PointMaze | large | state (2) | 2 | 1000 |
| giant | state (2) | 2 | 1000 | |
| AntMaze | large | state (29) / pixels ( ) | 8 | 1000 |
| giant | state (29) / pixels ( ) | 8 | 1000 | |
| HumanoidMaze | large | state (69) | 21 | 2000 |
| giant | state (69) | 21 | 4000 |
| Environment | Dataset | Variant | # Transitions | # Episodes | Data Episode Length |
| PointMaze | navigate | large | 1M | 1000 | 1000 |
| navigate | giant | 1M | 500 | 2000 | |
| stitch | large | 1M | 5000 | 200 | |
| stitch | giant | 1M | 5000 | 200 | |
| AntMaze | navigate | large | 1M | 1000 | 1000 |
| navigate | giant | 1M | 500 | 2000 |
| Hyperparameter | Value |
|---|---|
| – Optimization and Architecture – | |
| Learning rate | |
| Optimizer | Adam |
| Gradient steps | (state), (pixels) |
| Minibatch size | (state), (pixels) |
| Actor hidden dimensions | |
| Comparison | Raw | ||
|---|---|---|---|
| GITA vs. GCBC | |||
| GITA vs. GCIVL | |||
| GITA vs. GCIQL | |||
| GITA vs. QRL | |||
| GITA vs. CRL | |||
| GITA vs. HIQL |
| Comparison | Raw | ||
|---|---|---|---|
| GITA vs. Unconditioned Value | |||
| GITA vs. Highest Advantage Factor | |||
| GITA vs. Mean Advantage Before Exponentiation | |||
| GITA vs. Geometric |
| Environment | Type | Size | Best per Task | Best Global | |||
|---|---|---|---|---|---|---|---|
| AntMaze | navigate | large | |||||
| navigate | giant | ||||||
| stitch | large | ||||||
| stitch | giant | ||||||
| explore | large | ||||||
| HumanoidMaze | navigate | large |