Direction-Conditioned Policies for Online Goal-Conditioned Reinforcement Learning
Organizations: Department of Computer Science and Engineering Indian Institute of Technology Kharagpur, India · Department of Mechanical Engineering Indian Institute of Technology Kharagpur, India
Abstract
Contrastive Reinforcement Learning (CRL) learns representations that estimate goal reachability, yet its policy remains conditioned on raw goals and therefore does not directly exploit the geometry encoded by its critic. We introduce Direction-Conditioned Policies (DCP), a method built around a small modification to CRL: DCP selects previously visited states as waypoints during online training and conditions the policy on their direction and distance in representation space. At deployment, DCP applies the same interface directly to the final goal, requiring neither waypoint selection nor planning. Across nine navigation and manipulation tasks, DCP attains higher final success rates than CRL on seven tasks and spends more time near the goal on seven. Controlled maze experiments further show that DCP captures shortest-path geometry more accurately and that the supplied direction causally influences the actor's behavior. We identify waypoint coverage and ranking as limits to exploration, and show that learned candidate generation improves goal reaching in two controlled mazes.
Figures & tables
| Stage | Target | Policy input |
|---|---|---|
| Online collection | Scored visited waypoint | |
| Actor update | Achieved future state | |
| Deployment | Requested final goal |
| Success rate (%) | Time near goal | |||||
|---|---|---|---|---|---|---|
| Task | CRL | SSGC | DCP | CRL | SSGC | DCP |
| AntMaze Big | ||||||
| AntMaze Hardest | ||||||
| Ant U-Maze | ||||||
| Humanoid U-Maze | ||||||
| AntPush | ||||||
| Method | Collection target | Open room | U-maze |
|---|---|---|---|
| CRL | Final goal | ||
| Direct direction conditioning | Final goal | ||
| SSGC | Scored waypoint | ||
| DCP | Scored waypoint |
| U-maze measure | CRL | DCP |
|---|---|---|
| Spearman correlation with graph distance | ||
| Spearman correlation with Euclidean distance | ||
| Graph-correct local descent (%) |
| Task | Visited pool | Generator | Generator + projection |
|---|---|---|---|
| Two-route maze | |||
| Rare-gateway maze |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
|---|---|
| Actor / critic hidden width | 256 |
| Hidden layers | 2 |
| Representation dimension | 64 |
| Actor, critic, and temperature learning rates | |
| Discount | 0.99 |
| Minibatch size | 256 (Humanoid: 128) |
| Tasks | CRL actor | DCP actor | Actor increase | Total increase |
|---|---|---|---|---|
| AntMaze Big / Hardest / U-Maze | 78,096 | 94,224 | 20.65% | 6.37% |
| Humanoid U-Maze | 144,162 | 160,034 | 11.01% | 4.15% |
| AntPush / AntSoccer | 78,608 | 94,736 | 20.52% | 6.35% |
| PusherEasy / Hard / Hard Far | 75,534 | 91,406 | 21.01% | 6.40% |
| Task | Goal coordinates | Radius | Training steps |
|---|---|---|---|
| AntMaze Big | Ant position | 0.5 | 50M |
| AntMaze Hardest | Ant position | 0.5 | 50M |
| Ant U-Maze | Ant position | 0.5 | 50M |
| Humanoid U-Maze | Humanoid position | 0.5 | 100M |
| AntPush | Ant position (movable obstacle) | 0.5 | 100M |
| AntSoccer | Ball XY | 0.5 | 50M |
| Experiment | Training and evaluation |
|---|---|
| Nine-task comparison | Five seeds per method and task; 50M or 100M steps. |
| Four-way collection/input comparison | Five seeds, 10M steps, 20 evaluations, separately in each layout. |
| Representation geometry | Five encoders per method and layout, trained for 10M steps. |
| Candidate-source comparison | Five seeds per candidate source and maze, 10M steps. |
| Waypoint substitution | Three policies; 20,000 paired probes per policy. |
| Differential-drive intervention | Three policies; 42 directed pairs and four headings per policy. |
| Pool intervention | Pre-motion | Successful path | Paired wins |
|---|---|---|---|
| Reserve 128 of 512 slots | 32.39 | 38.17 | 3/3 |
| Add 128 to 512 online slots | 38.09 | 36.26 | 0/3 |
| Variant | Pretraining | Final success (%) | Mean training success (%) |
|---|---|---|---|
| Fresh joint DCP | 0 | 25.4 | 16.72 |
| Frozen CRL goal encoder | 50M | 22.3 | 11.60 |
| Both CRL encoders frozen | 50M | 23.0 | 12.46 |
| Environment | Method | 10M | 25M | 50M | 75M | 100M |
|---|---|---|---|---|---|---|
| AntMaze Big | DCP | 18.4 5.4 | 23.8 3.3 | 29.9 2.9 | – | – |
| CRL | 15.1 3.7 | 17.7 5.3 | 18.3 6.5 | – | – | |
| SSGC | 18.3 1.2 | 22.0 2.9 | 23.3 3.0 | – | – | |
| AntMaze Hardest | DCP | 9.5 2.5 | 14.5 2.1 | 16.1 2.3 | – | – |
| CRL | 5.8 2.5 | 10.5 3.8 | 12.2 5.6 | – | – | |
| SSGC | 7.7 1.6 | 13.2 4.8 | 14.2 8.0 | – | – |
| Environment | Method | 10M | 25M | 50M | 75M | 100M |
|---|---|---|---|---|---|---|
| AntMaze Big | DCP | 78.7 37.1 | 94.6 26.2 | 115.3 14.6 | – | – |
| CRL | 75.3 12.7 | 83.2 29.4 | 89.2 28.0 | – | – | |
| SSGC | 91.9 19.1 | 94.6 20.5 | 99.8 9.0 | – | – | |
| AntMaze Hardest | DCP | 30.3 12.6 | 65.1 12.9 | 63.2 13.3 | – | – |
| CRL | 28.4 15.9 | 39.9 24.0 | 48.1 23.7 | – | – | |
| SSGC | 32.6 8.1 | 59.0 21.0 | 66.0 41.4 | – | – |