cs.LGOct 4, 2026

Direction-Conditioned Policies for Online Goal-Conditioned Reinforcement Learning

Authors: S K Swaminathan, Damiya Gondha, Theyanesh Eswaramoorthy Rajahkrishnan, Aritra Hazra

Organizations: Department of Computer Science and Engineering Indian Institute of Technology Kharagpur, India · Department of Mechanical Engineering Indian Institute of Technology Kharagpur, India

Abstract

Contrastive Reinforcement Learning (CRL) learns representations that estimate goal reachability, yet its policy remains conditioned on raw goals and therefore does not directly exploit the geometry encoded by its critic. We introduce Direction-Conditioned Policies (DCP), a method built around a small modification to CRL: DCP selects previously visited states as waypoints during online training and conditions the policy on their direction and distance in representation space. At deployment, DCP applies the same interface directly to the final goal, requiring neither waypoint selection nor planning. Across nine navigation and manipulation tasks, DCP attains higher final success rates than CRL on seven tasks and spends more time near the goal on seven. Controlled maze experiments further show that DCP captures shortest-path geometry more accurately and that the supplied direction causally influences the actor's behavior. We identify waypoint coverage and ranking as limits to exploration, and show that learned candidate generation improves goal reaching in two controlled mazes.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Direction-Conditioned Policies via Compositional Subgoal Scoring for Online Goal-Conditioned Reinforcement Learning

    Jun 15, 2026Swaminathan S K, Damiya Gondha, Theyanesh Eswaramoorthy Rajahkrishnan +1Goal-Conditioned Reinforcement LearningGoal-Conditioned Value Function

  2. Do Better Goal Representations Improve Goal-Conditioned Reinforcement Learning?

    Sep 30, 2026Syed Nazmus Sakib, Abdul Monaf Chowdhury, Nafiul Haque +2Goal-Conditioned Reinforcement LearningObject Goal Navigation

  3. Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning

    Sep 28, 2026Hengrui Zhang, Yuhu Cheng, C. L. Philip Chen +1Goal-Conditioned Reinforcement LearningHigh-Level Subgoal Generation