stat.MLOct 1, 2026

Optimal Transport Meets Reinforcement Learning: A Survey

Authors: Yujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni Montana

Organizations: University of Warwick

Abstract

Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, commonly used divergences may become ineffective when these distributions overlap weakly, which is frequently encountered in imitation learning, offline RL, and deployment under distribution shift. Optimal transport (OT) offers an alternative by measuring the cost of \emph{moving} probability mass from one distribution to another under a ground cost that encodes task geometry. This survey covers how OT is used inside RL objectives and algorithms. For each method, we identify: the role OT plays, the distributions compared, the OT formulation used, and the treatment of temporal structure. Beyond categorising existing methods, we discuss the motivations behind different OT choices, practical considerations such as cost design and computational challenges, and highlight open problems including scalable trajectory-level transport, principled handling of mass mismatch, and theoretical analysis for OT-regularised RL.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reliable Modeling of Distribution Shifts via Displacement-Reshaped Optimal Transport

    May 6, 2026Philip Naumann, Jacob Kauffmann, Klaus-Robert Müller +1Distribution ShiftsTransport

  2. Sliced-Regularized Optimal Transport

    Apr 27, 2026Khai NguyenGradient

  3. Simultaneous Neural Optimal Transport

    Sep 29, 2026Milena Gazdieva, Kirill Sokolov, Jiawei Chen +2Optimal Transport ApproachImage Restoration