Optimal Transport Meets Reinforcement Learning: A Survey
Organizations: University of Warwick
Abstract
Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, commonly used divergences may become ineffective when these distributions overlap weakly, which is frequently encountered in imitation learning, offline RL, and deployment under distribution shift. Optimal transport (OT) offers an alternative by measuring the cost of \emph{moving} probability mass from one distribution to another under a ground cost that encodes task geometry. This survey covers how OT is used inside RL objectives and algorithms. For each method, we identify: the role OT plays, the distributions compared, the OT formulation used, and the treatment of temporal structure. Beyond categorising existing methods, we discuss the motivations behind different OT choices, practical considerations such as cost design and computational challenges, and highlight open problems including scalable trajectory-level transport, principled handling of mass mismatch, and theoretical analysis for OT-regularised RL.
Figures & tables
| Application | Start with | Distributions | Cost default | Key pitfall |
| Online imitation | (Sinkhorn div.) | occupancies vs | learned features | sensitive to cost scale |
| Offline reward labelling | entropic OT + mask | trajectory segments | cosine on encoder | spurious matches without temporal mask |
| Policy trust region | per-state Sinkhorn | vs | Euclidean (actions) | radius too tight no improvement |
| Offline constraint | penalty | vs | action-space metric | overly strong clones |
| World model | or | vs | latent features | optimising prediction, not control |
| Robustness | ball | transitions around | env. geometry | uncalibrated over-conservatism |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| MDP and RL | |
|---|---|
| State and action spaces. | |
| Discounted Markov decision process (MDP). | |
| Discount factor. | |
| Initial-state distribution (default ; when varying the start distribution). | |
| Environment transition kernel. | |
| Trajectory law generated by . | |
| Method (area) | OT role | Compared objects | OT form | Implementation / notes |
|---|---|---|---|---|
| Adversarial occupancy matching and discriminator-based IL | ||||
| WAIL ( Xiao et al., 2019 ) , WDAIL ( Zhang et al., 2020 ) ( generative adversarial IL ) | Discriminator / reward shaping | State-action occupancy measures | (Euclidean cost) | Kantorovich duality with 1-Lipschitz critic (WGAN-style training). |
| W-IRL ( Wu et al., 2023 ) ( IRL ) | Discriminator / reward shaping | State marginals | (Euclidean cost) | Kantorovich duality with a Lipschitz-constrained critic. |
| LWAIL ( Yang et al., 2026 ) ( adversarial IL; twin delayed DDPG (TD3) ) | Latent discriminator + reward | Latent state-pair (transition) occupancies (ICVF-pretrained features) | (Euclidean cost) | Duality-based critic in latent space; pretraining via ICVF as proposed. |
| SIL ( Papagiannis and Li, 2022 ) ( adversarial IL ) | Discriminator / reward shaping via primal OT | State-action occupancy measures | Entropy-regularised (OT-GAN-style cost ( Salimans et al., 2018 ) ) | Sinkhorn iterations in minibatches; entropic OT objective. |
| PWIL ( Dadashi et al., 2021a ) ( IL ) | Reward-from-transport (non-adversarial) | State-action occupancy measures | Upper bound on | Uses approximate primal coupling (e.g. POT ( Flamary et al., 2024 ) ); transport cost converted to rewards. |
| Method (setting) | OT role | Compared objects | OT form | Implementation / notes |
|---|---|---|---|---|
| Policy update constraints and regularisers | ||||
| WPO, SPO ( Song et al., 2023 ) ( on-policy; discrete ) | Trust-region constraint | Per-state action distributions vs. | / Sinkhorn | Per-state Lagrangian dual; closed-form distribution update under task-dependent action costs (e.g. – , ). |
| OT-TRPO ( Terpin et al., 2022 ) ( TRPO-style ) | Trust-region constraint | Action distributions (e.g. Gaussians) | Closed-form for Gaussians under quadratic costs; dual optimisation over the multiplier. | |
| BGPG, BGES ( Pacchiano et al., 2020 ) ( PG / ES ) | Behaviour regulariser | Embedding distributions vs. | Entropy-regularised | Dual OT test functions on embedding space; OT signal used as shaping/weights in PG or ES updates. |
| WPR ( Na et al., 2026 ) ( RLHF ) | Policy regularisation | learned policy vs. reference policy | Entropy-regularised | Dual formulation with Sinkhorn-Knopp iterations yields a tractable semantic regulariser compatible with PPO-style updates. |
| Wasserstein geometry: gradient flows and natural gradients | ||||
| Method (setting) | OT role | Compared objects | OT form | Implementation / notes |
|---|---|---|---|---|
| Value-awareness | ||||
| VAML ( Farahmand et al., 2017 ; Asadi et al., 2018 ) ( model-based ) | Value-aware transition model loss | vs. | Equivalence between Lipschitz test-function losses and a discrepancy on next-state distributions. | |
| Link to bisimulation metric | ||||
| ( Calo et al., 2024 ) (dynamic programming) | Efficient OT distances for Markov processes | Stationary Markov processes (occupancy of state pairs) | Discounted | Sinkhorn value iteration; Markov-structured couplings improve computation and convergence. |
| Latent MDP training | ||||
| WAE-MDP ( Delgrange et al., 2023a ) ( latent world model ) | OT-based model learning | Trace/trajectory distributions (encoded/decoded) | -type | WAE/WGAN-style objective; dual-critic realisation with gradient penalty in practice. |
| Method (area) | OT role | Compared objects | OT form | Implementation / notes |
|---|---|---|---|---|
| BRAC ( Wu et al., 2019 ) ( offline RL ) | Policy regularisation / value penalty | Action distributions vs. | Dual form with a discriminator (WGAN-style critic). | |
| Q-DOT ( Omura et al., 2025 ) ( offline RL ) | Policy regularisation via transport map | Action distributions (state-conditional) | ICNN parameterisation of a convex potential / transport map. | |
| VGF ( Xu et al., 2026 ) ( offline RL ) | Gradient flow | Action distributions (state-conditional) | JKO scheme, realised by an SVGD-style RKHS-restricted particle update rather than exact JKO minimisation; implicit behaviour control via a transport budget. | |
| PPL ( Asadulaev et al., 2024 ) ( offline RL ) | Partial behaviour cloning / selective matching | Action distributions under partial OT constraint | Partial OT | Maximin formulation with a tunable partiality parameter. |
| Givchi et al. (2021) ( RL (policy optimisation) ) | Unbalanced OT formulation | Projected state and action marginals of occupancies | Unbalanced OT | Multiple convex constraints handled via Dykstra’s algorithm; Bregman (KL) divergences used in practice. |
| Fatras et al. (2021) ( Domain adaptation ) | Domain-invariant representation learning | Source vs. target distributions | Unbalanced minibatch OT | POT and Geomloss packages |
| Method (area) | OT role | Compared objects | OT form | Implementation / notes |
|---|---|---|---|---|
| Distributionally robust control | ||||
| WR 2 L ( Abdullah et al., 2019 ) ( robust RL ) | Minimax constraint | Transition kernels within a Wasserstein ball | Approximates via a Hessian-of-expectation surrogate; zero-order optimisation as proposed. | |
| DRMDP ( Kordabad et al., 2022 ; Yang, 2020 ; Hou et al., 2020 ; Queeney et al., 2024 ; Yang, 2017 ) ( robust control / robust RL ) | Ambiguity set (Wasserstein ball) | Disturbance or transition distributions | OT-cost uncertainty set (generalises ) | Kantorovich duality yields tractable robust Bellman operators / reformulations. |
| FOM for DRMDP ( Clement and Kroer, 2021 ) ( robust control / robust RL ) | Ambiguity set (Wasserstein ball) | Disturbance or transition distributions | First-order methods. | |
| Safety | ||||
| ( Baheri, 2023a ; Shahrooei and Baheri, 2024 ) ( safe RL ) | Visitation shaping / regularisation | Risk distribution vs. state visitation distribution | Wasserstein with squared Euclidean cost | OT-derived costs used as penalties in value-based updates (as proposed). |
| Method (area) | OT role | Compared objects | OT form | Implementation / notes |
|---|---|---|---|---|
| Diversity and hierarchical RL | ||||
| WURL ( He et al., 2022 ) ( unsupervised RL ) | Intrinsic reward for diversity | State marginal distributions across policies/skills | Dual form and projected Wasserstein discrepancy amortise rewards from OT plans. | |
| PWSEP ( Yang et al., 2025 ) ( unsupervised RL ) | Skill separation / discovery | State distributions conditioned on skills | Wasserstein (sliced approximation) | Uses sliced Wasserstein distance for efficient estimation of separability objective. |
| WQDIL ( Yu et al., 2025 ) ( quality-diversity IL ) | Stable reward learning in latent space | Latent distributions of demonstrations vs. policy trajectories | WAE-WGAN | Duality-based adversarial training within WAE latent space (as proposed). |
| WDER ( Li et al., 2023 ) ( hierarchical RL ) | Diversity regulariser for subpolicies | Behavioural embeddings / action distributions | Smoothed with random-feature approximation | Duality with random-feature approximations enabling scalable updates. |
| HiPBOT ( Le et al., 2023 ) ( hierarchical control ) | Policy blending (mixture weights) | Expert-policy and agent-policy priors (temperature-weighted) | Unbalanced entropic OT | Sinkhorn-like scaling algorithm ( Chizat et al., 2018 ) . |
| Study (role) | Evid. | Environment & setting | OT form & implementation | Within-study comparison & outcome | Limitation noted |
|---|---|---|---|---|---|
| OTR ( Luo et al., 2023b ) (reward labelling) | E | D4RL locomotion, AntMaze and Adroit; offline, reward-free, one expert demonstration | Entropic OT between learner/expert state marginals; Sinkhorn; cosine cost on raw states | OTR+IQL against oracle-reward IQL and prior reward-learning baselines on D4RL locomotion (HalfCheetah, Hopper, Walker2d): near-oracle aggregate performance from one demonstration, avoiding the severe drops seen in some alternatives; not equal or superior on every task | Labels are trajectory-context dependent; state-marginal OT ignores ordering |
| Dong et al. (2026) (critique) | E, T | 32 benchmarks, offline and online (D4RL locomotion, AntMaze, Adroit); four learners: IQL, ReBRAC, TD3+BC, DrQ-v2 | Ablates two axes: Wasserstein vs. a point-to-set distance (MinDist), and with vs. without temporal constraints (SegMatch) | Negative/mixed : offline, the gains from Wasserstein approximation and temporal constraints largely vanish against a well-tuned learner; online, temporal constraints are essential while the proximity approximation is secondary | Temporal alignment does matter offline as the number of demonstrations grows |
| SinkhornDRL ( Sun et al., 2024 ) (critic loss) | E, T | 55 Atari 2600 games, online, fixed model capacity | Sinkhorn divergence between particle-based return distributions | Mixed : against QR-DQN and MMD-DQN, slower convergence early but better aggregate human-normalised score; largest gains with larger action spaces and multi-dimensional rewards | 20% more compute than MMD-DQN (a different overhead is reported against C51 and QR-DQN); sensitive to , iteration count and particle number |
| TemporalOT ( Fu et al., 2024 ) (reward labelling) | E | Nine Meta-world image-based manipulation tasks; online imitation from expert video, DrQ-v2 backbone | Masked (banded) entropic OT with context-window cosine cost on encoder features | Against an online OTR variant, ADS, GAIfO, BC and an oracle task reward: higher success rate without task rewards; an ablation shows both the context cost and the mask contribute | Requires online interaction; a temporally informed proxy rather than a full dynamic-OT formulation |
| SMMOTIL ( Sebag et al., 2023 ) (reward labelling) | E | Pendulum-v0 and CartPole-v0, with five experts differing in pole length or mass | Sliced multi-marginal OT; closed-form 1D solvers | Against concatenating the demonstrations into a single empirical measure: higher mean episodic reward under both length and mass diversity | Two low-dimensional control tasks only; slicing approximates the geometry |
| WPO / SPO ( Song et al., 2023 ) (trust region) | E, T | Tabular MDPs; continuous actions via an IQN target | Per-state / Sinkhorn trust region; 1D dual solve, no parametric form assumed | Against KL trust regions: monotonic improvement with exact advantages, and global convergence in the tabular setting under a decaying multiplier | Tabular guarantee needs finite spaces, non-negative rewards and full initial-state support; the bound degrades by under advantage error |