math.OCOct 8, 2026

Randomized Transport Maps for Model-Free Policy-Gradient Mean-Field Control

Authors: Adonis Jamal, Samy Mekkaoui, Yadh Hafsi, Huyên Pham

Organizations: École Polytechnique

Abstract

We develop a model-free policy gradient method for discrete-time mean-field control (MFC). In MFC, the policy affects the objective both through the controlled dynamics and through the population distribution. Standard REINFORCE estimators capture the first effect but not the second. We introduce Transport REINFORCE, a transport map-based approach that perturbs a suitable transformation of the population distribution to estimate this missing mean-field contribution. The method applies to both finite and continuous state spaces. In finite state spaces, we perturb the population distribution directly on the probability simplex through a convex combination of the current population weights and random weights. In continuous state spaces, we project the population distribution onto the manifold of Gaussian mixtures, and then randomize it via a transport map that ensures the perturbed law remains within this manifold. We prove consistency of the perturbed objective and gradient as the perturbation vanishes, and derive bias and mean-square error bounds for the resulting sample-based gradient estimator. Numerical experiments on several MFC benchmarks show that Transport REINFORCE improves over standard REINFORCE.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies

    Jul 13, 2026Ziheng Cheng, Xin Guo, Huyên Pham +1Policy GradientContinuous-Time RL

  2. Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data

    Jun 25, 2026Erhan Bayraktar, Martin Hernandez, Qinxin Yan +1Reinforcement LearningStochastic Optimal Control

  3. Mean Field Reinforcement Learning

    Jul 1, 2026René Carmona, Mathieu LaurièreReinforcement LearningQ-Learning