eess.SYOct 8, 2026

Policy Synthesis for Finite Populations of MDP Agents under Aggregate Reach-Avoid Chance Constraints

Authors: Jie Fu, Anamika Dubey

Organizations: Department of Electrical and Computer Engineering, University of Florida, Gainesville, FL 32611, USA · School of Electrical Engineering and Computer Science, Washington State University, Pullman, WA 99164, USA

Abstract

Consider a finite population of agents with decoupled Markov transition dynamics and empirical-density feedback, subject to the following constraints: with probability at least 1−δr1-δ_r, at least a fraction αrα_r of agents must reach a target region at some time t∗t^*, while, at each time up to t∗t^*, the unsafe population fraction must remain below βuβ_u with probability at least 1−δu1-δ_u. However, standard mean-field methods enforce these constraints only in expectation, which fails to account for stochastic fluctuations at finite fleet size NN. To address this control problem, we propagate the second-order moment (variance) of the empirical density alongside the mean-field trajectory via a discrete-time Lyapunov recursion, and apply the Cantelli inequality to convert chance constraints into tractable deterministic conditions on the moments of the empirical density. We then incorporate these moment-based surrogate constraints into a gradient-based sequential convex approximation procedure for density-feedback policy synthesis. We further introduce additional moment-error bounds to construct a rigorous finite-NN certificate. The method is evaluated on a gridworld environment and a power-system EV-charging aggregation problem and compared with a standard deterministic population-level LP baseline.

Figures & tables

Explore similar work

CardsList
  1. Reachability-Certified Subteam Decomposition for Locally Interacting Multi-Agent MDPs

    Sep 8, 2026Xiangwu Wang, Chengwei Cao, Hongyuan TangMulti-Agent System OptimizationMulti-Agent Coordination

  2. Mean Field Reinforcement Learning

    Jul 1, 2026René Carmona, Mathieu LaurièreReinforcement LearningQ-Learning

  3. Randomized Transport Maps for Model-Free Policy-Gradient Mean-Field Control

    Oct 8, 2026Adonis Jamal, Samy Mekkaoui, Yadh Hafsi +1Policy Gradient MethodsModel-Free RL