Safe Ergodic Control for Multi-Robot Systems via Quadratic Programming
Authors: Yo Toyomoto, Mohamed Elobaid, Bryce L. Ferguson, Alberto Quattrini Li, Shinkyu Park
Organizations: Electrical and Computer Engineering, King Abdullah University of Science and Technology (KAUST), Thuwal 23955, Saudi Arabia · Thayer School of Engineering at Dartmouth College, Hanover, NH, USA · Department of Computer Science, Dartmouth College, Hanover, NH, USA
Ergodic control drives robots to spend time in each region in proportion to a spatial distribution of interest, making it well suited for dense spatiotemporal environmental monitoring. Existing safe ergodic controllers rely on offline trajectory optimization or hierarchical architectures, which limit real-time applicability and decouple the ergodicity objective from the safety constraint. This paper presents a quadratic programming (QP)-based controller that treats ergodicity and safety jointly. We first introduce a Gaussian-kernel ergodic metric that, unlike the classical indicator-based metric, is time differentiable. This allows the exponential decay of the metric to be imposed as a time-varying control barrier function (CBF) constraint, relaxed by a slack variable, alongside hard CBF constraints for region containment and inter-robot collision avoidance. We establish that the resulting QP remains feasible from any safe initial configuration and that, up to the slack term, the ergodic metric decays exponentially. Simulations show improved performance over baseline methods, and experiments with three aerial vehicles validate the work on real hardware.
Figures & tables
Fig. 1 : Safe ergodic control of multiple robots. Three drones move within a region Q so that the time spent along their trajectories (white curves) is distributed according to a desired distribution ϕ (color scale), while remaining inside Q and keeping a mutual distance of at least rmin (green circles of radius rmin/2 ).
Parameter
Value
Parameter
Value
σmin
0.1
σmax
2.0
α(θ)
θ
γ
0.1
w
1012umax
βb(θ)
θ
βc(θ)
0.25θ
rmin
0.3
TABLE I : Parameters used in the simulations.
Fig. 2 : Time evolution of the ergodic metric E in ( 7 ) for a single robot under a uniform target distribution scenario with the proposed controller (blue) and SMC baseline (red). Thin lines show ten trials with random initial positions; bold lines show their averages.
Fig. 3 : Time evolution of the ergodic metric E in ( 7 ) for three robots under a two-peak target distribution scenario with the proposed controller (blue) and hierarchical SMC–CBF baseline (red). Thin lines show ten trials with random initial positions; bold lines show their averages.
Fig. 4 : Robot position along x , y , and z in the two-peak simulations with ℓϕ=1.0 , 2.0 , and 3.0m (top to bottom). Gray dash-dotted lines mark the coordinates of the two peaks.
Fig. 5 : Time evolution of the ergodic metric E in ( 7 ) for a single robot under two-peak target distribution with separations ℓϕ∈{1.0,2.0,3.0}\mathrm{m}$$ (solid) and under the uniform target distribution of the first scenario (dashed).
Fig. 6 : Snapshots of the three experiments with nϕ=2 , 3 , and 4 peaks (left to right) at t=30\text{,}\mathrm{s} (top) and $t=$150\text{\,}\mathrm{s} (bottom). Target distribution is overlaid in color (red high, blue low) and drone trajectories are shown in white.
Fig. 7 : Pairwise distances between drones ∥pj−pk∥2 for nϕ=2 , 3 , and 4 (top to bottom). Legend entries indicate the pair (j,k) . The dashed line marks rmin=0.3\text{,}\mathrm{m}$$ .
Fig. 8 : Time evolution of the ergodic metric E in ( 7 ) in the experiments with nϕ=2 , 3 , and 4 peaks.
Fig. 9 : Drone trajectories in the xy -plane in the two-peak experiment for σmax=0.3 , 1.0 , and 2.0m (left to right). Crosses mark the peak locations.
This paper presents a novel density control framework for multi-robot systems with spatial safety and energy sustainability guarantees. Stochastic robot motion is encoded through the Fokker-Planck Partial Differential Equation (PDE) at the density level. Control Lyapunov and control barrier functions are integrated with PDEs to enforce target density tracking, obstacle region avoidance, and energy sufficiency over multiple charging cycles. The resulting quadratic program enables fast in-the-loop implementation that adjusts commands in real-time. Multi-robot experiment and extensive simulations were conducted to demonstrate the effectiveness of the controller under localization and motion uncertainties.
Longchen Niu, Andrew Nasif, Gennaro Notomista
Department of Electrical and Computer Engineering, University of Waterloo, Waterloo, ON, Canada
Safe robot control requires maximizing return while satisfying safety constraints. In off-policy safe reinforcement learning, reward and safety Q-values are commonly learned by separate critic ensembles, with uncertainty handled independently for each objective. This objective-wise treatment neglects inter-objective correlation and can lead to overly conservative value estimates, thereby reducing sample efficiency. To address this issue, we propose Cholesky-Ordered Projection Q-learning (COP-Q), a safety-first method that incorporates inter-objective covariance into vector-valued Q-value estimation. COP-Q constructs a generalized confidence bound in the joint Q-value space and uses Cholesky factorization to encode objective priority in a sequential form. This preserves conservatism on safety while adaptively reducing excessive conservatism on the reward objective. The resulting estimate is used in both temporal-difference target computation and actor optimization. COP-Q incurs minimal computational overhead and is readily compatible with most existing deep Q-learning frameworks. Experiments on robot locomotion in Brax and safe navigation in Safety-Gymnasium, covering both hard- and soft-safety settings, demonstrate that COP-Q achieves strong safety performance together with competitive or improved sample efficiency relative to representative baselines.
Guopeng Li, Moritz A. Zanger, Matthijs T. J. Spaan +1
Department of Cognitive Robotics, Delft University of Technology · School of Transportation, Southeast University, China · Department of Intelligent Systems, Delft University of Technology, the Netherlands
As robot capabilities increase, quadratic programming (QP)-based controllers must account for a similarly increasing number of constraints to ensure safe, reliable operation. Yet, with each added constraint, this introduces more chances of momentary conflict: in which case, a QP solver that returns an "infeasible" status leaves the controller with nothing to execute. To address this, we introduce ElastiQP, a modified dual active-set QP solver that relaxes every inequality constraint with an exact, per-constraint l1 penalty while keeping equality constraints (dynamics) hard. Notably, ElastiQP does so by folding the slack variables into the solver analytically, maintaining a constant size of the condensed linear system. On a suite of robot control benchmarks, ElastiQP achieves microsecond-level performance, matching or outperforming leading modern solvers on feasible problems. On infeasible problems, ElastiQP handles these gracefully, confining violations to strictly the conflicting inequality terms, returning a usable solution up to 40x faster than the best alternative solvers. ElastiQP is available as an open-source C++ header-only library, with Python and JAX interfaces, at https://github.com/StanfordASL/elastiqp.
Daniel Morton, Jon Arrizabalaga, Zachary Manchester +1
Departments of Mechanical Engineering and Aeronautics & Astronautics, Stanford University, USA · Department of Aeronautics and Astronautics, Massachusetts Institute of Technology, USA