Organizations: Robotics Department, University of Michigan, Ann Arbor, MI 48109, USA · Contextual Robotics Institute, University of California San Diego, La Jolla, CA 92093, USA. · Toyota Motor North America, Research & Development, Ann Arbor, MI 48105, USA.
Model Predictive Path Integral (MPPI) control provides a sampling-based framework for optimal control, while Control Barrier Functions (CBFs) provide a principled means of enforcing safety constraints. We introduce BR-MPPI, which integrates CBF-like conditions into MPPI's control sampling procedure. CBFs impose inequality constraints that bound the rate of change of barrier functions using a class-K function of the barrier value. We instead impose the CBF condition as an equality constraint using a parametric linear class-K function and augment the system state with its parameter. The parameter's time derivative serves as an additional control input optimized by MPPI. We further design a cost function that promotes parameter values consistent with Nagumo's condition at the safe-set boundary, thereby encouraging safety. The resulting multiple state- and control-dependent equality constraints pose a challenge for random control sampling. We address this through state transformations and control projections inspired by manifold path planning that map sampled controls onto the constraint manifold. We also incorporate learned signed distance fields to represent robot geometry and reduce computation time. Simulations demonstrate improved sample efficiency over vanilla MPPI and higher navigation success rates across five robot models compared with safety-oriented MPPI variants. Hardware experiments on a quadrotor further demonstrate the method's ability to navigate constrained environments near safe-set boundaries.
Figures & tables
Fig. 2: (a),(b) Dynamic unicycle and (c), (d) Hexagonal single-integrator robots navigating an obstacle environment with BR-MPPI and MPPI. T denotes the minimum of two values: the time required to reach within 2 meters of the goal (which varies across simulations) and the total simulation duration.
Fig. 3: Representative trials comparing BR-MPPI with MPPI-CBF, Shield-MPPI, SC-MPPI, and GS-MPPI across five robot models: (a) single integrator, (b) unicycle, (c) dynamic unicycle, (d) planar quadrotor, and (e) mobile arm. Each panel shows trajectories with robot footprints and the corresponding body-point clearance histories. Red exclamation marks in the trajectory plots and red crosses in the clearance plots indicate collisions.
Analytic SDF
Optimized neural SDF
Robot model
R / C / T
Success
Avg. Comp. Time (ms)
R / C / T
Success
Avg. Comp. Time (ms)
Speedup
Unicycle
87 / 0 / 13
87%
38.82
84 / 11 / 5
84%
18.46
2.10 ×
Dynamic unicycle
100 / 0 / 0
100%
47.76
100 / 0 / 0
100%
18.97
2.52 ×
Planar quadrotor
98 / 0 / 2
98%
52.23
94 / 0 / 6
94%
23.48
2.22 ×
TABLE I: Analytic versus optimized neural SDF barriers in BR-MPPI over 100 trials per robot model. R/C/T denotes reached/collision/timeout; deadlocks are included in timeouts. Speedup is the ratio of analytic to neural latency.
Fig. 5: (a) Experimental setup showing the obstacle course and our custom-assembled quadrotor. (b) Visualization of a hardware run with BR-MPPI. Video available on our project page.
Sampling-based model predictive control methods, such as Model Predictive Path Integral (MPPI), offer derivative-free optimization and robustness in complex robotic systems. However, standard MPPI relies on cost-based soft penalties that cannot guarantee hard-constraint satisfaction, severely limiting its applicability to highly constrained tasks such as closed-chain manipulation. To address this, we propose Manifold-Constrained MPPI (MC-MPPI), a real-time sampling-based control framework that regulates nonlinear equality constraints while preserving the computational advantages of MPPI. The key idea is to decouple the constrained optimal control problem into latent-space planning and execution-level correction. At the planning stage, a Variational Autoencoder (VAE) learns a low-dimensional latent representation of the constraint manifold, enabling MPPI to efficiently generate near-feasible candidate trajectories without per-sample modification. Since this reference enables accurate linearization of the equality constraints, an execution-level Quadratic Programming (QP) controller resolves the residual manifold mismatch in a single solve rather than through iterative projection. Experiments on a 14-DoF closed-chain dual-arm system in both simulation and real-world settings demonstrate 100 Hz planning and 500 Hz execution, equality-residual regulation in static and dynamic environments, and a 95% success rate over 40 randomized hardware trials. Supplementary videos and implementation details are available at https://rcilab.github.io/mcmppi.
Seulchan Lee, Sanghyun Kim
School of Mechanical Engineering, Kyung Hee University, Yongin, South Korea
Model Predictive Path Integral (MPPI) control is widely used in manipulation for its gradient-free, parallel handling of non-convex costs. Manipulation tasks, however, often impose constraints that hold throughout the motion: a closed kinematic chain that two grasping arms keep exactly, or joint limits and obstacle clearances that are never crossed. MPPI handles such constraints only through the cost, as soft penalties that hold approximately and fail under a strong task cost. To address this, we propose Projection-Retraction MPPI (PR-MPPI), which enforces the constraints inside the sampled dynamics. At every rollout step, the sampled velocity is projected to satisfy both constraint types: the equality restricts it to a subspace, and each inequality to a half-space within that subspace, so inequality handling never breaks the equality. This projection, however, satisfies the constraints only to first order, and a finite step leaves a small drift off the equality. Therefore, we retract the returned command back onto the constraint to numerical tolerance and independent of task weighting. We validate PR-MPPI on 14-DoF dual-arm systems. In simulation, the returned commands satisfy the closed-chain equality to numerical tolerance through a joint-limit stress test and randomized obstacle avoidance. On real hardware, the arms of a Unitree H1-2 humanoid reactively avoid a moving obstacle. Code and experiment videos are available at https://rcilab.github.io/prmppi.
Seulchan Lee, Leesai Park, Minhyeong Kang +1
School of Mechanical Engineering, Kyung Hee University, Yongin-si 17104, Republic of Korea · Institutes of Convergence Technology, Suwon-si 16229, Republic of Korea
This paper proposes a novel method that extends the Model Predictive Path Integral (MPPI) method with motion primitives for additional structured sampling, which enhances the convergence towards a globally optimal solution. By evaluating motion primitives and perturbed control sequences in a real-time sampling-based optimization loop, this work addresses the limitations of the path planning capabilities of sampling-based controllers. The algorithm is implemented on a quadcopter simulator and tested on an obstacle field navigation task. It is demonstrated that the proposed approach enhances exploration of the control space while maintaining the fast, reactive behavior required for real-time control.
Marlon Mathisen, Aksel Vaaler, Olav Egeland +1
*Norwegian University of Science and Technology, Dept. of Mechanical and Industrial Engineering, Norway