RL-Guided PAC-NMPC for Probabilistically-Safe Perception-Based Navigation in Unknown Environments
Authors: Adam Polevoy, Dillon Capalongo, Katherine Tang, Mark Gonzales, Marin Kobilarov, Joseph Moore
Organizations: Johns Hopkins University Applied Physics Laboratory, Laurel, MD 20723, USA. · Department of Mechanical Engineering, Johns Hopkins University, Baltimore, MD 21218, USA.
In this paper, we present an approach for combining stochastic nonlinear model predictive control (SNMPC) and reinforcement learning (RL) to enable probabilistically-safe perception-based navigation in unknown environments. Our method first uses RL to train probabilistic actor-critic and sensor prediction models. We then leverage these probabilistic models in a sampling-based SNMPC framework known as Probably Approximately Correct (PAC)-NMPC, which uses hard constraints to enforce finite-time statistical guarantees on the probability of collision and value function improvement. By ensuring that our finite-horizon SNMPC policies decrease the value function in expectation, we can approach the long-horizon performance of the RL approach while satisfying probabilistic safety constraints. Through simulation experiments, we show that our approach can improve the safety of perception-based RL navigation policies and scale to high dimensional systems with large sensor input spaces and complex nonlinear dynamics. We also demonstrate our approach through hardware experiments, showing improved performance for vision-based navigation with an agile fixed-wing aerial vehicle in unknown environments.
Figures & tables
Fig. 1 : Actor-Critic PAC-NMPC for navigation and obstacle avoidance on a fixed-wing with an on-board depth camera.
Approach
Stochastic MPC
Stochastic Value Func.
Safety Guarantees
Perception (max dim)
Hardware Demo
DMPC [ 38 ] , CACTO [ 39 ]
✗
✗
✗
✗
✗
PETS [ 35 ] , MQP [ 37 ]
✓
✗
✗
✗
✗
DiffStack [ 44 ]
✗
✗
✗
✓ (5 × a) ∗
✗
AC-MPC [ 46 ]
✗
✗
✗
✓ (12 × 2)
✓
POLO [ 36 ]
✓
✓
✗
✗
✗
LOOP [ 42 ] , TD-MPC [ 43 ]
✓
✗
✗
✓ (16, 84 × 84 × 9)
✗
TABLE I : Comparison of MPC & RL hybrid approaches.
Fig. 2 : Overview of AC-PAC-NMPC.
Figure 4
Approach
Success
Stuck
Violation
RL Actor Network
90%
3%
7%
MPPI w/ Learned Vϕψ
70%
7%
23%
PAC-NMPC
Quad. Term. Cost
76%
24%
0%
Map & A*
89%
4%
7%
Learned Vϕψ
97%
3%
0%
TABLE II : Simulated cluttered environments results.
Fig. 5 : Cluttered and trap environment examples. Slices of the learned value functions are visualized as heatmaps.
Fig. 9 : Histogram of fixed-wing aerial vehicle model acceleration errors. Fitted Gaussian distribution shown in orange.
Fig. 10 : Image prediction on decimated RealSense D450 depth image collected during fixed-wing flight. Depth is visualized as blue (near) to yellow (far).
Model
FLOPs (M)
δ1
δ2
δ3
REL
RMSE
Geometric (Nearest)
N/A
0.133
0.217
0.306
2.838
0.634
Geometric (Splat)
N/A
0.186
0.309
0.419
2.240
0.548
Fully Connected
36.798
0.848
0.935
0.963
0.170
0.081
U-Net
42.283
0.923
0.965
0.983
0.087
0.058
TABLE V : Performance of depth image prediction models on simulated data. δi is the percentage of pixels satisfying max(dd^,d^d)<1.25i where d^ and d are predicted and ground-truth depths. REL and RMSE are the mean absolute relative error and root mean squared error respectively.
Fig. 11 : Simulation environment example. Ceiling/floor not shown. RL actor in yellow, PAC-NMPC with global planner in red, AC-PAC-NMPC in blue.
Approach
Success
Cost
RL Actor Network
81%
468.1
PAC-NMPC
Depth Image History
52%
757.7
Map
63%
810.2
Map & A*
74%
738.8
Map & A* Unknown
85%
718.14
Actor-Critic (ours)
90%
513.3
TABLE VI : Baseline Comparison Results
TABLE VII : Ablation Study Results
Fig. 12 : Out-of-distribution simulation environment example. Ceiling/floor not shown. RL actor (collision) in yellow, PAC-NMPC with global planner in red, AC-PAC-NMPC in blue.
Approach
Success
Cost
RL Actor Network
46%
481.8
PAC-NMPC
Map & A* Unknown
76%
732.6
Actor-Critic (ours)
76%
512.1
TABLE VIII : Out Of Distribution Experiment Results
Fig. 13 : Example hardware environments with time-lapse trajectories of the aerial vehicle controlled by our approach.
Fig. 14 : Optimized PAC bounds compared to Monte Carlo estimates at each planning interval for one trial.
Fig. 15 : Spurious noise in consecutive depth camera measurements, which caused a spike in Cα+ .
Approach
Success
Cost
RL Actor Network
40%
403.8
PAC-NMPC
Map & A* Unknown
40%
325.4
Actor-Critic (ours)
80%
252.2
TABLE IX : Fixed-wing aerial vehicle hardware results.
Sampling-Based Model-Predictive Control (MPC) algorithms are a flexible class of controllers used for navigation on a wide range of robotic systems. Historically, such approaches have lacked hard safety guarantees, a shortcoming which we remedy in this work by computing guaranteed reachable-set overapproximations online with a fast, interval-based pipeline. We show that our method achieves similar performance to a state-of-the-art reachability-based planner without the need for the expensive pre-computation step, and can be scaled to systems that are infeasible using existing approaches. Finally, we demonstrate that our technique reduces safety violations by over 99% in a racing simulation and successfully controls a model racecar on real hardware experiments without crashes.
Safe motion planning in uncertain, time-varying environments is challenging because the safe region can change unpredictably across planning steps, often causing a loss of recursive feasibility. In this work, we present a Probabilistic Recursively Feasible Model Predictive Control (PRF-MPC) framework that guarantees recursive feasibility with a specified probability. We introduce properties that an ideal predictor should satisfy to ensure distributional consistency, and use these properties to derive closed-form expressions for the means and covariances of trajectories predicted at future time steps. Building on this analysis, we construct safety constraints that ensure, with high probability, that the current safe set is contained within the safe sets at future time steps, thereby probabilistically guaranteeing recursive feasibility. Simulation results on a lane-change scenario demonstrate that the proposed method significantly improves recursive feasibility.
Hyeontae Sung, Hyeongchan Ham, Junyoung Park +2
School of Electrical Engineering, Korea Advanced Institute of Science and Technology, Daejeon, Korea · The SYCAMORE Lab, ´Ecole Polytechnique F´ed´erale de Lausanne (EPFL), Switzerland
We investigate interactive trajectory planning subject to uncertainty in the decisions of surrounding agents. To control the ego-agent, we aim to first learn the decision distribution and solve a Stochastic Model Predictive Control (SMPC) problem. To account for errors in the learned distribution, we show that it is possible to utilize Probably Approximately Correct (PAC) learning in combination with Distributionally Robust (DR) optimization to obtain a solution which accounts for the errors induced by the learning model. The results indicate that our PAC learning-based DR-MPC framework provides a method to interpolate between a robust MPC and an omnipotent SMPC, based on the available number of samples.
Erik Börve, Nikolce Murgovski, Morteza Haghir Chehreghani +1
Chalmers University of Technology, Chalmersgatan 4, 412 96 Göteborg, Sweden · Volvo Group Trucks Technology, Gropegårdsgatan 2, 417 15 Göteborg, Sweden