RL-Guided PAC-NMPC for Probabilistically-Safe Perception-Based Navigation in Unknown Environments
Authors: Adam Polevoy, Dillon Capalongo, Katherine Tang, Mark Gonzales, Marin Kobilarov, Joseph Moore
Organizations: Johns Hopkins University Applied Physics Laboratory, Laurel, MD 20723, USA. · Department of Mechanical Engineering, Johns Hopkins University, Baltimore, MD 21218, USA.
In this paper, we present an approach for combining stochastic nonlinear model predictive control (SNMPC) and reinforcement learning (RL) to enable probabilistically-safe perception-based navigation in unknown environments. Our method first uses RL to train probabilistic actor-critic and sensor prediction models. We then leverage these probabilistic models in a sampling-based SNMPC framework known as Probably Approximately Correct (PAC)-NMPC, which uses hard constraints to enforce finite-time statistical guarantees on the probability of collision and value function improvement. By ensuring that our finite-horizon SNMPC policies decrease the value function in expectation, we can approach the long-horizon performance of the RL approach while satisfying probabilistic safety constraints. Through simulation experiments, we show that our approach can improve the safety of perception-based RL navigation policies and scale to high dimensional systems with large sensor input spaces and complex nonlinear dynamics. We also demonstrate our approach through hardware experiments, showing improved performance for vision-based navigation with an agile fixed-wing aerial vehicle in unknown environments.
Figures & tables
Fig. 1 : Actor-Critic PAC-NMPC for navigation and obstacle avoidance on a fixed-wing with an on-board depth camera.
Approach
Stochastic MPC
Stochastic Value Func.
Safety Guarantees
Perception (max dim)
Hardware Demo
DMPC [ 38 ] , CACTO [ 39 ]
✗
✗
✗
✗
✗
PETS [ 35 ] , MQP [ 37 ]
✓
✗
✗
✗
✗
DiffStack [ 44 ]
✗
✗
✗
✓ (5 × a) ∗
✗
AC-MPC [ 46 ]
✗
✗
✗
✓ (12 × 2)
✓
POLO [ 36 ]
✓
✓
✗
✗
✗
LOOP [ 42 ] , TD-MPC [ 43 ]
✓
✗
✗
✓ (16, 84 × 84 × 9)
✗
TABLE I : Comparison of MPC & RL hybrid approaches.
Fig. 2 : Overview of AC-PAC-NMPC.
Figure 4
Approach
Success
Stuck
Violation
RL Actor Network
90%
3%
7%
MPPI w/ Learned Vϕψ
70%
7%
23%
PAC-NMPC
Quad. Term. Cost
76%
24%
0%
Map & A*
89%
4%
7%
Learned Vϕψ
97%
3%
0%
TABLE II : Simulated cluttered environments results.
Fig. 5 : Cluttered and trap environment examples. Slices of the learned value functions are visualized as heatmaps.
Fig. 9 : Histogram of fixed-wing aerial vehicle model acceleration errors. Fitted Gaussian distribution shown in orange.
Fig. 10 : Image prediction on decimated RealSense D450 depth image collected during fixed-wing flight. Depth is visualized as blue (near) to yellow (far).
Model
FLOPs (M)
δ1
δ2
δ3
REL
RMSE
Geometric (Nearest)
N/A
0.133
0.217
0.306
2.838
0.634
Geometric (Splat)
N/A
0.186
0.309
0.419
2.240
0.548
Fully Connected
36.798
0.848
0.935
0.963
0.170
0.081
U-Net
42.283
0.923
0.965
0.983
0.087
0.058
TABLE V : Performance of depth image prediction models on simulated data. δi is the percentage of pixels satisfying max(dd^,d^d)<1.25i where d^ and d are predicted and ground-truth depths. REL and RMSE are the mean absolute relative error and root mean squared error respectively.
Fig. 11 : Simulation environment example. Ceiling/floor not shown. RL actor in yellow, PAC-NMPC with global planner in red, AC-PAC-NMPC in blue.
Approach
Success
Cost
RL Actor Network
81%
468.1
PAC-NMPC
Depth Image History
52%
757.7
Map
63%
810.2
Map & A*
74%
738.8
Map & A* Unknown
85%
718.14
Actor-Critic (ours)
90%
513.3
TABLE VI : Baseline Comparison Results
TABLE VII : Ablation Study Results
Fig. 12 : Out-of-distribution simulation environment example. Ceiling/floor not shown. RL actor (collision) in yellow, PAC-NMPC with global planner in red, AC-PAC-NMPC in blue.
Approach
Success
Cost
RL Actor Network
46%
481.8
PAC-NMPC
Map & A* Unknown
76%
732.6
Actor-Critic (ours)
76%
512.1
TABLE VIII : Out Of Distribution Experiment Results
Fig. 13 : Example hardware environments with time-lapse trajectories of the aerial vehicle controlled by our approach.
Fig. 14 : Optimized PAC bounds compared to Monte Carlo estimates at each planning interval for one trial.
Fig. 15 : Spurious noise in consecutive depth camera measurements, which caused a spike in Cα+ .
Approach
Success
Cost
RL Actor Network
40%
403.8
PAC-NMPC
Map & A* Unknown
40%
325.4
Actor-Critic (ours)
80%
252.2
TABLE IX : Fixed-wing aerial vehicle hardware results.
School of Electrical Engineering, Korea Advanced Institute of Science and Technology, Daejeon, Korea · The SYCAMORE Lab, ´Ecole Polytechnique F´ed´erale de Lausanne (EPFL), Switzerland
Chalmers University of Technology, Chalmersgatan 4, 412 96 Göteborg, Sweden · Volvo Group Trucks Technology, Gropegårdsgatan 2, 417 15 Göteborg, Sweden