A Survey on Reinforcement Learning Applications in SLAM
Organizations: Computer Science and Engineering, University of North Texas, Denton, USA · Institute of Artificial Intelligence, University of Bremen, Bremen, Germany · Computer Science and Engineering, University of California Santa Cruz, Santa Cruz, USA · Electrical and Computer Engineering, University of Maine, Maine, USA
Abstract
Simultaneous localization and mapping (SLAM) allows a mobile robot or autonomous vehicle to build a map of an unknown environment while estimating its own pose within that map. Reinforcement learning (RL), in which an agent learns a decision policy from interaction and reward, has been applied to decide how such systems move, explore, and recognize places they have visited before. This survey reviews the applications of RL in SLAM. We first distinguish passive SLAM, in which the robot's motion is not chosen by the SLAM system, from active SLAM, in which it is, and summarize the sensors that provide the input to SLAM. We then introduce the RL methods used in this literature, from value-based and policy-based methods to actor-critic and deep RL. Next, we classify RL applications in SLAM into three categories: path planning, including environment exploration and obstacle avoidance; loop closure detection; and active SLAM. Thirteen representative studies are compared in terms of their simulation environment, deep learning method, SLAM method, and RL algorithm. Most of these studies are evaluated mainly in simulation, and value-based methods from the deep Q-network family are the most common. Finally, we discuss the challenges of applying RL to SLAM, namely computational demands, safety, generalization, high-dimensional state and action spaces, sample efficiency, and sensor and actuator delays, and we outline directions for future research.
Figures & tables
| Study | Year | Simulation environment | DL method | SLAM method | RL algorithm | Advantages | Disadvantages |
| Path planning | |||||||
| Wang [ 79 ] | 2024 | 2D grid map (RL); TUM RGB-D and recorded maze (SLAM) | Deep neural network | ORB-SLAM3 | Q-learning, DQN, SARSA | • Effective map conversion • Contribution to autonomous navigation | • Global planning only • Discrete actions • Simulation only |
| Nam et al. [ 80 ] | 2023 | Gazebo, ROS 2, TurtleBot3 | Deep neural network | SLAM + MCL | DQN, PPO, TD3 | • Comprehensive framework • Comparison of three DRL algorithms | • Limited generalization • Potential overfitting |
| Environment exploration | |||||||
| Chen et al. [ 78 ] | 2024 | Gazebo, ROS | CNN | SLAM (package not named) | DQN | • Effective training strategy • Addresses real-world challenges | • Slower exploration • No real-world implementation |
| Li et al. [ 81 ] | 2020 | ROS Stage; real robot (RoboMaster) | CNN | Karto SLAM (GMapping in experiments) | DQN (dueling FCQN) | • Modular framework • Efficient exploration strategy • Generalization performance | • Discrete action space |