A Survey on Reinforcement Learning Applications in SLAM
Authors: Mohammad Dehghani Tezerjani, Mohammad Khoshnazar, Mohammadhamed Tangestanizadeh, Arman Kiani, Qing Yang
Organizations: Computer Science and Engineering, University of North Texas, Denton, USA · Institute of Artificial Intelligence, University of Bremen, Bremen, Germany · Computer Science and Engineering, University of California Santa Cruz, Santa Cruz, USA · Electrical and Computer Engineering, University of Maine, Maine, USA
Simultaneous localization and mapping (SLAM) allows a mobile robot or autonomous vehicle to build a map of an unknown environment while estimating its own pose within that map. Reinforcement learning (RL), in which an agent learns a decision policy from interaction and reward, has been applied to decide how such systems move, explore, and recognize places they have visited before. This survey reviews the applications of RL in SLAM. We first distinguish passive SLAM, in which the robot's motion is not chosen by the SLAM system, from active SLAM, in which it is, and summarize the sensors that provide the input to SLAM. We then introduce the RL methods used in this literature, from value-based and policy-based methods to actor-critic and deep RL. Next, we classify RL applications in SLAM into three categories: path planning, including environment exploration and obstacle avoidance; loop closure detection; and active SLAM. Thirteen representative studies are compared in terms of their simulation environment, deep learning method, SLAM method, and RL algorithm. Most of these studies are evaluated mainly in simulation, and value-based methods from the deep Q-network family are the most common. Finally, we discuss the challenges of applying RL to SLAM, namely computational demands, safety, generalization, high-dimensional state and action spaces, sample efficiency, and sensor and actuator delays, and we outline directions for future research.
Figures & tables
Fig. 1: The agent–environment interaction loop in reinforcement learning.
Fig. 2: Classification of RL applications in SLAM.
Study
Year
Simulation environment
DL method
SLAM method
RL algorithm
Advantages
Disadvantages
Path planning
Wang [ 79 ]
2024
2D grid map (RL); TUM RGB-D and recorded maze (SLAM)
Deep neural network
ORB-SLAM3
Q-learning, DQN, SARSA
• Effective map conversion • Contribution to autonomous navigation
• Global planning only • Discrete actions • Simulation only
Nam et al. [ 80 ]
2023
Gazebo, ROS 2, TurtleBot3
Deep neural network
SLAM + MCL
DQN, PPO, TD3
• Comprehensive framework • Comparison of three DRL algorithms
• Limited generalization • Potential overfitting
Environment exploration
Chen et al. [ 78 ]
2024
Gazebo, ROS
CNN
SLAM (package not named)
DQN
• Effective training strategy • Addresses real-world challenges
• Slower exploration • No real-world implementation
Simultaneous localization and mapping (SLAM) is a foundational state estimation problem in robotics in which a robot accurately constructs a map of its environment while also localizing itself within this construction. We study the active SLAM problem through the lens of optimal stochastic control, thereby recasting it as a decision-making problem under partial information. After reviewing several commonly studied models, we present a general stochastic control formulation of active SLAM together with a rigorous treatment of motion, sensing, and map representation. We introduce a new exploration stage cost that encodes the geometry of the state when evaluating information-gathering actions. This formulation, constructed as a nonstandard partially observable Markov decision process (POMDP), is then analyzed to derive rigorously justified approximate solutions that are near-optimal. To enable this analysis, the associated regularity conditions are studied under general assumptions that apply to a wide range of robotics applications. For a particular case, we conduct an extensive numerical study in which standard learning algorithms are used to learn near-optimal policies.
Ilir Gusija, Fady Alajaji, Serdar Yüksel
Department of Mathematics and Statistics, Queen’s University, Kingston, ON, Canada
Simultaneous localization and mapping (SLAM) is one of the services running on an autonomous robot. It is typically run to assist other tasks such as planning, manipulation, etc. All these tasks are run on edge hardware and are subject to severe resource constraints. However, most SLAM systems are built and tested in isolation, and their performance is reported as if they are the only task running on a system. We observe that existing benchmarks lack a common mechanism for comparing SLAM systems under realistic resource constraints. To address this limitation, we have developed SLAMSqueezeBench, a framework that allows testing of SLAM systems under realistic workloads on edge hardware. It does so by imposing constraints on compute and memory resources available for the SLAM system during execution. It also simulates realistic camera frame acquisition with frame drops when a finite buffer is full. Using SLAMSqueezeBench, we compare nine SLAM systems spanning classical systems, learning-based systems, and approaches for Gaussian splatting. Our testing framework will be available for use by the community upon publication.
Visual SLAM is a well-established technology utilized in a wide range of real-world applications. However, its performance still degrades under challenging visual conditions, such as low texture, severe motion blur, and poor illumination. Systems based on deep learning outperform classical geometry-based ones and achieve state-of-the-art results by combining learned 2D data association and uncertainty with differentiable geometric optimization in recurrent architectures. Still, it remains unclear exactly which components are fundamentally responsible for this success. In this paper, we ask: Is the superior performance of deep learning-based systems driven primarily by learned 2D data association, the combination of learned 2D data association and uncertainty, or the recurrent architecture itself? We investigate this question empirically by conducting a controlled study. Our findings reveal that the success of DL-based V-SLAM systems hinges on learned 2D data association and uncertainty rather than their recurrent architecture, underscoring the necessity of learning-based paradigms for the design of these components.
Giovanni Cioffi, Davide Scaramuzza
Robotics and Perception Group, Department of Informatics, University of Zurich, Switzerland