Pow3R-SLAM: Real-Time RGB-D SLAM with 3D Reconstruction Priors
Authors: Christopher Kolios, Ishaan Mehta, Sasa Janjic, Yeganeh Bahoo, Sajad Saeedi
Organizations: Toronto Metropolitan University, Toronto, Canada. · University of Windsor, Windsor, Canada. · University College London, London, United Kingdom.
We present Pow3R-SLAM, a real-time RGB-D simultaneous localization and mapping (SLAM) system that uses Pow3R for tracking and mapping. Inspired by MASt3R-SLAM, a recent work on monocular SLAM using two-view 3D reconstruction priors, we extend the work to incorporate depth as a prior on the network's prediction, rather than as geometry to fuse. Where traditional RGB-D SLAM systems struggle with sparsity in the depth images, Pow3R utilizes the available depth to give a better-conditioned pointmap, while inferring the depths in empty regions from the two-view photometric, depth, and intrinsic data. Evaluated against MASt3R-SLAM following its protocol on 24 sequences from TUM, 7-Scenes, and Replica, Pow3R-SLAM runs 1.6x faster in wall time, has 15% lower mean trajectory error, a 3.1x lower unscaled error, and produces denser maps, with a 30% lower Chamfer distance. We also introduce a hybrid variant that runs 2.1x faster than MASt3R-SLAM at 25.3 frames per second (FPS), while maintaining improved tracking and mapping accuracy. Against ORB-SLAM3 in RGB-D mode, Pow3R-SLAM is more accurate on TUM, 7-Scenes, and ETH3D-SLAM, and completes every TUM sequence. While Pow3R-SLAM can struggle on a small set of self-similar scenes, its overall performance shows that adding depth as a prior for two-view 3D reconstruction SLAM can be beneficial. A project webpage is available at: https://ChrisKolios.github.io/Pow3R-SLAM , and code will be made open-source upon acceptance.
Figures & tables
Fig. 1: Sensor depth vs. Pow3R output . Given input frames ((a), (b)), with corresponding depth maps ((c), (d)) from the TUM fr1/desk scene [ 3 ] , (e) shows the pointmap result from the back-projected depthmaps, and (f) shows the pointmap result from passing the input RGB-D into Pow3R’s network. Pow3R maintains flat surfaces, and reconstructs the controller where the Kinect sensor lacks the data to do so (fills holes and regularizes surfaces).
Fig. 2: Pow3R-SLAM pipeline. Pow3R-SLAM is an RGB-D SLAM system that uses sensor depth as a prior on a two-view pointmap network. The front end (1-3) runs on every frame. (1) Pow3R predicts pointmaps and confidences for the frame and its keyframe, conditioned on sensor depth and intrinsics (Sec. III-B ). (2) Each prediction is made metric against the sensor depth (Eq. 1 ). (3) The frame is tracked in Sim(3) and may become a keyframe (Secs. III-C – III-D ). The backend (4, 5) runs on every new keyframe. (4) Loop closure on Pow3R’s encoder tokens adds edges in familiar locations, and relocalization locates frames the tracker has lost. (5) A global optimization refines every keyframe pose, with keyframe depths anchored to the sensor and confidence-calibrated edges (Sec. III-E3 , Eqs. 5 – 7 ). Blue : the sensor-depth path. Orange : optional components, including the hybrid variant, which tracks alternate frames by ICP instead of Pow3R (Sec. III-G ), and map-only keyframes (Sec. III-F ). ICP frames never become keyframes, so every keyframe comes from a Pow3R pass. Inputs and outputs: TUM fr1/desk .
ORB-SLAM3
Pow3R-SLAM
scene
DROID † ↓
Sim3 ↓
SE3 * ↓
MASt3R fresh ↓
Sim3 ↓
SE3 * ↓
hyb. Sim3 ↓
TUM RGB-D fr1
360
0.111
0.134
0.228
0.049
0.040
0.048
0.042
desk
0.018
0.018
0.018
0.016
0.019
0.035
0.017
desk2
0.042
X
X
0.024
0.024
0.093
0.022
floor
0.021
X
X
0.025
0.021
0.025
0.025
TABLE I: ATE root-mean square error (RMSE) (m) on TUM RGB-D fr1 and 7-Scenes . Bold , underline : best, second best. † : calibrated monocular result as reported in [ 7 ] . Fresh: our re-run. X: lost track, with means over completed scenes. ∗ : unscaled SE(3) alignment.
scene / metric
MASt3R fresh
Pow3R- SLAM
Pow3R- SLAM hyb.
office0
0.013
0.009
0.008
office1
0.007
0.009
0.013
office2
0.022
0.007
0.007
office3
0.027
0.010
0.012
office4
0.021
0.009
0.011
room0
0.006
0.007
0.007
TABLE II: Replica, ETH3D-SLAM and runtime. Keyframe ATE RMSE (m), Sim(3), with the unscaled SE(3) mean. ETH3D-SLAM: over each system’s completed sequences (count in brackets). Runtime: wall time summed over the 24-scene panel (TUM, 7-Scenes, Replica), at full and subsample-2 (s2). Mean of 2 exclusive repeats per system ( ± half-range), with MASt3R-SLAM timed on native images. FPS: frames per second at s2 of loop time excluding model load.
Fig. 3: Qualitative reconstruction. Rows: Replica office3 , 7-Scenes chess , ETH3D-SLAM mannequin_5 (full frame rate). Columns: (a) reference cloud from ground-truth depth, (b) MASt3R-SLAM and (c) Pow3R-SLAM maps after alignment, coloured by distance to the reference (capped at 10 cm), and (d) top view of the per-frame trajectories (ground truth (GT) black, MASt3R-SLAM orange, Pow3R-SLAM blue).
Fig. 4: The failure case . 7-Scenes stairs , with the same columns as Fig. 3 , (d) GT / M(ASt3R) / P(ow3R). MASt3R- vs. Pow3R-SLAM: 1.5 vs. 4.3 M points, keyframe ATE 0.016 vs. 0.044 m, accuracy 3.2 vs. 6.7 cm, completion 10.5 vs. 5.9 cm. Pow3R-SLAM’s map is more complete but less accurate, which we attribute to the self-similar treads (Sec. V ).
dataset
system
ATE ↓
Acc. ↓
Comp. ↓
Cham. ↓
Ch. RMSE ↓
7-Scenes
MASt3R-SLAM fresh
0.048
0.051
0.058
0.055
0.085
Pow3R-SLAM
0.044
0.048
0.038
0.043
0.058
Pow3R-SLAM hyb.
0.048
0.050
0.038
0.044
0.058
Replica
MASt3R-SLAM fresh
0.015
0.037
0.031
0.034
0.056
Pow3R-SLAM
0.009
0.022
0.018
0.020
0.027
Pow3R-SLAM hyb.
0.010
0.023
0.019
0.021
0.028
TABLE III: Reconstruction on 7-Scenes and Replica. ATE for reference. Mean accuracy (Acc.), completion (Comp.), Chamfer (Cham.), and RMSE Chamfer, in m.
configuration
Sim3 mean (24) ↓
Sim3 median (24) ↓
SE3 mean (24) ↓
Chamfer (15) ↓
stairs ↓
ORB-SLAM3 (ref., 21)
0.0342
0.0178
0.0411
–
0.045
MASt3R fresh (ref.)
0.0304
0.0228
0.1104
4.35
0.016
L1 Pow3R, no priors
0.0834
0.0449
0.4094
9.55
0.063
L2 + intrinsics
0.0711
0.0431
0.4058
8.70
0.052
L3 + cond
0.0319
0.0189
0.3640
6.58
0.115
L4 + scale
0.0260
0.0191
0.0337
3.07
0.044
TABLE IV: Ablations . Means and median ATE over the 24 scenes, Chamfer (cm) over the 15 map scenes, with stairs separate. Rows L1–L5 add one component at a time up to the shipped default (def.); L6 and L7 leave out depth sites (23 scenes, as fr1/rpy diverges without cond). ref.: reference systems, not ranked (ORB-SLAM3 over its 21 completions). KF: keyframes.
We present AMB3R-SLAM, a real-time monocular SLAM system capable of reconstructing kilometer-scale trajectories over 10k frames on a single consumer-grade GPU. Our model couples a lightweight front-end for low-latency online tracking with a hierarchical backend that progressively enforces local, mid-level, and global consistency. By avoiding bundle adjustment that relies on the static world assumption, our system naturally handles complex dynamic scenes out of the box. Furthermore, we demonstrate that our method can be extended to leverage stereo, RGB-D, and LiDAR as additional inputs. AMB3R-SLAM achieves strong camera tracking performance across 9 datasets, reducing the absolute trajectory error (ATE) of previous state-of-the-art methods on VBR and Oxford Spires by over 70%. With additional LiDAR input, our model further reduces ATE to sub-meter level on KITTI and VBR datasets.
Hengyi Wang, Lourdes Agapito
Department of Computer Science, University College London
Outdoor robot teams need a shared dense map despite limited overlap, independent reference frames, and uncertain monocular scale. Collaborative dense SLAM systems typically resolve this with depth sensors, which add payload, power, and calibration cost. We present CoMo3R-SLAM, a collaborative monocular dense SLAM system that places learned feed-forward 3D reconstruction priors at the center of the multi-agent problem: their dense pointmaps anchor scale across agents and supply correspondences strong enough to verify inter-agent links geometrically. Each agent tracks and fuses its own keyframes from a single RGB stream, while a coordinator retrieves cross-agent keyframes over the prior's encoder features, verifies them by bidirectional dense pointmap matching, synchronizes the independent similarity gauges in closed form, and refines every keyframe in one unified multi-agent sim(3) graph. Finally, a pose-depth alternation over geometry-aware segments lets inter-agent observations constrain dense structure as well as trajectories. Requiring neither measured depth nor supplied intrinsics, CoMo3R-SLAM attains the lowest trajectory error on three of four Tanks and Temples scenes, and competitive accuracy on Waymo driving sequences, while running at approximately 6-8 FPS on RTX 3080 Ti. A long-horizon traversal, independently captured day and night streams, and teams of up to four agents further map its operating range.
Zhihao Cao, Qi Shao, Shuhao Zhai +4
ETH Zurich, Switzerland · Harbin Engineering University, China · University of Liverpool, United Kingdom +4
Simultaneous localization and mapping (SLAM) based on Neural Radiance Fields (NeRF) enables dense, continuous scene reconstruction. However, existing systems operating with limited online resources struggle to simultaneously construct two types of constraints, namely, compact yet discriminative spatial constraints derived from scene representations and persistent temporal constraints derived from historical observations. To address this challenge, we propose CHOW-SLAM, a dense RGB-D SLAM framework that explicitly constructs these complementary spatial and temporal constraints. Spatially, we propose a compact parametric-hash (P-H) hybrid representation that organizes components based on planes and grids across scales in P and H branches. A unified multi-output decoder further aligns the ray termination distributions induced by TSDF and density, preserving geometry and appearance under a compact parameter budget. Temporally, we propose a complementary overlap-window strategy to prevent optimization from being dominated by short-term overlap or weakly related historical observations. Within a fixed budget, the strategy retains recent frames, selects high-overlap local frames, and introduces temporally distributed historical keyframes. Loss-aware keyframe insertion and bundle adjustment scheduling further adapt optimization to tracking quality. In addition, ORB-based tracking and geometric pose estimation are used for pose initialization, followed by neural rendering optimization to improve tracking stability. Extensive evaluations on multiple datasets demonstrate that CHOW-SLAM outperforms state-of-the-art methods in both scene reconstruction quality and camera tracking accuracy. The source code is available at https://github.com/jinjidexiaohuoban/CHOW-SLAM.
Wenxuan Ji, Jin Xiao, Xiaoguang Hu +3
School of Automation Science and Electrical Engineering, Beihang University, Beijing, 100191, China