Pow3R-SLAM: Real-Time RGB-D SLAM with 3D Reconstruction Priors
Authors: Christopher Kolios, Ishaan Mehta, Sasa Janjic, Yeganeh Bahoo, Sajad Saeedi
Organizations: Toronto Metropolitan University, Toronto, Canada. · University of Windsor, Windsor, Canada. · University College London, London, United Kingdom.
We present Pow3R-SLAM, a real-time RGB-D simultaneous localization and mapping (SLAM) system that uses Pow3R for tracking and mapping. Inspired by MASt3R-SLAM, a recent work on monocular SLAM using two-view 3D reconstruction priors, we extend the work to incorporate depth as a prior on the network's prediction, rather than as geometry to fuse. Where traditional RGB-D SLAM systems struggle with sparsity in the depth images, Pow3R utilizes the available depth to give a better-conditioned pointmap, while inferring the depths in empty regions from the two-view photometric, depth, and intrinsic data. Evaluated against MASt3R-SLAM following its protocol on 24 sequences from TUM, 7-Scenes, and Replica, Pow3R-SLAM runs 1.6x faster in wall time, has 15% lower mean trajectory error, a 3.1x lower unscaled error, and produces denser maps, with a 30% lower Chamfer distance. We also introduce a hybrid variant that runs 2.1x faster than MASt3R-SLAM at 25.3 frames per second (FPS), while maintaining improved tracking and mapping accuracy. Against ORB-SLAM3 in RGB-D mode, Pow3R-SLAM is more accurate on TUM, 7-Scenes, and ETH3D-SLAM, and completes every TUM sequence. While Pow3R-SLAM can struggle on a small set of self-similar scenes, its overall performance shows that adding depth as a prior for two-view 3D reconstruction SLAM can be beneficial. A project webpage is available at: https://ChrisKolios.github.io/Pow3R-SLAM , and code will be made open-source upon acceptance.
Figures & tables
Fig. 1: Sensor depth vs. Pow3R output . Given input frames ((a), (b)), with corresponding depth maps ((c), (d)) from the TUM fr1/desk scene [ 3 ] , (e) shows the pointmap result from the back-projected depthmaps, and (f) shows the pointmap result from passing the input RGB-D into Pow3R’s network. Pow3R maintains flat surfaces, and reconstructs the controller where the Kinect sensor lacks the data to do so (fills holes and regularizes surfaces).
Fig. 2: Pow3R-SLAM pipeline. Pow3R-SLAM is an RGB-D SLAM system that uses sensor depth as a prior on a two-view pointmap network. The front end (1-3) runs on every frame. (1) Pow3R predicts pointmaps and confidences for the frame and its keyframe, conditioned on sensor depth and intrinsics (Sec. III-B ). (2) Each prediction is made metric against the sensor depth (Eq. 1 ). (3) The frame is tracked in Sim(3) and may become a keyframe (Secs. III-C – III-D ). The backend (4, 5) runs on every new keyframe. (4) Loop closure on Pow3R’s encoder tokens adds edges in familiar locations, and relocalization locates frames the tracker has lost. (5) A global optimization refines every keyframe pose, with keyframe depths anchored to the sensor and confidence-calibrated edges (Sec. III-E3 , Eqs. 5 – 7 ). Blue : the sensor-depth path. Orange : optional components, including the hybrid variant, which tracks alternate frames by ICP instead of Pow3R (Sec. III-G ), and map-only keyframes (Sec. III-F ). ICP frames never become keyframes, so every keyframe comes from a Pow3R pass. Inputs and outputs: TUM fr1/desk .
ORB-SLAM3
Pow3R-SLAM
scene
DROID † ↓
Sim3 ↓
SE3 * ↓
MASt3R fresh ↓
Sim3 ↓
SE3 * ↓
hyb. Sim3 ↓
TUM RGB-D fr1
360
0.111
0.134
0.228
0.049
0.040
0.048
0.042
desk
0.018
0.018
0.018
0.016
0.019
0.035
0.017
desk2
0.042
X
X
0.024
0.024
0.093
0.022
floor
0.021
X
X
0.025
0.021
0.025
0.025
TABLE I: ATE root-mean square error (RMSE) (m) on TUM RGB-D fr1 and 7-Scenes . Bold , underline : best, second best. † : calibrated monocular result as reported in [ 7 ] . Fresh: our re-run. X: lost track, with means over completed scenes. ∗ : unscaled SE(3) alignment.
scene / metric
MASt3R fresh
Pow3R- SLAM
Pow3R- SLAM hyb.
office0
0.013
0.009
0.008
office1
0.007
0.009
0.013
office2
0.022
0.007
0.007
office3
0.027
0.010
0.012
office4
0.021
0.009
0.011
room0
0.006
0.007
0.007
TABLE II: Replica, ETH3D-SLAM and runtime. Keyframe ATE RMSE (m), Sim(3), with the unscaled SE(3) mean. ETH3D-SLAM: over each system’s completed sequences (count in brackets). Runtime: wall time summed over the 24-scene panel (TUM, 7-Scenes, Replica), at full and subsample-2 (s2). Mean of 2 exclusive repeats per system ( ± half-range), with MASt3R-SLAM timed on native images. FPS: frames per second at s2 of loop time excluding model load.
Fig. 3: Qualitative reconstruction. Rows: Replica office3 , 7-Scenes chess , ETH3D-SLAM mannequin_5 (full frame rate). Columns: (a) reference cloud from ground-truth depth, (b) MASt3R-SLAM and (c) Pow3R-SLAM maps after alignment, coloured by distance to the reference (capped at 10 cm), and (d) top view of the per-frame trajectories (ground truth (GT) black, MASt3R-SLAM orange, Pow3R-SLAM blue).
Fig. 4: The failure case . 7-Scenes stairs , with the same columns as Fig. 3 , (d) GT / M(ASt3R) / P(ow3R). MASt3R- vs. Pow3R-SLAM: 1.5 vs. 4.3 M points, keyframe ATE 0.016 vs. 0.044 m, accuracy 3.2 vs. 6.7 cm, completion 10.5 vs. 5.9 cm. Pow3R-SLAM’s map is more complete but less accurate, which we attribute to the self-similar treads (Sec. V ).
dataset
system
ATE ↓
Acc. ↓
Comp. ↓
Cham. ↓
Ch. RMSE ↓
7-Scenes
MASt3R-SLAM fresh
0.048
0.051
0.058
0.055
0.085
Pow3R-SLAM
0.044
0.048
0.038
0.043
0.058
Pow3R-SLAM hyb.
0.048
0.050
0.038
0.044
0.058
Replica
MASt3R-SLAM fresh
0.015
0.037
0.031
0.034
0.056
Pow3R-SLAM
0.009
0.022
0.018
0.020
0.027
Pow3R-SLAM hyb.
0.010
0.023
0.019
0.021
0.028
TABLE III: Reconstruction on 7-Scenes and Replica. ATE for reference. Mean accuracy (Acc.), completion (Comp.), Chamfer (Cham.), and RMSE Chamfer, in m.
configuration
Sim3 mean (24) ↓
Sim3 median (24) ↓
SE3 mean (24) ↓
Chamfer (15) ↓
stairs ↓
ORB-SLAM3 (ref., 21)
0.0342
0.0178
0.0411
–
0.045
MASt3R fresh (ref.)
0.0304
0.0228
0.1104
4.35
0.016
L1 Pow3R, no priors
0.0834
0.0449
0.4094
9.55
0.063
L2 + intrinsics
0.0711
0.0431
0.4058
8.70
0.052
L3 + cond
0.0319
0.0189
0.3640
6.58
0.115
L4 + scale
0.0260
0.0191
0.0337
3.07
0.044
TABLE IV: Ablations . Means and median ATE over the 24 scenes, Chamfer (cm) over the 15 map scenes, with stairs separate. Rows L1–L5 add one component at a time up to the shipped default (def.); L6 and L7 leave out depth sites (23 scenes, as fr1/rpy diverges without cond). ref.: reference systems, not ranked (ORB-SLAM3 over its 21 completions). KF: keyframes.