Feed-forward 3D models provide strong multi-view geometric priors, while on- line simultaneous localization and mapping (SLAM) relies mainly on local mea- surements and can accumulate drift over long sequences. Existing attempts to combine the two typically treat feed-forward predictions as an external geomet- ric state that is aligned or fused with the online estimate after the fact, which keeps broader multi-view evidence outside the optimizer that refines the SLAM state. We present F2SLAM, which instead converts feed-forward geometry di- rectly into optimization-native target-weight measurements attached to a persis- tent dense factor graph. A high-frequency stream maintains local tracking con- straints and graph connectivity, while a low-frequency stream uses wider multi- view context to selectively refresh existing measurements after a state-consistency check. Both streams constrain the same poses, inverse depths, and optional cam- era intrinsics through a single dense bundle adjustment. Experiments on multiple benchmarks demonstrate consistently strong trajectory estimation and improved dense reconstruction in both calibrated and uncalibrated settings. Notably, the uncalibrated configuration reduces the average ATE RMSE from 0.030 m for the strongest feed-forward baseline to 0.002 m on the Replica dataset.
Figures & tables
Figure 1: Motivation and overview of F 2 SLAM. Conventional dense SLAM (top) relies on local frame-to-frame measurements, while foundation-model-based approaches (middle) reconstruct and align local submaps. In contrast, F 2 SLAM (bottom) integrates high-frequency tracking and low-frequency feed-forward geometry in a persistent dense factor graph, where low-frequency predictions selectively refresh existing measurements for joint refinement through a single dense BA. The inset reports ATE RMSE on the TUM RGB-D dataset Sturm et al. (2012) .
Figure 2: Pipeline of F 2 SLAM. The high-frequency stream selects keyframes, maintains the graph topology, and frequently updates local target-weight measurements. The low-frequency stream refreshes measurements on selected existing edges using wider multi-view context, without adding new edges. Both streams constrain the same persistent geometric state through a single dense BA.
Method
Uncalib.
360
desk
desk2
floor
plant
room
rpy
teddy
xyz
Avg.
DROID-SLAM ( Teed and Deng, 2021 )
✗
0.111
0.018
0.042
0.021
0.016
0.049
0.026
0.048
0.012
0.038
DPV-SLAM ( Lipson et al., 2024 )
✗
0.112
0.018
0.029
0.057
0.021
0.330
0.030
0.084
0.010
0.076
DPV-SLAM++ ( Lipson et al., 2024 )
✗
0.132
0.018
0.029
0.050
0.022
0.096
0.032
0.098
0.010
0.054
GO-SLAM ( Zhang et al., 2023 )
✗
0.089
0.016
0.028
0.025
0.026
0.052
0.019
0.048
0.010
0.035
MASt3R-SLAM ( Murai et al., 2025 )
✗
0.049
0.016
0.024
0.025
0.020
0.061
0.027
0.041
0.009
0.030
F 2 SLAM (ours)
✗
0.070
0.017
0.026
0.021
0.014
0.049
0.023
0.034
0.010
0.029
Table 1: Trajectory evaluation (ATE RMSE [m] ↓ ) on the TUM RGB-D dataset ( Sturm et al., 2012 ) . ∗ denotes results reproduced on our hardware.
Method
Uncalib.
chess
fire
heads
office
pumpkin
kitchen
stairs
Avg.
DROID-SLAM ( Teed and Deng, 2021 )
✗
0.036
0.027
0.025
0.066
0.127
0.040
0.026
0.050
NICER-SLAM ( Zhu et al., 2024 )
✗
0.033
0.069
0.042
0.108
0.200
0.039
0.108
0.086
F 2 SLAM (ours)
✗
0.036
0.029
0.025
0.088
0.140
0.042
0.021
0.053
DROID-SLAM ( Teed and Deng, 2021 )
✓
0.047
0.038
0.034
0.136
0.166
0.080
0.044
0.078
MASt3R-SLAM ( Murai et al., 2025 )
✓
0.063
0.046
0.029
0.103
0.114
0.074
0.032
0.066
VGGT-SLAM ∗ ( Maggio et al., 2025 )
✓
0.037
0.026
0.018
0.104
0.133
0.061
0.093
0.067
Table 2: Trajectory evaluation (ATE RMSE [m] ↓ ) on the 7-Scenes dataset ( Shotton et al., 2013 ) . ∗ denotes results reproduced on our hardware.
Method
Uncalib.
R0
R1
R2
O0
O1
O2
O3
O4
Avg.
DROID-SLAM ( Teed and Deng, 2021 )
✗
0.003
0.001
0.003
0.003
0.004
0.003
0.005
0.004
0.003
NICER-SLAM ( Zhu et al., 2024 )
✗
0.013
0.016
0.011
0.021
0.032
0.021
0.014
0.020
0.019
F 2 SLAM (Ours)
✗
0.003
0.002
0.003
0.002
0.003
0.003
0.003
0.004
0.003
VGGT-SLAM ( Maggio et al., 2025 )
✓
0.030
0.167
0.086
0.042
0.064
0.095
0.039
0.043
0.071
VGGT-SLAM 2.0 ∗ ( Maggio and Carlone, 2026 )
✓
0.030
0.049
0.040
0.026
0.018
0.024
0.024
0.029
0.030
SLAM-Former ∗ ( Yuan et al., 2026 )
✓
0.031
0.033
0.025
0.030
0.027
0.038
0.033
0.036
0.032
Table 3: Trajectory evaluation (ATE RMSE [m] ↓ ) on the Replica dataset ( Straub et al., 2019 ) . ∗ denotes results reproduced on our hardware.
Method
Uncalib.
00
01
02
03
04
05
06
07
08
09
10
Avg.
ORB-SLAM2 ( Mur-Artal and Tardós, 2017 )
✗
40.65
502.20
47.82
0.94
1.30
29.95
40.82
16.04
43.09
38.77
5.42
69.73
LDSO ( Gao et al., 2018 )
✗
9.32
11.68
31.98
2.85
1.22
5.10
13.55
2.96
129.02
21.64
17.36
22.43
DROID-SLAM ( Teed and Deng, 2021 )
✗
92.10
5344.60
107.61
2.38
1.00
118.50
62.47
21.78
161.60
72.32
118.70
554.82
DPV-SLAM ( Lipson et al., 2024 )
✗
112.80
11.50
123.53
2.50
0.81
57.80
54.86
18.77
110.49
76.66
13.65
53.03
DPV-SLAM++ ( Lipson et al., 2024 )
✗
8.30
11.86
39.64
2.50
0.78
5.74
11.60
1.52
110.90
76.70
13.70
25.75
F 2 SLAM (ours)
✗
117.08
64.77
134.50
3.94
1.69
83.39
65.43
25.37
112.10
102.82
27.95
67.19
Table 4: Trajectory evaluation (ATE RMSE [m] ↓ ) on the KITTI Odometry dataset ( Geiger et al., 2012 ) . Baseline results are taken from VGGT-SLAM++ ( Mandal et al., 2026 ) .
Room0
Room1
Room2
Office0
Office1
Office2
Office3
Office4
Avg.
Method
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
DROID-SLAM + ( Teed and Deng, 2021 )
0.1218
0.0896
0.0835
0.0607
0.0326
0.1601
0.0301
0.1619
0.0239
0.1619
0.0566
0.1556
0.0449
0.0973
0.0465
0.0963
0.0550
0.1229
NICER-SLAM + ( Zhu et al., 2024 )
0.0253
0.0304
0.0393
0.0410
0.0340
0.0342
0.0549
0.0609
0.0345
0.0442
0.0402
0.0429
0.0334
0.0403
0.0303
0.0387
0.0365
0.0416
MASt3R-SLAM ∗ ( Murai et al., 2025 )
0.2468
0.2675
0.0393
0.0266
0.0258
0.0187
0.0336
0.0211
0.0359
0.0303
0.0609
0.0484
0.0881
0.0824
0.0742
0.0577
0.0756
0.0691
SLAM-Former ∗ ( Yuan et al., 2026 )
0.0339
0.0247
0.0248
0.0171
0.0244
0.0166
0.0208
0.0153
0.0178
0.0138
0.0345
0.0222
0.0296
0.0213
0.0305
0.0214
0.0270
0.0187
VGGT-SLAM 2.0 ∗ ( Maggio and Carlone, 2026 )
0.0305
0.0622
0.0564
0.0765
0.0277
0.0629
0.0475
0.0457
0.0340
0.0595
0.0249
0.0792
0.0322
0.0739
0.0408
0.0743
0.0367
0.0657
Table 5: Reconstruction accuracy (Acc. [m] ↓ ) and completion (Comp. [m] ↓ ) on the Replica dataset ( Straub et al., 2019 ) . + denotes results reported in NICER-SLAM, − denotes results reported in SLAM3R, and ∗ denotes results reproduced on our hardware.
Table 8
Figure 3: Qualitative reconstruction comparison on representative sequences. F 2 SLAM recovers more continuous geometry with fewer missing regions and artifacts.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Table 10
Method
Opt. Intr.
Chess
Fire
Heads
Office
Pumpkin
RedKitchen
Stairs
Avg.
w/ target motion 600
✓
0.036
0.024
0.014
0.092
0.143
0.048
0.025
0.054
w/ target motion 800 (Ours)
✓
0.036
0.024
0.014
0.087
0.140
0.041
0.020
0.051
w/ target motion 1000
✓
0.036
0.025
0.014
0.087
0.140
0.041
0.020
0.051
Appendix
Table 10: Ablation study of Flow3r target motion on 7-Scenes. We report Sim(3)-aligned ATE RMSE (m). Lower is better.
Figure 4: Ablation on the low-frequency context window size. Lower is better for all metrics. Log-scaled y-axes retain a continuous line plot while keeping the large errors at sizes 4 and 8 visible. We use a logarithmic coordinate system to better show the overall trend.
Figure 5: Qualitative trajectory comparison on representative sequences. All estimated trajectories are aligned to the ground truth using the same evaluation protocol. Compared with competing methods, F 2 SLAM follows the reference trajectory more closely throughout the sequence and exhibits less accumulated drift.
Figure 6: Two-dimensional trajectory projections on representative sequences. The projections emphasize global shape, turning behavior, and endpoint drift. F 2 SLAM better preserves the reference path and avoids the progressive deviation observed in competing methods.
Figure 7: Three-dimensional trajectory visualization. The full spatial trajectories reveal vertical drift and geometric distortion that can be hidden by planar projections. F 2 SLAM maintains a more coherent trajectory across all spatial directions.
Figure 8: Qualitative reconstruction results on representative Replica scenes. F 2 SLAM recovers coherent room layouts and preserves major scene structures across long sequences.
Figure 9: Qualitative reconstruction results on representative KITTI Odometry sequences. F 2 SLAM maintains improved global consistency and reduces accumulated drift over long driving sequences.
Recent geometric foundation models enable feed-forward inference for SLAM, but their predictions are strongly dependent on the input view set, which leads to geometric inconsistencies and trajectory drift when results are chained over long sequences. Online deployment further exposes a trade-off between the low latency of two-view tracking and the constraint richness of multi-view inference. We introduce UniSim-SLAM, an integrated system that runs lightweight two-view keyframe tracking in the frontend and performs periodic multi-view submap refinement in the backend. To combine predictions defined in heterogeneous local coordinates with inconsistent scales, we formulate a unified multi-level factor graph on Sim(3) that jointly optimizes global keyframe poses and submap poses. The graph integrates temporal view-to-view odometry edges, view-to-submap bridge edges with depth-statistics scale anchoring, and submap-to-submap tie and scale constraints to enforce consistent similarity relations across submaps. Experiments on TUM RGB-D and 7-Scenes show that UniSim-SLAM achieves state-of-the-art accuracy in the uncalibrated setting, reducing trajectory error by 38.5% on TUM RGB-D and 45.9% on 7-Scenes compared to prior best results. Project page: https://vision3d-lab.github.io/unisim-slam/
Inha Lee, Dongjae Jeong, Junhee Lee +1
lsan National Institute of Science and Technology, Ulsan, Korea
Recent works have explored unifying SLAM with geometric foundation models (GFMs). However, directly using GFM predictions for tracking is highly sensitive to model capability and uncertainty, as geometric inaccuracies in the predictions can adversely affect pose estimation. To address this limitation, we propose a decoupled framework that integrates classical feature-based SLAM with GFMs, which achieves higher quality and more consistent dense reconstruction. In brief, we use classical visual SLAM for robust low-latency tracking and use GFMs exclusively for mapping. By anchoring mapping to poses produced by the SLAM module and optimizing across depth scales, the proposed design avoids propagating inaccuracies from GFM predictions into pose estimation while imposing geometric constraints on the reconstruction. The system builds submaps from multiple posed keyframes and enforces scale consistency via lightweight frame and submap scale optimization. It also performs projection-based point cloud fusion within each submap, and updates submaps online to reflect trajectory updates from the feature-based SLAM. To evaluate tracking and reconstruction of our method, we introduce a loop-rich, building-scale indoor dataset with accurate sensor trajectories and LiDAR ground-truth. Experiments show that our approach achieves superior trajectory accuracy while improving reconstruction precision by 10%-20% over existing methods, with about 2 cm reconstruction error per 10 m chunk on building-scale dataset. On large-scale outdoor datasets, it attains 10 cm error per 30 m chunk (w.r.t LiDAR ground-truth models). Code and dataset: https://github.com/ori-drs/ScaRF-SLAM
Yuhao Zhang, Yifu Tao, Frank Dellaert +1
Oxford Robotics Institute, University of Oxford, UK · School of Interactive Computing, Georgia Institute of Technology, USA
SLAM methods based on 3D Gaussian Splatting (3DGS) have demonstrated impressive tracking and mapping performance, but typically require additional geometric information from external depth sensors. Meanwhile, recent SLAM systems that leverage geometric priors from pre-trained feed-forward models enable real-time dense reconstruction, yet often discard original RGB information during optimization, thus degrading overall reconstruction quality. We present GeoGS-SLAM, an online monocular dense reconstruction system that combines the 3DGS-based map representation with learned geometric priors. Given uncalibrated RGB input, we first employ a feed-forward visual geometry model to predict camera and scene priors. The Gaussian scene map is then expanded by directly sampling Gaussian primitives from both RGB input and geometric priors. Camera poses and the scene map are jointly optimized through a coarse-to-fine strategy that minimizes both photometric and geometric losses. To ensure global consistency, we further incorporate online loop closure detection and pose graph optimization. Extensive experiments across indoor and outdoor benchmarks demonstrate that GeoGS-SLAM achieves superior rendering quality and tracking accuracy compared to state-of-the-art methods while maintaining online real-time performance. Project page: https://rlgao.github.io/geogs_slam.
Ruilan Gao, Letian Jin, Yu Zhang
State Key Laboratory of Industrial Control Technology, College of Control Science and Engineering, Zhejiang University, Hangzhou, China, 310027. · Key Laboratory of Collaborative Sensing and Autonomous Unmanned Systems of Zhejiang Province, Hangzhou, China, 310027.