Feed-forward 3D models provide strong multi-view geometric priors, while on- line simultaneous localization and mapping (SLAM) relies mainly on local mea- surements and can accumulate drift over long sequences. Existing attempts to combine the two typically treat feed-forward predictions as an external geomet- ric state that is aligned or fused with the online estimate after the fact, which keeps broader multi-view evidence outside the optimizer that refines the SLAM state. We present F2SLAM, which instead converts feed-forward geometry di- rectly into optimization-native target-weight measurements attached to a persis- tent dense factor graph. A high-frequency stream maintains local tracking con- straints and graph connectivity, while a low-frequency stream uses wider multi- view context to selectively refresh existing measurements after a state-consistency check. Both streams constrain the same poses, inverse depths, and optional cam- era intrinsics through a single dense bundle adjustment. Experiments on multiple benchmarks demonstrate consistently strong trajectory estimation and improved dense reconstruction in both calibrated and uncalibrated settings. Notably, the uncalibrated configuration reduces the average ATE RMSE from 0.030 m for the strongest feed-forward baseline to 0.002 m on the Replica dataset.
Figures & tables
Figure 1: Motivation and overview of F 2 SLAM. Conventional dense SLAM (top) relies on local frame-to-frame measurements, while foundation-model-based approaches (middle) reconstruct and align local submaps. In contrast, F 2 SLAM (bottom) integrates high-frequency tracking and low-frequency feed-forward geometry in a persistent dense factor graph, where low-frequency predictions selectively refresh existing measurements for joint refinement through a single dense BA. The inset reports ATE RMSE on the TUM RGB-D dataset Sturm et al. (2012) .
Figure 2: Pipeline of F 2 SLAM. The high-frequency stream selects keyframes, maintains the graph topology, and frequently updates local target-weight measurements. The low-frequency stream refreshes measurements on selected existing edges using wider multi-view context, without adding new edges. Both streams constrain the same persistent geometric state through a single dense BA.
Method
Uncalib.
360
desk
desk2
floor
plant
room
rpy
teddy
xyz
Avg.
DROID-SLAM ( Teed and Deng, 2021 )
✗
0.111
0.018
0.042
0.021
0.016
0.049
0.026
0.048
0.012
0.038
DPV-SLAM ( Lipson et al., 2024 )
✗
0.112
0.018
0.029
0.057
0.021
0.330
0.030
0.084
0.010
0.076
DPV-SLAM++ ( Lipson et al., 2024 )
✗
0.132
0.018
0.029
0.050
0.022
0.096
0.032
0.098
0.010
0.054
GO-SLAM ( Zhang et al., 2023 )
✗
0.089
0.016
0.028
0.025
0.026
0.052
0.019
0.048
0.010
0.035
MASt3R-SLAM ( Murai et al., 2025 )
✗
0.049
0.016
0.024
0.025
0.020
0.061
0.027
0.041
0.009
0.030
F 2 SLAM (ours)
✗
0.070
0.017
0.026
0.021
0.014
0.049
0.023
0.034
0.010
0.029
Table 1: Trajectory evaluation (ATE RMSE [m] ↓ ) on the TUM RGB-D dataset ( Sturm et al., 2012 ) . ∗ denotes results reproduced on our hardware.
Method
Uncalib.
chess
fire
heads
office
pumpkin
kitchen
stairs
Avg.
DROID-SLAM ( Teed and Deng, 2021 )
✗
0.036
0.027
0.025
0.066
0.127
0.040
0.026
0.050
NICER-SLAM ( Zhu et al., 2024 )
✗
0.033
0.069
0.042
0.108
0.200
0.039
0.108
0.086
F 2 SLAM (ours)
✗
0.036
0.029
0.025
0.088
0.140
0.042
0.021
0.053
DROID-SLAM ( Teed and Deng, 2021 )
✓
0.047
0.038
0.034
0.136
0.166
0.080
0.044
0.078
MASt3R-SLAM ( Murai et al., 2025 )
✓
0.063
0.046
0.029
0.103
0.114
0.074
0.032
0.066
VGGT-SLAM ∗ ( Maggio et al., 2025 )
✓
0.037
0.026
0.018
0.104
0.133
0.061
0.093
0.067
Table 2: Trajectory evaluation (ATE RMSE [m] ↓ ) on the 7-Scenes dataset ( Shotton et al., 2013 ) . ∗ denotes results reproduced on our hardware.
Method
Uncalib.
R0
R1
R2
O0
O1
O2
O3
O4
Avg.
DROID-SLAM ( Teed and Deng, 2021 )
✗
0.003
0.001
0.003
0.003
0.004
0.003
0.005
0.004
0.003
NICER-SLAM ( Zhu et al., 2024 )
✗
0.013
0.016
0.011
0.021
0.032
0.021
0.014
0.020
0.019
F 2 SLAM (Ours)
✗
0.003
0.002
0.003
0.002
0.003
0.003
0.003
0.004
0.003
VGGT-SLAM ( Maggio et al., 2025 )
✓
0.030
0.167
0.086
0.042
0.064
0.095
0.039
0.043
0.071
VGGT-SLAM 2.0 ∗ ( Maggio and Carlone, 2026 )
✓
0.030
0.049
0.040
0.026
0.018
0.024
0.024
0.029
0.030
SLAM-Former ∗ ( Yuan et al., 2026 )
✓
0.031
0.033
0.025
0.030
0.027
0.038
0.033
0.036
0.032
Table 3: Trajectory evaluation (ATE RMSE [m] ↓ ) on the Replica dataset ( Straub et al., 2019 ) . ∗ denotes results reproduced on our hardware.
Method
Uncalib.
00
01
02
03
04
05
06
07
08
09
10
Avg.
ORB-SLAM2 ( Mur-Artal and Tardós, 2017 )
✗
40.65
502.20
47.82
0.94
1.30
29.95
40.82
16.04
43.09
38.77
5.42
69.73
LDSO ( Gao et al., 2018 )
✗
9.32
11.68
31.98
2.85
1.22
5.10
13.55
2.96
129.02
21.64
17.36
22.43
DROID-SLAM ( Teed and Deng, 2021 )
✗
92.10
5344.60
107.61
2.38
1.00
118.50
62.47
21.78
161.60
72.32
118.70
554.82
DPV-SLAM ( Lipson et al., 2024 )
✗
112.80
11.50
123.53
2.50
0.81
57.80
54.86
18.77
110.49
76.66
13.65
53.03
DPV-SLAM++ ( Lipson et al., 2024 )
✗
8.30
11.86
39.64
2.50
0.78
5.74
11.60
1.52
110.90
76.70
13.70
25.75
F 2 SLAM (ours)
✗
117.08
64.77
134.50
3.94
1.69
83.39
65.43
25.37
112.10
102.82
27.95
67.19
Table 4: Trajectory evaluation (ATE RMSE [m] ↓ ) on the KITTI Odometry dataset ( Geiger et al., 2012 ) . Baseline results are taken from VGGT-SLAM++ ( Mandal et al., 2026 ) .
Room0
Room1
Room2
Office0
Office1
Office2
Office3
Office4
Avg.
Method
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
Acc.
Comp.
DROID-SLAM + ( Teed and Deng, 2021 )
0.1218
0.0896
0.0835
0.0607
0.0326
0.1601
0.0301
0.1619
0.0239
0.1619
0.0566
0.1556
0.0449
0.0973
0.0465
0.0963
0.0550
0.1229
NICER-SLAM + ( Zhu et al., 2024 )
0.0253
0.0304
0.0393
0.0410
0.0340
0.0342
0.0549
0.0609
0.0345
0.0442
0.0402
0.0429
0.0334
0.0403
0.0303
0.0387
0.0365
0.0416
MASt3R-SLAM ∗ ( Murai et al., 2025 )
0.2468
0.2675
0.0393
0.0266
0.0258
0.0187
0.0336
0.0211
0.0359
0.0303
0.0609
0.0484
0.0881
0.0824
0.0742
0.0577
0.0756
0.0691
SLAM-Former ∗ ( Yuan et al., 2026 )
0.0339
0.0247
0.0248
0.0171
0.0244
0.0166
0.0208
0.0153
0.0178
0.0138
0.0345
0.0222
0.0296
0.0213
0.0305
0.0214
0.0270
0.0187
VGGT-SLAM 2.0 ∗ ( Maggio and Carlone, 2026 )
0.0305
0.0622
0.0564
0.0765
0.0277
0.0629
0.0475
0.0457
0.0340
0.0595
0.0249
0.0792
0.0322
0.0739
0.0408
0.0743
0.0367
0.0657
Table 5: Reconstruction accuracy (Acc. [m] ↓ ) and completion (Comp. [m] ↓ ) on the Replica dataset ( Straub et al., 2019 ) . + denotes results reported in NICER-SLAM, − denotes results reported in SLAM3R, and ∗ denotes results reproduced on our hardware.
Table 8
Figure 3: Qualitative reconstruction comparison on representative sequences. F 2 SLAM recovers more continuous geometry with fewer missing regions and artifacts.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Table 10
Method
Opt. Intr.
Chess
Fire
Heads
Office
Pumpkin
RedKitchen
Stairs
Avg.
w/ target motion 600
✓
0.036
0.024
0.014
0.092
0.143
0.048
0.025
0.054
w/ target motion 800 (Ours)
✓
0.036
0.024
0.014
0.087
0.140
0.041
0.020
0.051
w/ target motion 1000
✓
0.036
0.025
0.014
0.087
0.140
0.041
0.020
0.051
Appendix
Table 10: Ablation study of Flow3r target motion on 7-Scenes. We report Sim(3)-aligned ATE RMSE (m). Lower is better.
Figure 4: Ablation on the low-frequency context window size. Lower is better for all metrics. Log-scaled y-axes retain a continuous line plot while keeping the large errors at sizes 4 and 8 visible. We use a logarithmic coordinate system to better show the overall trend.
Figure 5: Qualitative trajectory comparison on representative sequences. All estimated trajectories are aligned to the ground truth using the same evaluation protocol. Compared with competing methods, F 2 SLAM follows the reference trajectory more closely throughout the sequence and exhibits less accumulated drift.
Figure 6: Two-dimensional trajectory projections on representative sequences. The projections emphasize global shape, turning behavior, and endpoint drift. F 2 SLAM better preserves the reference path and avoids the progressive deviation observed in competing methods.
Figure 7: Three-dimensional trajectory visualization. The full spatial trajectories reveal vertical drift and geometric distortion that can be hidden by planar projections. F 2 SLAM maintains a more coherent trajectory across all spatial directions.
Figure 8: Qualitative reconstruction results on representative Replica scenes. F 2 SLAM recovers coherent room layouts and preserves major scene structures across long sequences.
Figure 9: Qualitative reconstruction results on representative KITTI Odometry sequences. F 2 SLAM maintains improved global consistency and reduces accumulated drift over long driving sequences.
State Key Laboratory of Industrial Control Technology, College of Control Science and Engineering, Zhejiang University, Hangzhou, China, 310027. · Key Laboratory of Collaborative Sensing and Autonomous Unmanned Systems of Zhejiang Province, Hangzhou, China, 310027.