Floorplans provide compact and widely available geometric maps for indoor localization, but existing high-performing floorplan-based methods still convert them into dense scene-specific offline databases, tying accuracy, storage, and runtime to the sampling resolution of the discretized pose space. We present FreeLoc, an online RGB-based floorplan localization framework that treats the floorplan as a directly queryable geometric map. FreeLoc introduces an efficient online geometric querying and diffusion-aided refinement scheme, which retrieves plausible pose anchors through on-the-fly floorplan ray querying and refines them into accurate continuous pose estimates. For sequential localization, FreeLoc develops an online likelihood construction strategy that bridges single-frame localization and probabilistic temporal fusion by constructing likelihoods from coarse-sampled candidates and refined pose hypotheses, enabling histogram-filter-based temporal fusion without offline databases. Experiments demonstrate real-time online inference and state-of-the-art performance in both single-frame and sequential localization, while real-world results validate practical deployability in indoor robotic localization scenarios.
Figures & tables
Figure 1: FreeLoc enables database-free online floorplan localization, directly using the provided floorplan to support accurate single-frame and sequential localization in new environments.
Figure 2: Overview of FreeLoc. We retrieve top- K coarse anchors by matching image-side rays with online queried floorplan rays, then refine and re-score them for the final estimate. For sequential localization, coarse and refined candidates construct online likelihoods for temporal fusion.
Method
Gibson(f)
Structured3D
@0.1m↑
@0.5m↑
@1m↑
@1m30∘↑
@0.1m↑
@0.5m↑
@1m↑
@1m30∘↑
PF-Net
0
1.0
4.1
1.1
0.1
1.2
4.4
1.4
LASER
0.2
3.9
9.8
6.2
0.5
5.2
10.0
7.8
F 3 Loc
5.3
30.7
37.9
36.3
1.7
15.0
23.0
21.9
Ours
10.0
39.4
44.6
42.9
3.1
23.9
30.5
29.1
Table 1: Single-frame localization recall (%) on Gibson(f) and Structured3D. Bold denotes the best result.
Method
SR@1m (%) ↑
RMSEsucc (m) ↓
RMSEall (m) ↓
PF-Net
8.0
0.43
4.27
LASER
35.0
0.31
2.69
F 3 Loc
86.5
0.14
0.87
Ours
100.0
0.15
0.15
Table 2: Sequential localization results on Gibson(t).
Figure 3: Real-world performance of FreeLoc. The center panel shows the floorplan with converged predicted and ground-truth trajectories; surrounding panels show time-ordered RGB observations and posterior probabilities. Blue indicates predictions and green indicates ground truth.
Initial error
Before (m)
After (m)
≤0.5m
0.25
0.18
≤1.0m
0.35
0.25
≤1.5m
0.42
0.32
All
3.14
3.09
Table 3: Pose refinement under different coarse-initialization quality.
Refinement
SR@1m (%) ↑
SR@0.5m (%) ↑
SR@0.2m (%) ↑
RMSE all (m) ↓
ICP Optimization
97.3
97.3
43.2
0.21
Diffusion-Aided Pose Refinement
100.0
100.0
51.4
0.15
Table 4: Sequential localization comparison between ICP optimization and diffusion-aided pose refinement on Gibson(t).
Setting
Level
SR@1m
RMSE succ
RMSE all
Clean
–
100.0
0.145
0.145
Gaussian
Mild
100.0
0.147
0.147
Moderate
100.0
0.171
0.171
Severe
97.3
0.235
0.307
Motion-dep.
Mild
100.0
0.145
0.145
Moderate
100.0
0.150
0.150
Table 5: Robustness to inter-frame ego-motion uncertainty on Gibson(t).
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Qualitative results of diffusion-aided pose refinement. Starting from coarse pose anchors retrieved by online floorplan sampling, our diffusion-aided pose refinement module progressively predicts residual pose corrections guided by paired image-side and floorplan-side ray geometry. The color transition from light blue to dark blue indicates the pose recovery process, while green denotes the ground truth. Coarse anchors from different initial locations are refined toward poses closer to the ground truth, validating the effectiveness of diffusion-aided pose refinement.
Figure 5: Qualitative comparison of single-frame localization results. We compare FreeLoc with representative baselines on Gibson and Structured3D. Compared with PF-Net, LASER, and F 3 Loc, FreeLoc produces pose estimates that are better aligned with the ground truth, benefiting from accurate online coarse pose retrieval and diffusion-aided pose refinement.
Figure 6: Qualitative comparison of sequential localization results. We compare FreeLoc with representative baselines on the Gibson(t) sequential localization dataset. The results show the temporal evolution of posterior distributions. Benefiting from more reliable single-frame localization evidence, FreeLoc concentrates the posterior around the correct region more quickly and maintains more accurate localization over time.
Refinement
SR@1m (%) ↑
SR@0.5m (%) ↑
SR@0.2m (%) ↑
RMSE all (m) ↓
Direct Regression
97.3
97.3
48.6
0.20
Diffusion-Aided Pose Refinement
100.0
100.0
51.4
0.15
Appendix
Table 7: Ablation study of diffusion-aided pose refinement against direct regression on Gibson(t).
Method
Gibson(f)
Structured3D
@0.1m↑
@0.5m↑
@1m↑
@1m30∘↑
@0.1m↑
@0.5m↑
@1m↑
@1m30∘↑
UnLoc
13.5
46.1
49.6
47.8
3.6
27.7
33.6
32.6
Ours
16.6
46.9
50.2
48.7
4.9
29.5
34.3
33.0
Appendix
Table 8: Controlled comparison with UnLoc under matched depth predictions. FreeLoc uses the same UnLoc depth predictions at test time without retraining the diffusion refinement model. Bold denotes the better result.