Floorplans provide compact and widely available geometric maps for indoor localization, but existing high-performing floorplan-based methods still convert them into dense scene-specific offline databases, tying accuracy, storage, and runtime to the sampling resolution of the discretized pose space. We present FreeLoc, an online RGB-based floorplan localization framework that treats the floorplan as a directly queryable geometric map. FreeLoc introduces an efficient online geometric querying and diffusion-aided refinement scheme, which retrieves plausible pose anchors through on-the-fly floorplan ray querying and refines them into accurate continuous pose estimates. For sequential localization, FreeLoc develops an online likelihood construction strategy that bridges single-frame localization and probabilistic temporal fusion by constructing likelihoods from coarse-sampled candidates and refined pose hypotheses, enabling histogram-filter-based temporal fusion without offline databases. Experiments demonstrate real-time online inference and state-of-the-art performance in both single-frame and sequential localization, while real-world results validate practical deployability in indoor robotic localization scenarios.
Figures & tables
Figure 1: FreeLoc enables database-free online floorplan localization, directly using the provided floorplan to support accurate single-frame and sequential localization in new environments.
Figure 2: Overview of FreeLoc. We retrieve top- K coarse anchors by matching image-side rays with online queried floorplan rays, then refine and re-score them for the final estimate. For sequential localization, coarse and refined candidates construct online likelihoods for temporal fusion.
Method
Gibson(f)
Structured3D
@0.1m↑
@0.5m↑
@1m↑
@1m30∘↑
@0.1m↑
@0.5m↑
@1m↑
@1m30∘↑
PF-Net
0
1.0
4.1
1.1
0.1
1.2
4.4
1.4
LASER
0.2
3.9
9.8
6.2
0.5
5.2
10.0
7.8
F 3 Loc
5.3
30.7
37.9
36.3
1.7
15.0
23.0
21.9
Ours
10.0
39.4
44.6
42.9
3.1
23.9
30.5
29.1
Table 1: Single-frame localization recall (%) on Gibson(f) and Structured3D. Bold denotes the best result.
Method
SR@1m (%) ↑
RMSEsucc (m) ↓
RMSEall (m) ↓
PF-Net
8.0
0.43
4.27
LASER
35.0
0.31
2.69
F 3 Loc
86.5
0.14
0.87
Ours
100.0
0.15
0.15
Table 2: Sequential localization results on Gibson(t).
Figure 3: Real-world performance of FreeLoc. The center panel shows the floorplan with converged predicted and ground-truth trajectories; surrounding panels show time-ordered RGB observations and posterior probabilities. Blue indicates predictions and green indicates ground truth.
Initial error
Before (m)
After (m)
≤0.5m
0.25
0.18
≤1.0m
0.35
0.25
≤1.5m
0.42
0.32
All
3.14
3.09
Table 3: Pose refinement under different coarse-initialization quality.
Refinement
SR@1m (%) ↑
SR@0.5m (%) ↑
SR@0.2m (%) ↑
RMSE all (m) ↓
ICP Optimization
97.3
97.3
43.2
0.21
Diffusion-Aided Pose Refinement
100.0
100.0
51.4
0.15
Table 4: Sequential localization comparison between ICP optimization and diffusion-aided pose refinement on Gibson(t).
Setting
Level
SR@1m
RMSE succ
RMSE all
Clean
–
100.0
0.145
0.145
Gaussian
Mild
100.0
0.147
0.147
Moderate
100.0
0.171
0.171
Severe
97.3
0.235
0.307
Motion-dep.
Mild
100.0
0.145
0.145
Moderate
100.0
0.150
0.150
Table 5: Robustness to inter-frame ego-motion uncertainty on Gibson(t).
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Qualitative results of diffusion-aided pose refinement. Starting from coarse pose anchors retrieved by online floorplan sampling, our diffusion-aided pose refinement module progressively predicts residual pose corrections guided by paired image-side and floorplan-side ray geometry. The color transition from light blue to dark blue indicates the pose recovery process, while green denotes the ground truth. Coarse anchors from different initial locations are refined toward poses closer to the ground truth, validating the effectiveness of diffusion-aided pose refinement.
Figure 5: Qualitative comparison of single-frame localization results. We compare FreeLoc with representative baselines on Gibson and Structured3D. Compared with PF-Net, LASER, and F 3 Loc, FreeLoc produces pose estimates that are better aligned with the ground truth, benefiting from accurate online coarse pose retrieval and diffusion-aided pose refinement.
Figure 6: Qualitative comparison of sequential localization results. We compare FreeLoc with representative baselines on the Gibson(t) sequential localization dataset. The results show the temporal evolution of posterior distributions. Benefiting from more reliable single-frame localization evidence, FreeLoc concentrates the posterior around the correct region more quickly and maintains more accurate localization over time.
Refinement
SR@1m (%) ↑
SR@0.5m (%) ↑
SR@0.2m (%) ↑
RMSE all (m) ↓
Direct Regression
97.3
97.3
48.6
0.20
Diffusion-Aided Pose Refinement
100.0
100.0
51.4
0.15
Appendix
Table 7: Ablation study of diffusion-aided pose refinement against direct regression on Gibson(t).
Method
Gibson(f)
Structured3D
@0.1m↑
@0.5m↑
@1m↑
@1m30∘↑
@0.1m↑
@0.5m↑
@1m↑
@1m30∘↑
UnLoc
13.5
46.1
49.6
47.8
3.6
27.7
33.6
32.6
Ours
16.6
46.9
50.2
48.7
4.9
29.5
34.3
33.0
Appendix
Table 8: Controlled comparison with UnLoc under matched depth predictions. FreeLoc uses the same UnLoc depth predictions at test time without retraining the diffusion refinement model. Bold denotes the better result.
Visual Floorplan Localization (FLoc) has emerged as a promising solution for indoor localization by matching egocentric images against minimalist structural maps. However, due to cross-modal information asymmetry and repetitive indoor layouts, visual FLoc is fundamentally challenged by multimodal pose distributions, where visually identical observations map to distinct, spatially separated locations. Existing ray-matching-based methods tackle this by explicitly predicting sparse geometric or semantic rays, which inherently incur information loss and demand resource-intensive preprocessing alongside exhaustive matching during inference. In this paper, we bypass the intermediate ray-matching paradigm and propose a coarse-to-fine visual FLoc framework that progresses from uncertainty to determinism. In the coarse stage, we design an image-conditioned pose diffusion model to parameterize the continuous multimodal pose distribution, effectively routing stochastically initialized pose particles toward distinct candidate modes. In the refinement stage, we propose a localized refiner that predicts bounded sub-meter pose residuals from candidate-centered floorplan crops, where structural ambiguities are largely eliminated. Our method effectively balances global multi-hypothesis tracking and local sub-meter refinement without requiring any offline map preprocessing or test-time lookup tables. Comprehensive results on the S3D (full) and ZInD benchmarks demonstrate that our approach achieves state-of-the-art accuracy and robustness.
Shiyong Meng, Bolei Chen, Ping Zhong +4
School of Computer Science and Engineering, Central South University
Many public buildings provide floorplans with a "you are here" indicator to help visitors orient themselves. Floorplan localization seeks to computationally replicate this capability by determining where visual observations were captured within a floorplan. However, existing methods typically assume controlled small-scale environments and precise vectorized floorplans, limiting their ability to operate in large-scale buildings and rasterized floorplans. In this work, we present an approach for performing floorplan localization in the wild by grounding the task in a reconstructed 3D representation of the scene. Given an unconstrained image collection, our method reconstructs a gravity-aligned 3D scene and projects it into a 2D density map that serves as a floorplan proxy. Floorplan localization is then formulated as aligning this proxy with the input floorplan via a 2D similarity transform. To bridge the appearance gap between density maps and architectural floorplans, we adapt a 2D foundation model to learn cross-modal correspondences, introducing a fine-tuning scheme that encourages semantically aligned matches while preserving structural consistency. Extensive experiments demonstrate substantial improvements over prior methods, including in extremely sparse settings with as little as a single input image. Our code and data will be publicly available.
Junhyeong Cho, Ruojin Cai, Hadar Averbuch-Elor
Cornell University · Kempner Institute, Harvard University
Floorplans are compact, appearance-invariant maps ideal for indoor localization, yet existing methods rely on depth networks that are brittle in cluttered scenes. We propose GALoc, a geometry-first framework that replaces depth prediction with gravity-aligned wireframes that satisfy verticality and coplanarity by construction. Given monocular RGB, camera intrinsics, relative poses, and IMU orientation, GALoc constructs a linear constraint matrix encoding verticality and coplanarity, and finds the camera gauge minimizing its smallest singular value via global search. The rectified wireframes are projected into bird's-eye-view layouts through a closed-form, FOV-consistent transformation and matched against the floorplan via metric-free SE(2) search. We evaluate end-to-end on Structured3D, with calibrated noise on Gibson, and on real-world author-collected sequences. When sufficient wall geometry is visible, GALoc matches or outperforms depth-based baselines -- achieving 88% sequential localization success at 0.1m over 100-step sequences on Gibson vs the baseline's 68% -- while abstaining in structure-blind scenes.
Jeahn Han, Minji Kim, Jeongbin Sohn +3
GIST · KAIST · Zurich University of Applied Sciences