Scene Coordinate Regression (SCR) has recently emerged as a promising approach for LiDAR-based localization, achieving accurate localization without requiring an explicit 3D map. Despite their effectiveness, existing SCR methods rely on scene classification-based global embedding that struggles to provide fine-grained discrimination among nearby locations. Moreover, their reliance on uniform sampling of local features during training assigns equal importance to all points, thereby inadvertently propagating features from dynamic objects or unstable regions and potentially degrading training stability. In this paper, we present ReLoc, a revamped SCR architecture that can effectively address these limitations. First, we redesign the global embedding module by combining learnable context tokens with a feature aggregator to capture richer and more discriminative scene context. Second, we introduce an attention-based local feature enhancement module to mitigate the impact of noisy local features while encouraging context-consistent structures, yielding more robust local feature representations. Experimental results on two large-scale outdoor datasets demonstrate that our approach achieves state-of-the-art accuracy over previous SCR-based methods while maintaining real-time inference performance.
Figures & tables
Figure 1 : Comparison of localization accuracy across recent LiDAR-based SCR methods . Our method achieves the lowest translation and rotation errors on the QEOxford and NCLT datasets while maintaining real-time inference capability of over 90Hz.
Figure 2 : Comparison of global embeddings. We evaluate the discriminative ability of global embeddings within the same cluster (colored black on left) by measuring their cosine similarity to a scene cluster center ( ⋆ ) on the 15-13-06-37 in the QEOxford dataset. Compared to classification logits-based embeddings (LightLoc [ 15 ] ) that lack fine-grained discrimination by maintaining erroneously high similarity for spatially distant samples, our proposed global embedding provides more discriminative information to the SCR model, yielding high cosine similarity for nearby scans while naturally decreasing as the spatial distance increases.
Figure 3 : Analysis of LightLoc [ 15 ] . We analyze the relationship between the regression error & inlier ratio and semantic meaning in the 15-13-06-37 sequence of the QEOxford. To categorize the query points according to semantic meaning, we employ a pretrained SphereFormer [ 32 ] . Features from the dynamic categories (e.g. vehicles, pedestrians) suffer from higher errors and lower inlier ratios, whereas features from static regions show the opposite trend.
Figure 4 : Overview of our ReLoc framework. We adopt a two-stage training strategy. In Stage 1, the global embedding module is trained using 6-DOF pose supervision and contrastive regularization to learn context-aware scene representations. The learned weights are then transferred to Stage 2, where the global embeddings are concatenated with enhanced features through the Local Feature Enhancement Module, and an MLP-based regressor is trained to predict scene coordinates.
Figure 5 : Architecture of the global embedding module. Learnable context tokens interact with voxel features via cross-attention to capture scene-aware context and are then aggregated by an MLP-Mixer to produce a compact global embedding for downstream regression.
Figure 6 : Attention behavior of the feature enhancement module. Red lines indicate the top-100 referenced features, indicating that during the self-attention, the query feature ( ⋆ ) primarily refers to features from geometrically consistent structures (e.g. walls).
Method
15-13-06-37
17-13-26-39
17-14-03-00
18-14-14-42
Avg. ( ↓ )
SGLoc [ 13 ]
1.79/1.67
1.81/1.76
1.33/1.59
1.19/1.39
1.53/1.60
LiSA [ 14 ]
0.94/1.10
1.17/1.21
0.84/1.15
0.85/1.11
0.95/1.14
LightLoc [ 15 ]
0.82/1.12
0.85/1.07
0.81/1.11
0.82 /1.16
0.83/1.12
GTRLoc [ 16 ]
0.77/1.02
0.77/1.01
0.67/1.01
0.80 / 1.07
0.75/1.03
ReLoc (ours)
0.73/0.87
0.64/0.83
0.57/0.85
0.80/0.94
0.69/0.87
Table I : Quantitative localization results on the QEOxford dataset. Mean position error (m) and mean orientation error ( ∘ ) are reported. Best and second results are in bold and underlined.
Figure 7 : Qualitative localization result on 17-13-26-39 of the QEOxford dataset. Position errors are visualized using a heatmap capped at 5m, where our method exhibits consistently lower errors (dots colored in blue ) along the trajectory compared to prior approaches.
Method
2012-02-12
2012-02-19
2012-03-31
2012-05-26
Avg. ( ↓ )
SGLoc [ 13 ]
1.20/3.08
1.20/3.05
1.12/3.28
3.48/4.43
1.75/3.46
LiSA [ 14 ]
0.97/ 2.23
0.91/ 2.09
0.87/ 2.21
3.11/ 2.72
1.47/ 2.31
LightLoc [ 15 ]
0.98/2.76
0.89/2.51
0.86/2.67
3.10/3.26
1.46/2.80
GTRLoc [ 16 ]
0.95 /2.53
0.82 /2.45
0.82 /2.52
3.01 /2.99
1.40 /2.62
ReLoc (ours)
0.88 / 2.25
0.76/2.01
0.65/2.18
2.46/2.66
1.19/2.28
Table II : Quantitative localization results on the NCLT dataset . Mean position error (m) and mean orientation error ( ∘ ) are reported. Best and second results are in bold and underlined.
Figure 8 : Qualitative localization results on 2012-02-12 of the NCLT dataset. Position errors are visualized as done in Fig. 7 . Our method shows the reduced position error (dots colored in blue ) along the trajectory compared to prior approaches.
Global Embedding
Local Feature Enhancement
QEOxford ( ↓ )
NCLT ( ↓ )
mean error(m/°)
mean error(m/°)
1.32/1.69
2.72/4.18
✓
0.74/0.94
1.25/2.50
✓
✓
0.69/0.87
1.19/2.28
Table III : Ablation study on the QEOxford and NCLT datasets.
Median regression error ( Lscr ) / Inlier ratio (m / %) ( ↓ / ↑ )
Class
w/o fE
w/ fE
Class
w/o fE
w/ fE
manmade
2.73/55.8
2.24/66.9
terrain
2.40/59.6
2.03/68.8
vegetation
2.09/69.8
1.81/78.4
sidewalk
2.44/61.5
2.04/71.9
car
3.53/41.4
3.07/48.1
truck
4.71/29.2
4.53/31.9
Table IV : Impact of the local feature enhancement module ( fE ) on the 17-13-26-39 sequence of the QEOxford dataset.
K
128 (default)
256
512
β
0.1 (default)
0.5
1.0
m/ ∘
0.69/0.87
0.70/0.89
0.75/0.88
m/ ∘
0.69/0.87
0.82/0.87
0.92/0.89
Table V : Ablation studies of hyperparameters ( K and β ) on the QEOxford dataset. We report mean rotation and translation error of all the test sequences.
Fujian Key Laboratory of Urban Intelligent Sensing and Computing, Xiamen University, Xiamen, China · School of Engineering Mathematics and Technology, University of Bristol, Bristol, United Kingdom