Text-to-point-cloud localization estimates a position in a city-scale 3D map from descriptions of surrounding objects. Existing coarse-to-fine methods retrieve submaps using aggregate learned compatibility and then localize within a selected submap. However, repetitive or similar urban objects can inflate the embedding similarity between the query and multiple submaps, even when the instance layout within a submap violates the query description. Meanwhile, query-relevant instances often span submap boundaries, leaving the retrieved submap with incomplete contextual evidence. We term these failure modes layout-inconsistent aliasing and boundary evidence incompleteness, respectively. To address them, we propose PARC-Loc, a coarse-to-fine localization framework built on Partial Assignment with Relational Consistency (PARC). PARC jointly models hint-object compatibility and pairwise spatial relations, allowing unmatched elements while favoring assignments consistent with the queried layout. At the coarse stage, its candidate-level assessment complements neural similarity for layout-consistent submap selection. At the fine stage, the context is expanded with query-relevant instances from adjacent submaps, while PARC yields object-level matching weights that guide cross-modal attention. Extensive experiments on KITTI360Pose and CityLoc show that PARC-Loc outperforms conventional coarse-to-fine baselines. On KITTI360Pose, our method improves Top-1 localization recall at 5 m from 0.50 to 0.67, achieving a 34% relative gain over the strongest baseline.
Figures & tables
Figure 1: Two failure modes in the standard coarse-to-fine decision structure.
Figure 2: Effect of added query-matching instances on candidate scores. Top: standardized score change; bottom: promotion rate among initially non-Top-1 candidates. Curves compare relation-satisfied and relation-violated placements; error bars denote 95% cluster-bootstrap confidence intervals.
(a) Prevalence of Boundary-Truncated Evidence
Analysis unit
Affected
Total
Rate
Boundary-truncated instances
7,907
19,629
40.28%
Hints referencing them
11,905
18,459
64.49%
Queries
3,179
3,187
99.75%
Table 1: Boundary-related evidence on the KITTI360Pose validation split. (a) Prevalence of boundary-truncated instances and affected hints and queries. (b) Exact Top-1 recall, Close@1 recall within 15m , and the partition of Top-1 retrieval errors into close non-GT and other submaps.
Figure 3: Overview of PARC-Loc. In the coarse stage, the relational consistency between hints and each submap is explicitly assessed through partial assignment to produce a reliable submap selection. In the fine stage, neighboring-object aggregation first makes cross-boundary context available; PARC is then reapplied to produce object weights that bias cross-attention before offset prediction.
Method
Venue
Localization Recall ( τ=5/10/15m ) ↑
Validation Set
Test Set
k=1
k=5
k=10
k=1
k=5
k=10
Text2Pos ( 2022 )
CVPR’22
0.14/0.25/0.31
0.36/0.55/0.61
0.48/0.68/0.74
0.13/0.21/0.25
0.33/0.48/0.52
0.43/0.61/0.65
RET ( 2023 )
AAAI’23
0.19/0.30/0.37
0.44/0.62/0.67
0.52/0.72/0.78
0.16/0.25/0.29
0.35/0.51/0.56
0.46/0.65/0.71
Text2Loc ( 2024 )
CVPR’24
0.37/0.57/0.63
0.68/0.85/0.87
0.77/0.91/0.93
0.33/0.48/0.52
0.61/0.75/0.78
0.71/0.84/0.86
CMMLoc ( 2025 )
CVPR’25
0.44/0.62/0.68
0.75/0.88/0.90
0.83/0.93/0.95
0.39/0.53/0.56
0.67/0.80/0.82
0.77/0.87/0.89
Table 2: Full localization results on KITTI360Pose. Each entry reports Top- k recall at 5/10/15m . Best results are in bold , and second-best results are underlined .
Method
Submap Retrieval Recall ↑
Validation Set
Test Set
k=1
k=3
k=5
k=1
k=3
k=5
Text2Pos ( 2022 )
0.14
0.28
0.37
0.12
0.25
0.33
Text2Loc ( 2024 )
0.32
0.56
0.67
0.28
0.49
0.58
CMMLoc ( 2025 )
0.35
0.61
0.73
0.32
0.53
0.63
PMSH ( 2025 )
0.37
0.63
0.73
0.34
0.56
0.65
Table 3: Coarse submap retrieval recall on KITTI360Pose. Best results are in bold , and second-best results are underlined .
Variant
Validation Set
Test Set
k=1
k=3
k=5
k=1
k=3
k=5
Neural retriever
0.49
0.75
0.83
0.46
0.70
0.78
PARC w/o unary
0.53
0.80
0.89
0.51
0.76
0.82
PARC w/o relation
0.52
0.79
0.87
0.50
0.74
0.82
PARC w/o partial
0.51
0.77
0.85
0.48
0.72
0.79
Full PARC
0.54
0.81
0.89
0.51
0.76
0.83
Table 4: Retrieval-level ablation of PARC on KITTI360Pose. All variants share our MNCL-style neural retriever. The full model is highlighted in bold .
Cumulative variant
Localization Recall ↑
Validation
Test
Neural pipeline
0.54/0.82/0.89
0.51/0.78/0.85
+ Coarse PARC
0.62/0.89/0.93
0.61/0.84/0.88
+ Neighbor context
0.66/0.90/0.93
0.64/0.85/0.89
+ Query-conditioned priority
0.69/0.90/0.93
0.66/0.85/0.90
Full PARC-Loc ( + bias)
0.71/0.91/0.94
0.67/0.87/0.90
Table 5: Cumulative end-to-end ablation on KITTI360Pose. Each entry reports Top-1/5/10 localization recall within 5m . The full model is highlighted in bold .
Context construction
@5m
@10m
@15m
Single selected submap
62.22
82.65
86.41
Five-cell context
67.74
84.00
86.57
Distance-prioritized 3×3
66.68
83.68
86.32
Table 6: Validation Top-1 localization recall (%) under controlled context construction. Each column gives the distance threshold; higher is better.
Candidate rank
β=0
β=0.3
Top-1
68.92±0.73
71.10±0.49
Top-3
84.96±0.51
85.90±0.79
Top-5
89.11±0.54
90.97±0.68
Top-10
92.15±0.46
93.95±0.50
Table 7: Matched-seed localization recall within 5m (%). Values are mean ± sample standard deviation over seeds 2027–2029 at epoch 3. Higher is better.
Figure 4: Qualitative examples of the two stage-specific mechanisms. Query boxes reproduce the relations while omitting the repeated subject “Pose is.” (a) PARC replaces the MNCL neural Top-1 distant alias with a candidate within 15m. (b) A close non-GT selected submap omits neighboring evidence; the 3×3 context reduces localization error. Red and green outlines mark incorrect and correct results, respectively. The yellow marker denotes the ground-truth pose, and the blue marker denotes the predicted pose.
Method
CityLoc-K Val.
CityLoc-K Test
R@5m
R@10m
R@15m
R@5m
R@10m
R@15m
Text2Pos
16.48
40.69
62.92
14.62
38.27
59.55
Text2Loc
18.91
45.26
64.28
17.97
41.22
61.50
MNCL
19.30
45.94
64.50
18.76
42.63
62.58
CMMLoc
20.77
48.65
67.89
21.71
46.67
66.00
VLM-Loc
36.23
63.66
77.77
35.91
63.81
76.79
Table 8: Localization on CityLoc-K. Recalls are reported as percentages under 5m , 10m , and 15m thresholds. Best results are in bold , and second-best results are underlined . VLM-Loc ( 2026 ) is included for reference and uses vision-language reasoning over BEV and scene-graph inputs, whereas the remaining methods follow the conventional T2P setting.
Method
R@5m ↑
R@10m ↑
R@15m ↑
Text2Pos
8.11
27.21
50.01
Text2Loc
9.45
29.44
51.17
MNCL
13.68
35.93
53.78
CMMLoc
11.68
34.79
54.71
VLM-Loc
21.37
49.12
68.26
PARC-Loc
16.56
48.96
73.57
Table 9: Zero-shot cross-domain localization on CityLoc-C. Recalls are reported as percentages. PARC-Loc is trained on CityLoc-K and evaluated on CityLoc-C. Best results are in bold , and second-best results are underlined .
Figure 5: Robustness to controlled gallery growth. (a) Gallery expansion. (b) Top-1 submap retrieval recall at each radius. (c) Recall retention relative to the 200m gallery.