How to Reduce Localization Ambiguity? Geometry-Semantic Constrained BEV Representation Learning for Satellite-Ground Localization
Authors: Junming Feng, Panwang Xia, Qiong Wu, Xudong Lu, Zeyu Jiao, Kun Lv, Zherong Wu, Yi Wan, +3 more
Organizations: The Hong Kong Polytechnic University, Hong Kong · Southern University of Science and Technology · Wuhan University, Wuhan, China · The Chinese University of Hong Kong, Hong Kong, China · Huawei Technologies Co., Ltd
Satellite-ground localization estimates the planar position and yaw orientation of a ground camera within a geo-referenced satellite image. Most recent methods map ground and satellite features into a shared bird's-eye-view (BEV) space and establish spatial correspondences. However, insufficient depth constraints can assign one ground feature to different distances along a viewing direction, creating geometric ambiguity in BEV feature placement. Similar appearances at different locations can also create descriptor matching ambiguity, while existing descriptor learning lacks explicit semantic supervision to distinguish them. We propose GeoSem-BEV, a geometry-semantic constrained BEV representation learning method. Radial depth supervision constrains distance assignment, and vertical height supervision constrains height aggregation. Shared explicit semantic supervision promotes consistent semantic predictions across views and helps distinguish locations with similar semantics. These constraints improve feature placement and descriptor discriminability, enhancing state-of-the-art BEV localization models. On VIGOR with unknown orientation, GeoSem-BEV reduces mean orientation error by 37.2% and 38.1% in the cross-area and same-area settings, respectively. The corresponding errors are reduced by 10.8% and 15.6% on DReSS-D. On KITTI-CVL, it reduces same-area mean orientation error by 26.8% under 10 degree orientation noise.
Figures & tables
Figure 1: Two ambiguities in BEV-space feature matching. Left: Insufficient radial-depth constraints leave BEV feature placement ambiguous along viewing rays. (a) Upper right: Similar appearances at different locations can produce ambiguous descriptors without explicit semantic supervision. (b) Lower right: Explicit semantic supervision helps distinguish them.
Figure 2: GeoSem-BEV with FG 2 as the base model. Geometric supervision constrains radial assignment and height aggregation of ground-view features. A shared semantic head supervises both descriptor fields, and same-class hard negatives penalize spatially incorrect matches.
Cross-area
Same-area
Localization (m) ↓
Orientation ( ∘ ) ↓
Localization (m) ↓
Orientation ( ∘ ) ↓
Method
Mean
Median
Mean
Median
Mean
Median
Mean
Median
SliceMatch
7.220
3.310
25.970
4.510
6.490
3.130
25.460
4.710
CCVPE
5.410
1.890
27.780
13.580
3.740
1.420
12.830
6.620
DenseFlow
7.670
3.670
17.630
2.940
4.970
1.900
11.200
1.590
GS-DenseFlow
5.647 (-26.38%)
2.344 (-36.13%)
15.466 (-12.27%)
2.157 (-26.63%)
4.677 (-5.90%)
1.906 (+0.32%)
9.340 (-16.61%)
1.589 (-0.06%)
Table 1: VIGOR unknown-orientation test results. Lower is better. GeoSem-BEV implementations are shown in black bold; best, second-best, and third-best distinct values within each area split are marked in red bold, blue bold, and black bold, respectively. Ties share the same rank. FG 2 reports original single-stage results; published baselines follow the cited comparisons. For each GeoSem-BEV implementation, percentages in parentheses after every metric report the relative change from its corresponding base model, computed as (EGS−Ebase)/Ebase×100% .
Method
Loc. mean ↓
Loc. median ↓
Ori. mean ↓
Ori. median ↓
R@1m/5 ∘ ↑
R@3m/10 ∘ ↑
R@5m/20 ∘ ↑
(m)
(m)
( ∘ )
( ∘ )
FG 2†
5.751
2.943
18.908
1.337
13.07
50.11
67.04
Sem-FG 2
6.339
3.335
22.740
1.397
11.06
45.25
62.76
Geo-FG 2
5.014
2.469
15.980
1.240
15.57
57.75
74.26
GS-FG 2
4.662
2.364
13.850
1.241
16.53
59.81
76.72
Table 2: Ablation on VIGOR cross-area with unknown orientation. All results use RANSAC. Best values in each column are bold. FG 2† denotes our reproduced FG 2 baseline.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Cross-area
Same-area
Localization (m) ↓
Orientation ( ∘ ) ↓
Localization (m) ↓
Orientation ( ∘ ) ↓
Mean
Median
Mean
Median
Mean
Median
Mean
Median
CCVPE
6.050
2.230
37.390
10.270
3.010
1.020
14.440
7.940
FG 2†
6.422
3.814
16.466
2.411
5.508
2.850
14.124
2.129
GS-FG 2
5.926 (-7.72%)
3.396 (-10.96%)
14.696 (-10.75%)
2.234 (-7.34%)
4.886 (-11.29%)
2.475 (-13.16%)
11.921 (-15.60%)
1.939 (-8.92%)
ViewBridge
4.220
2.370
12.580
2.600
2.940
1.490
8.470
1.620
Appendix
Table 3: DReSS-D test results under unknown orientation. All results are reported without RANSAC refinement, following the published DReSS-D comparison. CCVPE and ViewBridge values follow that comparison, and FG 2† denotes our re-implementation. Best, second-best, and third-best distinct values within each metric are marked in red bold, blue bold, and black bold, respectively; ties share the same rank. Percentages in parentheses after every metric report the relative change from the corresponding base model, computed as (EGS−Ebase)/Ebase×100% .
Area
Ori.
Method
Loc. (m) ↓
Ori. ( ∘ ) ↓
Ori. (%) ↑
Mean
Median
Mean
Median
R@1 ∘
R@5 ∘
Cross-area
±10∘
GGCVT
–
–
–
–
98.98
100.00
CCVPE
9.160
3.330
1.550
0.840
57.72
96.19
HC-Net
8.470
4.570
3.220
1.630
33.58
83.78
FG 2
7.310
4.150
3.620
2.370
23.03
77.84
GS-FG 2
7.026 (-3.89%)
4.372 (+5.35%)
3.485 (-3.73%)
2.003 (-15.49%)
28.07 (+21.88%)
80.48 (+3.39%)
Appendix
Table 4: KITTI-CVL test results. Orientation noise is sampled uniformly within ±10∘ during training and testing. Lateral and longitudinal recalls are omitted. Best, second-best, and third-best distinct values within each area and orientation setting are marked in red bold, blue bold, and black bold, respectively; ties share the same rank. For GS-FG 2 and GS-ViewBridge, percentages in parentheses after every metric report relative changes from the corresponding base model. A dash denotes an unreported source metric.
Figure 3: Qualitative visualization of geometry-constrained feature placement and explicitly supervised descriptor learning. In the left block, the orange overlay is the RGB content sampled from the selected ground-image region and projected onto the satellite images; the two columns of satellite images show the projection produced by FG 2† and Geo-FG 2 , respectively. The cyan + and arrow mark the ground-truth camera position and orientation used for alignment. In the right block, each heatmap shows the normalized descriptor-matching response for one ground-view BEV query over satellite locations: warmer colors indicate higher matching scores. The green triangle marks the ground-truth satellite location corresponding to the query cell, and the white star marks the highest-scoring predicted location; the reported peak error is the distance between these two markers.
Figure 4: Qualitative correspondence comparison on challenging VIGOR samples. Each row shows the same ground–satellite image pair for FG 2 (left) and GS-FG 2 (right). Lines visualize selected correspondences; green and red arrows indicate the ground-truth and predicted camera poses, respectively.
Method
Valid-cell ratio
Corr. error mean ↓ (m)
Corr. error median ↓ (m)
Corr.@1m ↑
Corr.@3m ↑
Corr.@5m ↑
FG 2†
0.726
9.095
3.662
9.29
41.20
60.40
Geo-FG 2
0.726
8.427
3.244
10.64
45.52
64.86
Appendix
Table 5: Ground-truth-aligned BEV correspondence localization on 5,000 VIGOR cross-area validation samples with unknown orientation. Correspondence errors are measured from the highest-scoring satellite-view BEV cell to the aligned location. Recalls are percentages. Best values are in bold.
Geometric differences between cross-view images, such as drone and satellite views, significantly increase the challenge of Cross-View Geo-Localization (CVGL), which aims to acquire the geolocation of images by image retrieval. To further enhance the CVGL performance, this paper proposes a parameter-efficient adaptation framework for bridging the geometric gap across images based on the vision foundation model (VFM) (e.g., DINOv3), termed BGG. BGG not only effectively leverages the general visual representations of VFM and captures the robust and consistent features from cross-view images, but also utilizes the generalization capabilities of the VFM, significantly improving the CVGL performance. It mainly contains a Multi-granularity Feature Enhancement Adapter (MFEA) and a Frequency-Aware Structural Aggregation (FASA) module. Specifically, MFEA enhances the scale adaptability and viewpoint robustness of features by multi-level dilated convolutions, effectively bridging the cross-view geometric gap with small training costs. Additionally, considering the [CLS] token lacks spatial details for precise image retrieval and localization, the FASA module modulates patch tokens in the frequency domain and performs adaptive aggregation for local structural feature enhancement. Finally, BGG fuses the enhanced local features with the [CLS] token for more accurate CVGL. Extensive experiments on University-1652 and SUES-200 datasets demonstrate that BGG has significant advantages over other methods and achieves state-of-the-art localization performance with low training costs.
Wei Wang, Dou Quan, Ning Huyan +4
Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education of China, Xidian University · Department of Automation, Tsinghua University, Beijing 100084, China · School of Telecommunications, Xidian University, Xi’an 710071, China
Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same scene across extreme viewpoint gaps. Cross-view localization naturally provides a promising pathway toward this ability, as it requires a model to align ground-view imagery with geo-referenced satellite-view imagery despite drastic appearance changes to estimate camera poses. Recent visual foundation models have made this long-standing localization problem increasingly feasible by providing rich 2D representations for cross-view matching. However, we argue that cross-view localization should not be viewed merely as 2D matching or pose estimation. In this work, we revisit cross-view localization as more than pose estimation and investigate how it can help the model develop consistent cross-view understanding under extreme viewpoint changes, including stable semantics, reliable structure, and transferable geometry. We identify three key limitations of existing methods that prevent them from achieving this. They usually lack explicit 3D grounding, rely on strict point-wise matching that can weaken semantic consistency, and learn from an absolute objective that provides limited guidance for geometric reasoning. To address these limitations, we propose CROSS, a unified cross-view localization framework built upon 3D-grounded alignment, structure-aware matching, and hypothesis ranking. This formulation makes structure learning an intrinsic requirement, encourages semantic representations to remain stable, and enables the model to acquire transferable geometry. Extensive experiments on the KITTI and VIGOR datasets show that CROSS achieves state-of-the-art performance in cross-view localization. More importantly, CROSS effectively learns stable semantics, reliable structure, and transferable geometry across extremely different viewpoints.
Mao Chen, Xiangkai Zhang, Zhiyong Liu +2
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, and also with the School of Artificial Intelligence, University of Chinese Academy of Sciences · Beijing Aerospace Control Center
Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imagery. Although existing methods often use geometric warping to expose co-visible cues, such transformations rely on restrictive spatial assumptions and inevitably introduce severe visual distortions under view-dependent visibility, yielding noisy supervision and fragile correspondences. To overcome this, we propose a novel joint-view consensus-guided learning framework that entirely bypasses explicit geometric warping. Instead of forcing rigid spatial alignment, we dynamically mine and adaptively strengthen a semantic consensus directly within the feature space. Specifically, an auxiliary joint-view pathway during training enables direct cross-view interaction, allowing each view to selectively aggregate corroborative evidence into a unified consensus representation. To resolve feature heterogeneity among the single- and joint-view streams, we introduce global pattern probes acting as a semantic dictionary to project divergent modalities into a strictly aligned metric space. Guided by a consensus-mediated contrastive objective, single-view embeddings are explicitly pulled toward the joint-view anchor during training, distilling this consensus-mining capability into the single-view encoders for robust retrieval at inference. Extensive experiments demonstrate that our method achieves state-of-the-art performance across four standard benchmarks, underscoring the importance of discovering cross-view semantic consensus for reliable geo-localization.
Zhuo Song, Lian Xu, Runqing Jiang +4
Sun Yat-sen University, Shenzhen, China · The University of Western Australia, Perth, Australia