GeoBridge++: Fact-Guided Geo-Semantic Bridging for Unified Cross-View Geo-Localization
Authors: Zixuan Song, Jing Zhang, Di Wang, Zhiming Luo, Wenbin Liu, Haonan Guo, En Wang, Bo Du, +1 more
Organizations: College of Computer Science and Technology and Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University, Changchun 130012, China · Zhongguancun Academy, Beijing 100094, China · School of Computer Science, Wuhan University, Wuhan, China · State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, China
Cross-view geo-localization infers a location by retrieving geo-tagged reference images matching a query image. However, the traditional satellite-centric paradigm limits robustness when high-resolution or up-to-date satellite imagery is unavailable and underexploits complementary cues across views and modalities. To address these challenges, we propose GeoBridge, a novel model that performs bidirectional matching across views and supports language-to-image retrieval. GeoBridge builds on a novel semantic-anchor mechanism that bridges multi-view features through textual descriptions for robust, flexible localization. We further extend GeoBridge to propose GeoBridge++, a fact-guided geo-semantic bridging framework incorporating real-world geographic knowledge to reduce the ambiguity and instability in appearance-dominated supervision. It integrates structured geographic attributes with visual observations to construct factual descriptions and applies targeted guidance based on modality-specific observable content, thereby enhancing geographic discriminability. GeoBridge++ exploits explicit spatial structures encoded by static maps to build a geo-semantic bridge that adaptively aggregates complementary multi-view information and promotes cross-view consistency. In support of this task, we further construct GeoLoc-MM, a million-scale, multi-view, and multi-scale dataset with aligned drone, satellite, street-view, and static-map imagery at six spatial extents per location, enabling systematic evaluation of arbitrary cross-view retrieval, scale robustness, and cross-view generalization. Extensive experiments show that GeoBridge supports robust cross-view and cross-modal geo-localization, while GeoBridge++ achieves consistent improvements across multiple benchmarks. Code and dataset will be released at https://github.com/MiliLab/GeoBridge.
Figures & tables
Fig. 1: Schematic diagram of GeoBridge and GeoBridge++. Cross-view geo-localization aims to match images with geo-referenced coordinates based on images, while cross-modal geo-localization aims to match images with geo-referenced coordinates based on language descriptions.
Fig. 2: Overall workflow of GeoBridge. A unified location description is generated from aligned drone, street-view panorama, and satellite images and serves as the semantic anchor. GeoBridge jointly performs image-image and text-image alignment to learn a shared representation space for cross-view and cross-modal geo-localization.
Fig. 3: Overall framework of GeoBridge++ with multi-level location fact guidance and geo-semantic bridging. Geographic, visual, and shared geo-visual facts provide modality-aware supervision for different view groups, while the geo-semantic bridge adaptively aggregates drone, street-view, satellite, and static-map representations into a shared location-level representation for unified cross-view alignment.
Fig. 4: Architecture of the geo-semantic bridge.
Dataset
Year
Platform
Region(s)
Size(k)
Multi-scale
GPS-tag
Altitudes
CVUSA [ 36 ]
2015
T+S
Nationwide (USA)
44.4 + 44.4
✗
✓
—
Tian et al. [ 40 ]
2017
G+A
Multi-city (USA)
35.4 + 35.4
✗
✓
—
CVACT [ 37 ]
2019
T+S
City-scale(Canberra, Australia)
128.3 + 128.3
✗
✓
—
VIGOR [ 39 ]
2021
T+A
Multi-city (USA)
105.2 + 90.6
✗
✓
—
University–1652 [ 15 ]
2020
D+G+S
Campus/City, multi-source
37.9 + 2.6 + 0.7
✓
✗
✗
DenseUAV [ 6 ]
2023
D+S
City-scale (Zhejiang, China)
9.1 + 31.7
✓
✗
✓
TABLE I: Comparison of cross-view geo-localization datasets. Platform : A = Aerial, D = Drone, G = Street, T = Street-view panorama, S = Satellite, M = Map. Multi-scale : multiple scales per location. GPS-tag : geo-coordinates provided. Altitudes : multiple drone heights.
Fig. 5: Multi-view data processing for the GeoLoc dataset. (a) Global distribution of multi-view image groups (darker shades indicate higher density). (b) Counts of drone images per ground-footprint bin as a function of covered area (m 2 ).
Method
Drone to Satellite
Satellite to Drone
R@1
AP
R@1
AP
SAIG-D [ 46 ]
78.85
81.62
86.45
78.48
DWDR [ 47 ]
86.41
88.41
91.30
86.02
MBF [ 48 ]
89.05
90.61
92.15
84.45
MCCG [ 49 ]
89.64
91.32
94.30
89.39
SeGCN [ 50 ]
89.18
90.89
94.29
89.65
TABLE II: Comparison on the University–1652 [ 15 ] dataset; best results are in bold, second-best are underlined.
Drone to Satellite
Method
150m
200m
250m
300m
R@1
AP
R@1
AP
R@1
AP
R@1
AP
MBF [ 48 ]
85.62
88.21
87.43
90.02
90.65
92.53
92.12
93.63
MCCG [ 49 ]
82.22
85.47
89.38
91.41
93.82
95.04
95.07
96.20
CCR [ 51 ]
87.08
89.55
93.57
94.90
95.42
96.28
96.82
97.39
SeGCN [ 50 ]
90.80
92.32
91.93
93.41
92.53
93.90
93.33
94.61
TABLE III: Comparison on the SUES–200 [ 42 ] dataset; best results are in bold, second-best are underlined.
Method
CVUSA
VIGOR-Same
VIGOR-Cross
R@1
R@1%
R@1
Hit
R@1
Hit
TransGeo [ 8 ]
94.08
99.77
61.48
73.09
18.99
21.21
FRGeo [ 55 ]
97.06
99.85
71.26
82.41
37.54
40.66
SAIG-D [ 46 ]
96.08
99.86
65.23
74.11
33.05
36.71
VimGeo [ 56 ]
96.19
99.52
55.24
57.43
19.31
20.72
Sample4Geo [ 29 ]
98.68
99.87
77.86
89.82
61.70
69.87
TABLE IV: Comparison on the CVUSA [ 36 ] and VIGOR [ 39 ] datasets for street-to-satellite retrieval; best results are shown in bold, second-best are underlined.
Method
Regional val
Regional test
R@1
R@5
R@10
R@1%
R@1
R@5
R@10
R@1%
GeoDTR [ 58 ]
47.42
68.43
78.11
99.05
25.43
42.07
51.66
84.74
SAIG-D [ 46 ]
71.92
92.83
96.05
99.81
34.72
61.53
71.08
91.47
Sample4Geo [ 29 ]
97.20
99.43
99.69
99.93
84.66
92.66
94.42
98.10
Panorama-BEV [ 45 ]
97.78
99.63
99.79
99.93
85.68
92.91
94.77
98.21
GeoBridge (ours)
97.38
99.64
99.85
99.93
86.14
94.45
96.95
98.46
TABLE V: Comparison on the CVGlobal [ 45 ] dataset for street-to-satellite retrieval; best results are shown in bold, second-best are underlined.
Method
Regional val
Regional test
R@1
R@5
R@10
R@1%
R@1
R@5
R@10
R@1%
GeoDTR [ 58 ]
11.95
24.33
34.07
89.32
4.57
10.47
16.07
53.45
SAIG-D [ 46 ]
37.27
70.61
80.01
97.53
4.64
14.80
22.16
61.55
Sample4Geo [ 29 ]
75.38
92.80
95.67
99.54
46.94
70.46
77.91
91.47
Panorama-BEV [ 45 ]
81.31
95.84
97.68
99.71
49.41
71.04
78.33
91.96
GeoBridge++ (ours)
83.05
96.98
98.74
99.84
58.96
85.78
88.29
93.35
TABLE VI: Comparison on the CVGlobal [ 45 ] dataset for street-to-map retrieval; best results are shown in bold, second-best are underlined.
Method
D2S
S2D
D2T
T2D
R@1
AP
R@1
AP
R@1
AP
R@1
AP
Sample4Geo [ 29 ]
27.27
39.69
28.70
40.32
29.51
31.17
15.56
29.68
MEAN [ 52 ]
21.52
27.08
21.38
26.97
13.08
17.76
1.87
7.74
DAC [ 53 ]
6.19
8.41
13.91
15.16
13.74
15.16
19.34
23.01
CAMP [ 59 ]
19.60
24.39
14.88
18.75
11.31
12.34
11.46
19.16
MCCG [ 49 ]
12.23
13.22
14.73
17.70
15.51
19.11
12.90
15.75
TABLE VII: Comparison on the GeoLoc dataset. D, T, and S denote Drone, Street-View Panorama, and Satellite, respectively. The best results are shown in bold, the second-best results are underlined.
Method
Scale
D2S
S2D
D2T
T2D
D2M
M2D
S2M
M2S
R@1
AP
R@1
AP
R@1
AP
R@1
AP
R@1
AP
R@1
AP
R@1
AP
R@1
AP
Sample4Geo [ 29 ]
80
25.95
39.17
24.20
37.61
27.51
33.07
20.27
34.82
14.61
28.05
14.33
27.46
20.53
35.67
18.89
33.75
MEAN [ 52 ]
23.26
30.32
21.44
30.96
13.06
17.25
2.89
7.34
12.61
24.16
12.18
24.09
12.74
24.18
13.11
25.28
CAMP [ 59 ]
20.50
28.24
17.61
22.87
12.31
14.25
13.12
20.71
17.26
22.92
17.32
22.10
13.24
18.92
12.90
18.32
MCCG [ 49 ]
14.38
18.91
13.60
17.78
13.77
19.14
13.96
19.78
16.46
29.50
17.52
30.71
15.51
27.46
15.71
29.03
GeoBridge++(ours)
55.01
69.33
55.69
70.09
58.88
73.02
57.44
71.74
51.46
66.54
47.72
63.31
50.58
66.33
46.51
63.04
TABLE VIII: Comparison on the GeoLoc-MM dataset with drone-to-satellite methods extended to other cross-view combinations. D, T, S, and M denote Drone, Street-View Panorama, Satellite, and Static Map, respectively. The best results are shown in bold, the second-best results are underlined.
Fig. 6: Qualitative results of GeoBridge on GeoLoc. Blue boxes denote queries, and red boxes indicate true matches.
Method
Scale
T2S
S2T
D2T
T2D
T2M
M2T
S2M
M2S
R@1
AP
R@1
AP
R@1
AP
R@1
AP
R@1
AP
R@1
AP
R@1
AP
R@1
AP
Panorama-BEV [ 45 ]
80
12.24
15.91
20.41
23.01
18.18
22.18
16.33
20.62
10.99
13.79
10.85
13.52
14.48
21.01
14.35
20.46
Sample4Geo [ 29 ]
16.70
20.48
15.52
21.15
27.51
33.07
20.27
34.82
13.13
26.16
11.15
23.23
20.53
35.67
18.89
33.75
AuxGeo [ 57 ]
15.85
22.63
14.76
20.86
7.51
11.85
10.55
15.85
15.17
18.41
11.39
24.11
12.47
16.86
12.91
14.13
FRGeo [ 55 ]
18.79
24.47
16.74
21.04
15.33
21.98
16.35
17.07
6.31
15.07
6.20
14.62
13.39
19.68
18.98
25.11
GeoBridge++(ours)
50.19
65.80
51.87
67.23
58.88
73.02
57.44
71.74
51.56
66.58
48.29
63.85
50.58
66.33
46.51
63.04
TABLE IX: Comparison on the GeoLoc-MM dataset with street-to-satellite methods extended to other cross-view combinations D, T, S, and M denote Drone, Street-View Panorama, Satellite, and Static Map, respectively. The best results are shown in bold, the second-best results are underlined.
Fig. 7: Qualitative results of GeoBridge++ on GeoLoc-MM. Blue boxes denote queries, and red boxes indicate truth matches. The six rows from top to bottom correspond to spatial extents ranging from 80 to 200 m .
Method
R@1
R@5
R@10
ViLT [ 64 ]
2.00
16.00
38.00
BLIP-B [ 33 ]
0.00
1.00
2.00
EVA2-CLIP-B/16 [ 65 ]
15.00
38.00
54.00
EVA2-CLIP-L/14 [ 65 ]
19.00
43.00
55.00
CLIP-L/14 [ 61 ]
25.00
60.00
74.00
CLIP-B/16 [ 61 ]
18.00
51.00
61.00
TABLE X: Comparison results on the RSIEval [ 63 ] dataset. The best results are shown in bold, the second-best results are underlined.
Fig. 8: Qualitative cross-modal geo-localization results of GeoBridge on GeoLoc. Using street view descriptions to match drone perspectives, the top three results are reported; red boxes indicate correct matches.
Fig. 9: Qualitative cross-modal geo-localization results of GeoBridge++ on map-augmented GeoLoc. Descriptions generated from map observations are used to retrieve images from street-view, satellite, and drone perspectives. The top-ranked retrieved results are shown, with correct matches highlighted in red.
Method
Street Description
Satellite Description
Drone Description
Satellite Image
Drone Image
Street Image
Drone Image
Street Image
Satellite Image
R@1
L@50
R@1
L@50
R@1
L@50
R@1
L@50
R@1
L@50
R@1
L@50
ViLT [ 64 ]
1.40
6.41
1.05
5.72
6.3
13.21
1.05
5.96
6.20
13.49
1.16
5.76
BLIP-B [ 33 ]
0.84
4.6
0.78
5.42
0.93
5.06
0.73
5.59
0.82
5.14
0.95
5.64
EVA2-CLIP-B/16 [ 65 ]
1.01
5.28
0.92
4.30
1.21
5.74
0.90
4.26
1.18
5.63
0.99
5.23
EVA2-CLIP-L/14 [ 65 ]
1.16
4.97
1.83
6.95
1.23
5.57
1.23
5.42
1.21
5.55
1.18
4.88
TABLE XI: Comparison of cross-modal geo-localization methods on the GeoLoc dataset. Descriptions are generated from a single view; comparative results include two additional viewpoints. Best results are in bold; second-best are underlined.
Desc.
Target
Metric
ViLT [ 64 ]
CLIP -L/14 [ 61 ]
CLIP -B/16 [ 61 ]
CrossText 2Loc [ 21 ]
GeoBridge++ (Ours)
Map
Street
R@1
7.41
12.37
8.54
13.53
23.48
L@50
14.57
22.84
16.55
23.24
27.49
Satellite
R@1
1.73
2.65
2.33
2.60
8.99
L@50
6.63
7.30
7.15
7.55
10.60
Drone
R@1
1.49
2.57
1.95
2.44
9.20
L@50
5.47
8.00
7.65
8.48
15.66
TABLE XII: Comparison of GeoBridge++ with cross-modal geo-localization methods on the map-augmented GeoLoc dataset. Desc. denotes descriptions generated from a single source view, with retrieval performed over images from other target views. The best results are shown in bold, and the second-best results are underlined.
Method
D2S
S2D
T2S
S2T
D2T
T2D
Image-only
38.20
34.63
6.43
6.95
7.16
4.90
Text-only
42.83
42.83
35.40
36.40
39.00
38.63
GeoBridge
45.06
44.81
38.87
39.21
41.23
41.15
TABLE XIII: Ablation R@1 results of GeoBridge with different alignment strategies on the GeoLoc dataset. D, T, and S denote Drone, Street-View Panorama, and Satellite, respectively. The best results are shown in bold, the second-best results are underlined.
Fact
Bridge
Pairwise
D2S
S2D
T2S
S2T
D2T
T2D
D2M
M2D
S2M
M2S
T2M
M2T
✓
52.19
53.46
43.99
41.31
44.51
44.96
45.63
43.00
43.43
37.71
27.92
28.09
✓
25.83
25.83
7.14
6.52
6.69
7.59
13.87
14.11
17.44
16.71
4.34
3.66
✓
21.94
20.31
1.18
0.64
0.58
0.64
16.58
16.91
14.46
16.03
0.60
15.61
✓
✓
42.53
44.33
57.72
25.32
25.49
29.88
29.13
27.99
35.54
30.93
24.29
18.78
✓
✓
58.01
57.88
49.69
45.84
55.86
59.24
49.09
44.22
46.10
45.45
39.58
36.72
TABLE XIV: Ablation R@1 results of GeoBridge++ with different strategies on the map-augmented GeoLoc dataset. Fact, Bridge, Pairwise denote location fact guidance, geo-semantic bridge alignment, and direct pairwise image alignment, respectively. D, T, S, and M denote Drone, Street-View Panorama, Satellite, and Static Map, respectively. The best results are shown in bold, the second-best results are underlined.
G
V
SH
D2S
S2D
T2S
S2T
D2T
T2D
D2M
M2D
S2M
M2S
T2M
M2T
✓
46.81
49.47
–
–
–
–
40.46
38.91
35.86
33.90
–
–
✓
52.19
51.15
39.49
39.60
42.55
42.18
–
–
–
–
–
–
✓
47.39
51.65
25.58
26.57
27.92
28.69
37.08
39.92
38.24
36.95
21.70
22.93
✓
✓
51.99
53.30
17.72
29.47
21.53
19.59
31.17
25.21
26.69
22.31
8.56
7.92
✓
✓
✓
58.01
57.88
49.69
45.84
55.86
59.24
49.09
44.22
46.10
45.45
39.58
36.72
TABLE XV: Ablation R@1 results of GeoBridge++ with different types of location fact guidance on the map-augmented GeoLoc dataset. G, V, and SH denote geographic, visual, and shared geo-visual facts, respectively. D, T, S, and M denote Drone, Street-View Panorama, Satellite, and Static Map, respectively. The best results are shown in bold, the second-best results are underlined.
TABLE XVI: Further validation of the text description for GeoBridge on the GeoLoc dataset. D, T, and S denote Drone, Street-View Panorama, and Satellite, respectively. The best results are shown in bold, the second-best results are underlined.
Method
D2S
S2D
T2S
S2T
D2T
T2D
R@1
AP
R@1
AP
R@1
AP
R@1
AP
R@1
AP
R@1
AP
Sample4Geo [ 29 ]
13.70
22.22
7.40
16.34
7.41
14.81
8.64
12.35
13.58
17.68
12.35
17.37
MEAN [ 52 ]
12.35
17.37
13.58
19.60
-
-
-
-
4.41
7.27
5.23
8.94
MCCG [ 49 ]
10.27
14.22
14.17
17.89
-
-
-
-
2.63
6.31
5.58
7.16
panorama-BEV [ 45 ]
-
-
-
-
4.94
8.64
11.11
17.28
6.17
14.81
9.88
15.60
AuxGeo [ 57 ]
-
-
-
-
3.07
18.93
7.48
21.80
13.74
14.95
10.00
12.49
Appendix
TABLE XVII: Comparison on the CVUSA subdataset. D, T, and S denote Drone, Street-View Panorama, and Satellite, respectively. The best results are shown in bold, the second-best results are underlined.
Method
D2S
S2D
T2S
S2T
D2T
T2D
Image-only
43.76
47.83
14.12
15.98
16.26
12.25
Text-only
46.49
46.49
38.70
39.50
41.21
41.05
GeoBridge
49.05
48.76
42.10
41.96
43.54
43.41
Appendix
TABLE XVIII: Ablation AP results of GeoBridge with different alignment strategies on GeoLoc dataset. D, T, and S denote Drone, Street-View Panorama, and Satellite, respectively. The best results are shown in bold, the second-best results are underlined.
Fact
Bridge
Pairwise
D2S
S2D
T2S
S2T
D2T
T2D
D2M
M2D
S2M
M2S
T2M
M2T
✓
62.42
62.92
59.05
57.79
59.78
60.41
61.17
58.00
59.35
53.46
41.87
41.44
✓
39.18
39.75
16.44
14.79
14.69
16.47
26.38
26.41
31.18
29.81
11.20
9.84
✓
35.69
33.93
3.38
2.52
2.47
2.69
29.48
30.26
26.99
28.84
2.53
23.35
✓
✓
57.72
59.59
42.46
40.88
41.61
46.03
45.83
44.45
26.15
48.06
39.56
34.46
✓
✓
71.01
70.66
64.69
61.35
69.29
71.64
63.21
58.31
60.67
59.85
55.23
52.62
Appendix
TABLE XIX: Ablation AP results of GeoBridge++ with different strategies on the map-augmented GeoLoc dataset. Fact, Bridge, Pairwise denote location fact guidance, geo-semantic bridge alignment, and direct pairwise image alignment, respectively. D, T, S, and M denote Drone, Street-View Panorama, Satellite, and Static Map, respectively. The best results are shown in bold, the second-best results are underlined.
G
V
SH
D2S
S2D
T2S
S2T
D2T
T2D
D2M
M2D
S2M
M2S
T2M
M2T
✓
60.88
63.17
–
–
–
–
55.45
54.62
51.47
49.11
–
–
✓
64.71
63.47
54.72
55.52
57.29
57.79
–
–
–
–
–
–
✓
61.54
65.22
39.85
41.69
42.44
43.25
51.83
55.15
53.25
52.02
35.05
37.40
✓
✓
65.18
66.53
30.49
44.79
35.96
34.00
46.61
40.17
41.90
36.35
18.85
18.77
✓
✓
✓
71.01
70.66
64.69
61.35
69.29
71.64
63.21
58.31
60.67
59.85
55.23
52.62
Appendix
TABLE XX: Ablation AP results of GeoBridge++ with different types of location fact guidance on the map-augmented GeoLoc dataset. G, V, and SH denote geographic, visual, and shared geo-visual facts, respectively. D, T, S, and M denote Drone, Street-View Panorama, Satellite, and Static Map, respectively. The best results are shown in bold, the second-best results are underlined.
Fig. 10: Examples of original drone images.
Fig. 11: Examples of basic validity screening.
Fig. 12: Examples of blurry drone subimages.
Fig. 13: Examples of low global-contrast drone subimages
Fig. 14: Examples of uniform-texture and noisy pseudo-texture drone subimages
Fig. 15: Examples of aligned tri-view images.
Fig. 16: Example of multi-view and multi-scale observations in GeoLoc-MM.
Fig. 17: Tri-View instruction protocol for generating unified semantic descriptions. The blue text box denotes the instruction prompt. The first column presents the street-view panorama, the second column shows the drone-view image, the third column displays the satellite image, and the fourth column contains the generated textual description.
Fig. 18: Cross-modal geo-location drone image description instructions, with blue text boxes indicating the prompts.
Fig. 19: Cross-modal geo-location street-panorama image description instructions, with blue text boxes indicating the prompts.
Fig. 20: Cross-modal geo-location satellite image description instructions, with blue text boxes indicating the prompts.
Fig. 21: Examples of geographic fact construction from OSM data. Yellow boxes indicate the LLM instruction, red boxes show the extracted geographic facts, and blue boxes present the corresponding geographic factual descriptions.
Fig. 22: Examples of visual fact construction from cross-views. Yellow boxes indicate the LLM instructions, and blue boxes present the corresponding visual factual descriptions. The first column presents the drone-view image, the second column shows the satellite image, and the third column displays the street-view panorama.
Fig. 23: Examples of shared geo-visual fact construction. Yellow boxes indicate the LLM instruction, red boxes show the extracted shared geo-visual facts, and blue boxes present the corresponding shared geo-visual factual descriptions. The four images correspond to the drone, satellite, static-map, and street-view panorama views, respectively.
Fig. 24: Qualitative results of GeoBridge on the GeoLoc dataset for cross-modal geo-location. Using street view descriptions to match satellite perspectives, the top three results are reported; red boxes indicate correct matches.
Fig. 25: Qualitative results of GeoBridge on the GeoLoc dataset for cross-modal geo-location. Using satellite view descriptions to match drone perspectives, the top three results are reported; red boxes indicate correct matches.
Fig. 26: Qualitative results of GeoBridge on the GeoLoc dataset for cross-modal geo-location. Using drone view descriptions to match street perspectives, the top three results are reported; red boxes indicate correct matches.
Fig. 27: Qualitative results of GeoBridge++ on GeoLoc-MM. Blue boxes denote queries, and red boxes indicate truth matches. The six rows from top to bottom correspond to spatial extents ranging from 80 to 200 m .
Fig. 28: Qualitative results of GeoBridge++ on GeoLoc-MM. Blue boxes denote queries, and red boxes indicate truth matches. The six rows from top to bottom correspond to spatial extents ranging from 80 to 200 m .
Fig. 29: Qualitative results of GeoBridge++ on GeoLoc-MM. Blue boxes denote queries, and red boxes indicate truth matches. The six rows from top to bottom correspond to spatial extents ranging from 80 to 200 m .
Fig. 30: Qualitative results of GeoBridge++ on GeoLoc-MM. Blue boxes denote queries, and red boxes indicate truth matches. The six rows from top to bottom correspond to spatial extents ranging from 80 to 200 m .
Fig. 31: Qualitative cross-modal geo-localization results of GeoBridge++ on map-augmented GeoLoc. Descriptions generated from street-view observations are used to retrieve images from drone, satellite, and map perspectives. The top-three retrieved results are shown, with correct matches highlighted in red.
Fig. 32: Qualitative cross-modal geo-localization results of GeoBridge++ on map-augmented GeoLoc. Descriptions generated from satellite observations are used to retrieve images from drone, street-view, and map perspectives. The top-ranked retrieved results are shown, with correct matches highlighted in red.
Fig. 33: Qualitative cross-modal geo-localization results of GeoBridge++ on map-augmented GeoLoc. Descriptions generated from drone observations are used to retrieve images from map, satellite, and street-view perspectives. The top-ranked retrieved results are shown, with correct matches highlighted in red.
Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.g., ground or drone) within a geo-tagged reference image (e.g., satellite). Existing approaches heavily rely on 2D appearance matching and are constrained by limited datasets lacking geometric metadata, diverse prompts, and standard field-of-view imagery. To address these intertwined challenges, we first introduce \dataset, a large-scale, high-fidelity building dataset comprising over 220,000 ground-satellite and drone-satellite pairs. It provides multi-modal prompts (points, boxes, masks) and camera poses to enable flexible target referring and explicit spatial modeling. Furthermore, we propose a novel single-stage Geometry-Aware Geo-localization framework (GAGeo), built upon the permutation-equivariant 3D foundation model π3. By seamlessly integrating visual features, referring prompts, and learnable task tokens, our model adapts the inherited 3D prior to jointly predict bounding boxes, segmentation masks, and camera poses in a single forward pass. Additionally, we introduce a contrastive loss that utilizes the satellite view as a universal anchor, implicitly aligning ground and drone representations to enable zero-shot ground-to-drone localization without requiring triplet training data. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods, exhibiting exceptional generalization ability in unseen scenes and novel cross-view setups.
Liyao Wang, Ruipu Wu, Haojun Xu +3
Beihang University, Beijing, China · Meituan, Beijing, China
Geometric differences between cross-view images, such as drone and satellite views, significantly increase the challenge of Cross-View Geo-Localization (CVGL), which aims to acquire the geolocation of images by image retrieval. To further enhance the CVGL performance, this paper proposes a parameter-efficient adaptation framework for bridging the geometric gap across images based on the vision foundation model (VFM) (e.g., DINOv3), termed BGG. BGG not only effectively leverages the general visual representations of VFM and captures the robust and consistent features from cross-view images, but also utilizes the generalization capabilities of the VFM, significantly improving the CVGL performance. It mainly contains a Multi-granularity Feature Enhancement Adapter (MFEA) and a Frequency-Aware Structural Aggregation (FASA) module. Specifically, MFEA enhances the scale adaptability and viewpoint robustness of features by multi-level dilated convolutions, effectively bridging the cross-view geometric gap with small training costs. Additionally, considering the [CLS] token lacks spatial details for precise image retrieval and localization, the FASA module modulates patch tokens in the frequency domain and performs adaptive aggregation for local structural feature enhancement. Finally, BGG fuses the enhanced local features with the [CLS] token for more accurate CVGL. Extensive experiments on University-1652 and SUES-200 datasets demonstrate that BGG has significant advantages over other methods and achieves state-of-the-art localization performance with low training costs.
Wei Wang, Dou Quan, Ning Huyan +4
Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education of China, Xidian University · Department of Automation, Tsinghua University, Beijing 100084, China · School of Telecommunications, Xidian University, Xi’an 710071, China
Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imagery. Although existing methods often use geometric warping to expose co-visible cues, such transformations rely on restrictive spatial assumptions and inevitably introduce severe visual distortions under view-dependent visibility, yielding noisy supervision and fragile correspondences. To overcome this, we propose a novel joint-view consensus-guided learning framework that entirely bypasses explicit geometric warping. Instead of forcing rigid spatial alignment, we dynamically mine and adaptively strengthen a semantic consensus directly within the feature space. Specifically, an auxiliary joint-view pathway during training enables direct cross-view interaction, allowing each view to selectively aggregate corroborative evidence into a unified consensus representation. To resolve feature heterogeneity among the single- and joint-view streams, we introduce global pattern probes acting as a semantic dictionary to project divergent modalities into a strictly aligned metric space. Guided by a consensus-mediated contrastive objective, single-view embeddings are explicitly pulled toward the joint-view anchor during training, distilling this consensus-mining capability into the single-view encoders for robust retrieval at inference. Extensive experiments demonstrate that our method achieves state-of-the-art performance across four standard benchmarks, underscoring the importance of discovering cross-view semantic consensus for reliable geo-localization.
Zhuo Song, Lian Xu, Runqing Jiang +4
Sun Yat-sen University, Shenzhen, China · The University of Western Australia, Perth, Australia