Global video geo-localization aims to infer the geographic location of a video worldwide, evaluating performance across four geographic hierarchies: city, state/province, country, and continent. Existing methods typically employ one-way uniform sampling to process video frames and train independent classifiers for each hierarchy, which leads to the loss of key geographic cues and prediction conflicts between hierarchies, especially for complex multi-shot edited videos. To address these limitations, we propose GeoGAT, which integrates bidirectional temporal sampling with graph attention networks (GATs). Specifically, GeoGAT extracts forward and offset-reversed frame sequences to construct complementary spatiotemporal features. These fused features are then fed into a predefined geographical hierarchy graph, where GATs perform structure-aware message passing, while a dual-constraint mechanism prunes predictions to eliminate cross-hierarchy conflicts. We construct GeoGAT10k, comprising 9,720 multi-shot edited videos from 166 cities worldwide, specifically to benchmark generalization ability on complex video structures. Experimental results on CityGuessr68k and GeoGAT10k demonstrate that GeoGAT eliminates hierarchical conflicts entirely and achieves state-of-the-art performance across all four geographic hierarchies. On CityGuessr68k, GeoGAT outperforms the strongest baseline, evaluated under both classification and retrieval protocols, by 2.6 percentage points at the city level. On the more challenging GeoGAT10k with multi-shot edited videos, the accuracy improvement exceeds 24 percentage points, validating strong generalization to complex real-world scenarios.
Figures & tables
Figure 1: Research motivation.
Figure 2: Overview of GeoGAT. (a) Bidirectional Temporal Sampling extracts complementary features from forward and offset-reversed frame sequences. (b) GAT-based Hierarchy Classifier propagates information over a directed geographic graph to model cross-granularity dependencies. (c) Multi-granularity Loss enforces geographic consistency through mutual information and hierarchical constraints.
Figure 3: Video frame samples from 10 different countries in the GeoGAT10k dataset.
Figure 4: Data distribution. GeoGAT10k covers most regions of the world and maintains a uniform spread across the globe, ensuring balanced geographic diversity.
Method
Venue
City
State
Country
Continent
Image-based
PlaNet
ECCV’16
55.8
56.3
60.8
74.1
ISNs
ECCV’18
59.5
59.9
64.1
75.9
GeoDecoder
CVPR’23
64.2
64.5
69.5
79.9
GeoCLIP
NeurIPS’23
57.8
60.5
75.9
90.8
GeoReasoner
ICML’24
38.5
42.8
64.4
81.9
Table 1: Comparison with state-of-the-art methods. VidTAG-cls † : visual encoder adapted as a classifier under our training protocol. VidTAG-ret ‡ : native retrieval paradigm, k -NN ( k=5 ) over training embeddings without GPS queries. GeoBayes ∗ is reproduced following the original paper. Best baseline results are underlined ; best overall results are in bold . Details are in the supplementary material.
Figure 5: Top-1 accuracy comparison and gain of GeoGAT over the best baseline across geographic hierarchies.
CityGuessr68k Top 1 Acc.(%)
GeoGAT10k Top 1 Acc.(%)
Configuration
Backbone
Sampling
Soft
Hard
LMI
City
State
Country
Continent
City
State
Country
Continent
Linear
Linear
OUS
∘
∘
∘
61.9
62.2
64.1
73.4
12.7
12.4
18.4
42.6
Linear + BTS
Linear
BTS
∘
∘
∘
57.8
58.0
62.4
73.4
12.0
11.9
16.7
47.5
w/ OUS
GATv2
OUS
∘
∘
∘
74.7
75.3
75.6
84.1
30.7
30.9
34.1
62.4
w/ BRS
GATv2
BRS
∘
∘
∘
82.2
82.6
82.5
88.2
41.8
41.9
47.7
70.5
w/ BTS
GATv2
BTS
∘
∘
∘
83.1
83.2
84.7
89.3
42.0
42.1
48.5
71.3
Table 2: Ablation study on CityGuessr68k and GeoGAT10k. ∘ : disabled; ✓ : enabled. OUS: one-way uniform sampling; BRS: bidirectional random sampling; BTS: bidirectional temporal sampling. The best results are in bold .
Figure 6: Normalized attention weight distribution across GAT layers. Cell (i,j) shows attention proportion from source hierarchy i (rows) to target j (columns). Diagonal cells reflect self-attention, with dominance growing coarser from city to continent. Off-diagonal warmth along adjacent levels indicates cross-hierarchy flow, while distant cells remain cold, revealing locality in geographic dependencies.
Figure 7: Sparse-sample region performance. Lines show Top-1 accuracy trends across geographic hierarchies; the right panel reports hierarchy conflict rates.
Worldwide image geo-localization aims to infer the geographic location of an image captured anywhere on Earth, spanning street, city, regional, national, and continental scales. Existing methods rely on visual features that are sensitive to environmental variations (e.g., lighting, season, and weather) and lack effective post-processing to filter outlier candidates, limiting localization accuracy. To address these limitations, we propose DualGeo, a two-stage framework for worldwide image geo-localization. First, it establishes a geo-representational foundation by fusing image and semantic segmentation features via bidirectional cross-attention. The fused features are then aligned with GPS coordinates through dual-view contrastive learning to build a global retrieval database. Second, it performs geo-cognitive refinement by re-ranking retrieved candidates using geographic clustering. It then feeds them into large multimodal models (LMMs) for final coordinate prediction. Experiments on IM2GPS, IM2GPS3k, and YFCC4k show that DualGeo outperforms state-of-the-art methods, improving street-level (<1 km) and city-level (<25 km) localization accuracy by 3.6%-16.58% and 1.29%-8.77%, respectively. Our code and datasets are available : https://github.com/CJ310177/DualGeo.
Junchao Cui, Wenqi Shi, Shaoyong Du +4
Henan Key Laboratory of Cyberspace Situation Awareness, Zhengzhou, China · Information Engineering University, Zhengzhou, China
Worldwide image geo-localization aims to determine where on Earth a single image was captured. However, visually similar scenes may lie thousands of kilometers apart, so methods that localize primarily by appearance often mistake a distant look-alike for the true location. We attribute this failure to a structural cause: in existing methods, GPS coordinates serve only as training supervision, and the distance relationships among locations never enter the learned representation. To address this, we propose GeoMetric, a retrieval-based framework that encodes GPS coordinates relationally rather than in isolation, injecting the distance structure among locations into both representation learning and inference. GeoMetric comprises three components: (1) a Transformer-based GPS encoder with distance-aware location attention that modulates inter-sample aggregation by great-circle proximity; (2) a trimodal contrastive objective that aligns images, geo-textual descriptions, and GPS embeddings in a unified space; and (3) a retrieval-augmented inference stage that supplies large multimodal models (LMMs) with contrastive candidate context for grounded coordinate reasoning. Extensive experiments on IM2GPS, IM2GPS3k, YFCC4k, and YFCC26k demonstrate that GeoMetric consistently outperforms state-of-the-art methods, improving street-level accuracy (within 1 km) by 1.5%, 0.9%, 6.9%, and 2.5%, respectively. Controlled ablations confirm that the gains originate from the proposed geographic encoding rather than any specific LMM.
Junchao Cui, Xuanzi Ma, Wenqi Shi +3
Information Engineering University Zhengzhou, China
Cross-view geo-localization (CVGL) retrieves geo-tagged satellite imagery for a ground-view query. Most systems exhaustively search a flat, fixed-resolution gallery, incurring high cost over large areas and adapting poorly to satellite resolution changes. Autoregressive coarse-to-fine alternatives reduce comparisons but bind later predictions to earlier decisions and a predefined hierarchy. We introduce GeoMoE, a sparse mixture-of-experts dual encoder that decouples global multi-scale representation learning from local hierarchical search. Global multi-scale supervision and content-adaptive routing map ground and satellite images across resolutions into a globally comparable embedding space. At inference, each image is encoded once, and probabilistic beam search follows parent--child links to score a small candidate subset. Later levels reuse these descriptors rather than features generated by preceding levels, limiting feature-level error propagation and hierarchy coupling. We further introduce VIGOR-M, a four-city benchmark with an explicit parent--child satellite hierarchy and held-out half-step galleries for single-resolution, cross-resolution, and hierarchical evaluation. GeoMoE achieves 95.78% R@40m on Just Zoom In, 2.77 percentage points above the previous best, and 62.39% R@1 on VIGOR-M. The latter requires 0.885 MMAC/query for descriptor matching, 5.27% of an exhaustive L3 scan, while exceeding the strongest exhaustive baseline by 3.12 percentage points in R@1. One model trained on L1, L2, and L3 also outperforms a matched dense control across all six galleries and transfers to three withheld resolutions. By decoupling globally trained embeddings from local hierarchical search, GeoMoE jointly improves localization accuracy, search efficiency, and cross-resolution transfer.
Ruijie Fan, Junyan Ye, Qi Zhu +1
Tsinghua Shenzhen International Graduate School, Tsinghua University · School of Geospatial Engineering and Science, Sun Yat-sen University