cs.CVOct 8, 2026

S3^3Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-Localization

Authors: Ziqian Mo, Hill Zhang, Haosheng Tan, Ling Li, Jiaheng Wei

Organizations: The Hong Kong University of Science and Technology (Guangzhou) Guangzhou, China · Claremont McKenna College Claremont, California, USA

Abstract

Cross-view geo-localization (CVGL) aims to estimate geographic locations by matching images captured from different viewpoints, such as drone and satellite views. Existing methods mainly rely on visual representations, but often fail to jointly model fine-grained structural correspondences and semantic priors, making them prone to confusion between visually similar but semantically different regions, and thus limiting robustness under large viewpoint variations. To address these challenges, we propose \textbf{S3^3Geo}, a structure-semantic synergistic learning framework for cross-view matching. Specifically, we first introduce a Decoupled Query Pooling (DQP) module to extract a compact set of region-aware features from dense tokens, enabling explicit modeling of local structural patterns. We then design a query-level contrastive learning scheme with an optimal transport (OT)-based formulation to establish soft correspondences under cross-view spatial misalignment. Furthermore, we incorporate a Semantic Knowledge Distillation (SKD) strategy from a frozen CLIP teacher to transfer semantic priors and relational structures, thereby improving discrimination on hard negatives. By operating synergistically, the semantic priors provide robust contextual filtering, which guides the structural module to establish precise spatial alignments. Experiments on the University-1652 and SUES-200 datasets demonstrate that \textbf{S3^3Geo} consistently outperforms state-of-the-art approaches without increasing inference complexity, validating the effectiveness of jointly modeling structural and semantic information for CVGL.

Explore similar work

CardsList
  1. Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization

    Jun 29, 2026Liyao Wang, Ruipu Wu, Haojun Xu +3Cross-View Geo-LocalizationCamera Pose Estimation

  2. BGG: Bridging the Geometric Gap between Cross-View images by Vision Foundation Model Adaptation for Geo-Localization

    May 11, 2026Wei Wang, Dou Quan, Ning Huyan +4Cross-View Geo-LocalizationGeospatial Representation Learning

  3. Warp-free Cross-view Geo-localization via Feature-space Consensus Mining

    Aug 10, 2026Zhuo Song, Lian Xu, Runqing Jiang +4Multi-View ConsistencyCross-View Geo-Localization