cs.CVOct 5, 2026

What Words Keep of a Place: Zero-Shot Language Reasoning for Cross-View Geo-Localization

Authors: Ayesh Abu Lehyeh, Jay Hwasung Jung, Safwan Wshah

Organizations: Department of Computer Science, University of Vermont

Abstract

Cross-view geo-localization is commonly solved as an image retrieval problem, matching a ground-level image against a database of satellite tiles through a jointly trained embedding. Such models are accurate, but they need large paired supervision and cannot show what evidence supports a match. In this paper, we study a different question: how much of this task can be solved through language alone? We prompt a multimodal large language model (MLLM) to describe each ground panorama and each satellite tile as structured text, and localize by comparing these descriptions. No component is trained. We evaluate on 9,826 VIGOR pairs from four U.S. cities, in three settings. First, the descriptions are faithful but not discriminative. They agree closely across the two views, yet ranking the full pool by description similarity almost never returns the correct tile (0.39% Recall@1). Second, we narrow the pool to ten neighboring tiles, as a coarse prior would do. The same descriptions now become useful: an MLLM judge that scores structural consistency doubles random ranking and matches a strong lexical baseline. It also states which fields of the two descriptions agree and which conflict, which an embedding distance cannot do, and which we see as a step toward interpretable localization. Third, we place the judge on a trained visual retriever. On the queries it ranks wrongly, reranking from images works, while reranking from our descriptions does not (23.5% against 10.7% Recall@1). Scene structure survives the conversion into language, while the fine appearance detail needed to separate nearby places does not. Code and prompts are publicly available at https://github.com/AyeshAbuLehyeh/GeoLingual.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GeoBridge++: Fact-Guided Geo-Semantic Bridging for Unified Cross-View Geo-Localization

    Dec 2, 2025Zixuan Song, Jing Zhang, Di Wang +6Cross-View Geo-LocalizationGeogs-Slam

  2. MAPS: Multi-Anchor Projection Similarity for Joint Vision-Language Geo-Localization

    Jun 21, 2026Yutong Hu, Siyuan Tan, Shaocheng Yan +3Cross-View Geo-LocalizationVision-Language Alignment

  3. ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models

    May 29, 2026Kaiwen Xue, Tao Wei, Guoxin Zhang +5Cross-View Geo-LocalizationMultimodal Large Language Models