cs.CVJun 15, 2026

Training-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos

Authors: Ke LiDi WangYongshan ZhuTing WangWeiping NiTao LeiQuan WangXinbo Gao

Organizations: School of Computer Science and Technology, Xidian University, Xi’an 710071, China · Interdisciplinary Institute of Artificial Intelligence, Xidian University, Xi’an, Shaanxi 710126, China · School of Artificial Intelligence, Xidian University, Xi’an 710071, China · Northwest Institute of Nuclear Technology, Xi’an 710024, China · School of Physics and Information Engineering, Fuzhou University, Fuzhou 350108, China

Abstract

Remote sensing visual grounding (RSVG) aims to localize a referred target in a remote sensing image or video according to a natural language expression. Existing RSVG methods usually rely on task-specific manual annotations, which are costly to collect and inevitably limited in covering the diversity of real-world geospatial scenarios. As a result, they often struggle to generalize to open-vocabulary queries involving novel objects, fine-grained attributes, complex spatial relationships, and functional semantics. In this paper, we propose RSVG-ZeroOV, a training-free framework that leverages frozen generic foundation models for zero-shot open-vocabulary RSVG. RSVG-ZeroOV follows an Overview-Focus-Evolve paradigm, which exploits the distinct yet complementary attention patterns of vision-language models (VLMs) and diffusion models (DMs) to progressively generate precise grounding results. Specifically, (i) Overview utilizes a VLM to extract cross-attention maps that capture semantic correlations between the referring expression and visual regions; (ii) Focus leverages the fine-grained modeling priors of a DM to compensate for object structure and shape information often overlooked by VLM attention; and (iii) Evolve introduces a simple yet effective attention evolution module to suppress irrelevant activations, yielding purified object masks. To handle video inputs, we further present Video RSVG-ZeroOV, which extends image-level grounding to spatio-temporal grounding through a query-relevant key-frame selector and a temporal propagator, enabling efficient and temporally coherent video grounding without video annotations or fine-tuning. Extensive experiments on six image and video grounding benchmarks show that RSVG-ZeroOV consistently outperforms existing zero-shot baselines and achieves competitive or superior performance compared with weakly- and fully-supervised methods.

Explore similar work

CardsList