cs.CVMar 25, 2026

OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding

Authors: Xiaoyu Tang, Jun Dong, Jintao Cheng, Rui Fan

Organizations: School of Electronics and Information Engineering, and Xingzhi College, South China Normal University, Foshan 528225, China · School of Data Science and Engineering, and Xingzhi College, South China Normal University, Shanwei, 516600, China · Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology, Hong Kong SAR, China · College of Electronics & Information Engineering, Shanghai Research Institute for Intelligent Autonomous Systems, the State Key Laboratory of Intelligent Autonomous Systems, and Frontiers Science Center for Intelligent Autonomous Systems, Tongji University, Shanghai 201804, China

Abstract

Remote sensing visual grounding (RSVG) aims to localize specific targets in remote sensing images using natural language expressions. However, existing methods are restricted to single-sensor domains, i.e., either optical or synthetic aperture radar (SAR), limiting their real-world applicability. In this paper, we introduce the Cross-Domain RSVG (CD-RSVG) task and construct OptSAR-RSVG, the first large-scale benchmark dataset for this setting. To tackle the challenges of cross-domain feature modeling, computational inefficiency, and fine-grained semantic discrimination, we propose OptiSAR-Net++. Our framework features a patch-level Low-Rank Adaptation Mixture of Experts (PL-MoE) for efficient cross-domain feature decoupling. To mitigate the substantial computational overhead of Transformer decoding frameworks, we adopt a CLIP-based contrastive paradigm and further incorporate dynamic adversarial negative sampling, thereby transforming generative regression into an efficient cross-modal matching process. Additionally, a text-guided dual-gate fusion module (TGDF-SSA) and a region-aware auxiliary head are introduced to enhance semantic-visual alignment and spatial modeling. Extensive experiments demonstrate that OptiSAR-Net++ achieves SOTA performance on both OptSAR-RSVG and DIOR-RSVG benchmarks, offering significant advantages in localization accuracy and efficiency. The model and dataset have been made publicly available at https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus.

Figures & tables

Explore similar work

CardsList
  1. VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation

    Sep 15, 2026Quanwei Liu, Tao Huang, Jiaqi Yang +1Remote Sensing Image SegmentationVision-Language Model Adaptation

  2. Training-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos

    Jun 15, 2026Ke Li, Di Wang, Yongshan Zhu +5Weak Visual GroundingVisual Grounding Benchmarks