OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding
Authors: Xiaoyu Tang, Jun Dong, Jintao Cheng, Rui Fan
Organizations: School of Electronics and Information Engineering, and Xingzhi College, South China Normal University, Foshan 528225, China · School of Data Science and Engineering, and Xingzhi College, South China Normal University, Shanwei, 516600, China · Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology, Hong Kong SAR, China · College of Electronics & Information Engineering, Shanghai Research Institute for Intelligent Autonomous Systems, the State Key Laboratory of Intelligent Autonomous Systems, and Frontiers Science Center for Intelligent Autonomous Systems, Tongji University, Shanghai 201804, China
Remote sensing visual grounding (RSVG) aims to localize specific targets in remote sensing images using natural language expressions. However, existing methods are restricted to single-sensor domains, i.e., either optical or synthetic aperture radar (SAR), limiting their real-world applicability. In this paper, we introduce the Cross-Domain RSVG (CD-RSVG) task and construct OptSAR-RSVG, the first large-scale benchmark dataset for this setting. To tackle the challenges of cross-domain feature modeling, computational inefficiency, and fine-grained semantic discrimination, we propose OptiSAR-Net++. Our framework features a patch-level Low-Rank Adaptation Mixture of Experts (PL-MoE) for efficient cross-domain feature decoupling. To mitigate the substantial computational overhead of Transformer decoding frameworks, we adopt a CLIP-based contrastive paradigm and further incorporate dynamic adversarial negative sampling, thereby transforming generative regression into an efficient cross-modal matching process. Additionally, a text-guided dual-gate fusion module (TGDF-SSA) and a region-aware auxiliary head are introduced to enhance semantic-visual alignment and spatial modeling. Extensive experiments demonstrate that OptiSAR-Net++ achieves SOTA performance on both OptSAR-RSVG and DIOR-RSVG benchmarks, offering significant advantages in localization accuracy and efficiency. The model and dataset have been made publicly available at https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus.
Figures & tables
Fig. 1: Comparison between existing methods and the proposed CD-RSVG paradigm. (a) Existing cross-domain methods rely on co-registered image pairs and are mainly designed for detection, lacking language grounding. (b) Mainstream RSVG methods depend on heavy Transformers and single-source data, resulting in limited cross-domain generalization. (c) Our OptiSAR-Net++ employs a MoE backbone to uniformly handle multi-source inputs, integrates visual-linguistic features, and adopts a CLIP-style matching paradigm for efficient referring visual grounding.
Fig. 2: Overall architecture of OptiSAR-Net++. The framework processes multi-source images and language queries for CD-RSVG via three main components: (1) A shared CNN backbone with PL-MoE, which adaptively routes image patches to domain-specific experts for cross-domain feature modeling. (2) A vision-language fusion neck (TGDF-SSA) that injects semantic information into multi-scale visual features. Text colors denote target categories ( orange / blue ) and attributes ( green ). (3) Detection heads comprising a regression head for candidate generation, a CLIP-based contrastive head for efficient retrieval matching, and an auxiliary region-aware classification head for spatial distribution modeling. During training, adversarial negatives ( red text ) are dynamically sampled to enhance fine-grained cross-domain grounding. Best viewed in color.
Fig. 3: Statistical analysis of the OptSAR-RSVG dataset. (a)-(c) Normalized bounding box height, width, and area distributions for optical (blue) and SAR (green) samples. Optical targets show broader scale variations, whereas SAR targets concentrate on medium-to-small scales. (d) Cross-domain sample quantity, reflecting practical data acquisition ratios. (e) Caption word count distribution (average: 11.13 words). (f) Average pixel area per category, highlighting inter-category scale diversity. (g) Sample distribution across 16 categories. Best viewed in color.
Fig. 4: Word cloud visualizations of textual descriptions in OptSAR-RSVG. (a) Target attributes, highlighting size and color descriptors. (b) Overall vocabulary, dominated by terms like ”image”, ”ship”, and modality identifiers. (c) Category names. (d) Spatial directional vocabulary, providing crucial semantic cues for localization.
Fig. 5: Representative samples from the OptSAR-RSVG dataset, covering diverse optical (top row) and SAR (bottom row) scenarios. Each image is paired with a bounding box and a textual description containing target attributes, categories, and spatial cues. Optical samples exhibit rich semantic details, while SAR samples demonstrate target localization under challenging low-contrast conditions.
Method
Params (M)
Optical Testset
SAR Testset
All Testset
Pr@0.5
Pr@0.7
Pr@0.9
meanIoU
cumIoU
Pr@0.5
Pr@0.7
Pr@0.9
meanIoU
cumIoU
meanIoU
cumIoU
Transformer-Based:
TransVG [ 18 ]
149.7
43.06
30.10
3.23
36.87
39.86
65.99
41.47
1.64
51.72
20.04
40.69
40.90
LQVG [ 44 ]
156.8
89.64
80.06
33.48
77.29
82.75
94.13
87.39
32.14
80.76
81.04
78.04
82.15
TACMT [ 8 ]
150.9
85.51
79.53
42.20
75.74
79.98
92.87
89.74
38.29
81.59
82.46
77.24
81.40
CSDNet [ 14 ]
154.6
86.64
77.86
34.67
75.48
83.20
93.55
90.13
29.67
81.00
74.71
77.01
82.96
TABLE I: Performance comparison on OptSAR-RSVG benchmark. We evaluate our method against SOTA approaches across optical and SAR domains. Our OptSAR-Net++ achieves superior performance with significantly fewer parameters. The best two results are highlighted in red and blue .
Fig. 6: CD-RSVG result comparison in optical and SAR scenes. Green bounding boxes denote the ground-truth annotations, and red bounding boxes indicate model predictions. The text prompts cover diverse semantic cues, including target category, attributes, and spatial location. OptiSAR-Net++ exhibits clear advantages in both localization accuracy and bounding-box regression quality.
Fig. 7: Comparison of data efficiency on OptSAR-RSVG under different training data ratios (top) and comparison of spatial grounding performance under different description masking ratios (bottom).
Method
Params (M)
Optical Testset
SAR Testset
30
60
All
30
60
All
Transformer-Based:
LQVG [ 44 ]
156.8
70.58
73.11
77.29
76.83
77.95
80.76
TACMT [ 8 ]
150.9
59.39
66.38
75.74
72.08
77.15
81.59
CSDNet [ 14 ]
154.6
70.86
75.05
75.48
74.42
77.10
81.00
Contrastive Learning-Based:
TABLE II: Performance Comparison with Different Training Data (%)
Method
Params (M)
Optical Testset
SAR Testset
0
10
30
0
10
30
Transformer-Based:
LQVG [ 44 ]
156.8
77.29
73.30
73.23
80.76
79.95
78.90
TACMT [ 8 ]
150.9
75.74
73.35
69.20
81.59
79.67
77.22
CSDNet [ 14 ]
154.6
75.48
73.64
73.21
81.00
75.78
75.75
Contrastive Learning-Based:
TABLE III: Performance Comparison with Different Masking Caption Ratio (%)
Method
Params (M)
Pr@0.5
Pr@0.7
Pr@0.9
mean IoU
cum IoU
One-Stage Dense Prediction:
ZSGNet [ 46 ]
-
51.67
42.30
10.15
44.12
51.65
FAOA [ 47 ]
-
70.86
62.04
36.44
62.86
67.28
ReSC [ 48 ]
179.9
72.71
63.01
33.37
64.24
68.10
LBYL-Net [ 49 ]
163.8
73.78
65.36
19.52
65.23
76.37
Transformer-Based:
TABLE IV: Performance comparison on DIOR-RSVG benchmark.
Fig. 8: Visualization of last-layer feature activation heatmaps across ablation stages for optical (top two rows) and SAR (bottom two rows) samples. (a) Input image. (b) Baseline model, showing dispersed activations across the scene. (c) Model with fine-grained adversarial sampling, exhibiting improved attention concentration. (d) Full OptiSAR-Net++ (integrating PL-MoE, TGDF-SSA, and the auxiliary head), yielding highly focused activations that precisely align with target boundaries. Green boxes denote localization results.
Fine. Sample
PL- MoE
Sem. Injection
Params (M)
mean IoU
cum IoU
91.0
76.42
82.15
✓
91.0
81.59 +5.17
88.03 +7.88
✓
✓
91.3
82.34 +0.75
88.96 +0.93
✓
✓
✓
95.6
82.76 +0.42
90.70 +1.74
TABLE V: Ablation Study of Key Components
Fig. 9: Visualization of expert routing patterns under different MoE granularity settings. From top to bottom, the models use 2, 4, and 8 experts for cross-domain learning, respectively.
MoE (n, k, p)
2, 2, 2
4, 2, 2
8, 1, 2
8, 3, 2
8, 2, 1 (G. Lv.)
8, 2, ∞ (I. Lv.)
8, 2, 4 (P. Lv.)
8, 2, 2 (P. Lv.)
Params (M)
94.9
95.6
95.6
95.6
95.6
95.6
95.6
95.6
meanIoU
81.04
81.48
81.97
82.64
82.73
81.37
82.71
82.76
cumIoU
87.76
89.82
88.56
89.35
89.84
87.96
90.48
90.70
TABLE VI: Ablation Study of MoE Configuration
Fig. 10: Additional visual grounding predictions of OptiSAR-Net++ on the OptSAR-RSVG test set, covering a variety of representative scenes in optical (first row) and SAR (second row) imagery, as well as several failure cases (third row). For each example, the text query is shown next to the image; green bounding boxes denote the ground-truth annotations, and red bounding boxes indicate the predicted results.
Fig. 11: Zero-shot image annotation comparison between OptiSAR-Net++ and GLIP. Both methods annotate unseen scenes by matching detections against training text embeddings (confidence >0.85 ). Red text indicates misclassifications. OptiSAR-Net++ outperforms GLIP in category accuracy, attribute discrimination, and spatial-relation richness, demonstrating its potential for automated remote-sensing annotation.
Xiaoyu Tang (Member, IEEE) received the B.S. degree from South China Normal University, Shanwei, China, in 2003, and the M.S. degree from Sun Yat-sen University, Guangzhou, China, in 2011. He is currently pursuing the Ph.D. degree with South China Normal University. He is working with Xingzhi College, South China Normal University, where he is engaged in information system development. His research interests include machine vision, intelligent control, and the Internet of Things. Mr. Tang is a member of the IEEE ICICSP Technical Committee.
Table 18
Jun Dong (Student Member, IEEE) is currently pursuing the bachelor’s degree in Internet of Things (IoT) engineering with the School of Data Science and Engineering, South China Normal University, Shanwei, China.
Table 19
Jintao Cheng received his bachelor’s degree from the School of Physics and Telecommunications Engineering, South China Normal University, in 2021. He is currently pursuing an MPhil degree at The Hong Kong University of Science and Technology.
Table 20
Rui Fan (Senior Member, IEEE) received the B.Eng. degree in Automation from the Harbin Institute of Technology in 2015 and the Ph.D. degree in Electrical and Electronic Engineering from the University of Bristol in 2018. He worked as a Research Associate at the Hong Kong University of Science and Technology from 2018 to 2020 and a Postdoctoral Scholar-Employee at the University of California San Diego between 2020 and 2021. He began his faculty career as a Full Research Professor in the College of Electronics & Information Engineering at Tongji University in 2021. He was promoted to Full Professor in 2022 and attained tenure in 2024, both in the same college and at the Shanghai Research Institute for Intelligent Autonomous Systems. His research interests include computer vision, deep learning, and robotics, with a specific focus on humanoid visual perception under the two-streams hypothesis. Prof. Fan served as an associate editor for ICRA’23/25 and IROS’23/24, an area chair for ICIP’24, and a senior program committee member for AAAI’23/24/25/26. He organized several impactful workshops and special sessions in conjunction with WACV’21, ICIP’21/22/23, ICCV’21/25, and ECCV’22. He was honored by being included in the Stanford University List of Top 2% Scientists Worldwide between 2022 and 2025, recognized on the Forbes China List of 100 Outstanding Overseas Returnees in 2023, acknowledged as one of Xiaomi Young Talents in 2023, and awarded the Shanghai Science & Technology 35 Under 35 honor in 2024 as its youngest recipient.
Remote sensing visual grounding (RSVG) aims to locate specific objects in high-resolution RS imagery using free-form natural language descriptions. While recent advances in multimodal large language models (MLLMs) show great potential for such open-vocabulary RSVG, their training-free adaptation is hindered by the modality gap between abstract linguistic semantics and fine-grained visual cues. In cluttered RS scenes, this gap inevitably causes severe localization drift. To bridge this gap, we propose Exemplar-driven Calibrated Refinement (ExACT), a novel training-free framework driven by a one-shot visual prompting mechanism to explicitly provide discriminative structural guidance for precise pixel-level localization. Specifically, we propose a Vision Exemplar-based Calibrator (VEC) that extracts fine-grained visual correspondences from the given exemplar to rectify the rough cross-modal priors from frozen MLLMs, effectively suppressing background artifacts and accurately outlining target boundaries. Subsequently, a Structure-Aware Refiner (SAR) employs an iterative merge-and-select clustering strategy to consolidate the calibrated priors into high-quality positive and negative geometric prompts. These prompts then guide the Segment Anything Model (SAM) to achieve precise pixel-level predictions. Extensive experiments confirm the superiority of ExACT over existing training-free and weakly-supervised methods.
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches and spectral variations, alongside textual logic drift from unconstrained, variable user-input granularities. To mitigate these bottlenecks, this paper establishes the first cross-domain RRSIS benchmark, designated as the Vaihingen-Potsdam Referring (VPRef) dataset, comprising 46,972 language-image-annotation triplets organized into a three-tier linguistic hierarchy. Building upon this benchmark, we develop a tailored parameter-efficient domain adaptation baseline anchored on the Segment Anything Model (SAM3) via Low-Rank Adaptation (LoRA). Our framework counteracts visual distribution discrepancies through pseudo-label-driven self-training and addresses textual logic drift via random multi-granularity text prompt mixing. Crucially, the distribution of empirical metrics across ablative variants suggests a potential decoupling between cross-modal semantic robustification and visual domain alignment, demonstrating that linguistic variance drives fine-grained semantic invariance while pseudo-label propagation governs macro-scale spatial grid alignment. Extensive benchmarks demonstrate the proposed framework achieves superior cross-domain segmentation boundaries while modifying merely 1.08% of the foundational parameter footprint, establishing a robust baseline for future multi-modal remote sensing domain adaptation research. The dataset and code will be available at https://github.com/quanweiliu/VPRef.
Quanwei Liu, Tao Huang, Jiaqi Yang +1
College of Science and Engineering, James Cook University, Cairns, 4878, Australia
Remote sensing visual grounding (RSVG) aims to localize a referred target in a remote sensing image or video according to a natural language expression. Existing RSVG methods usually rely on task-specific manual annotations, which are costly to collect and inevitably limited in covering the diversity of real-world geospatial scenarios. As a result, they often struggle to generalize to open-vocabulary queries involving novel objects, fine-grained attributes, complex spatial relationships, and functional semantics. In this paper, we propose RSVG-ZeroOV, a training-free framework that leverages frozen generic foundation models for zero-shot open-vocabulary RSVG. RSVG-ZeroOV follows an Overview-Focus-Evolve paradigm, which exploits the distinct yet complementary attention patterns of vision-language models (VLMs) and diffusion models (DMs) to progressively generate precise grounding results. Specifically, (i) Overview utilizes a VLM to extract cross-attention maps that capture semantic correlations between the referring expression and visual regions; (ii) Focus leverages the fine-grained modeling priors of a DM to compensate for object structure and shape information often overlooked by VLM attention; and (iii) Evolve introduces a simple yet effective attention evolution module to suppress irrelevant activations, yielding purified object masks. To handle video inputs, we further present Video RSVG-ZeroOV, which extends image-level grounding to spatio-temporal grounding through a query-relevant key-frame selector and a temporal propagator, enabling efficient and temporally coherent video grounding without video annotations or fine-tuning. Extensive experiments on six image and video grounding benchmarks show that RSVG-ZeroOV consistently outperforms existing zero-shot baselines and achieves competitive or superior performance compared with weakly- and fully-supervised methods.
Ke Li, Di Wang, Yongshan Zhu +5
School of Computer Science and Technology, Xidian University, Xi’an 710071, China · Interdisciplinary Institute of Artificial Intelligence, Xidian University, Xi’an, Shaanxi 710126, China · School of Artificial Intelligence, Xidian University, Xi’an 710071, China +2