RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA
Authors: Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Sihang Zhao, Chun Yuan, Jing Li
Organizations: Tsinghua University · Zhejiang University · Central University of Finance and Economics · East China Normal University · Key Laboratory of Geographic Information Science
Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at inference time. We instead investigate whether the benefit of zoom-in visual privilege can be internalized into the model. We introduce RS-OPSD, a reliable privileged on-policy self-distillation (OPSD) framework for UHR remote sensing VQA. To provide high-quality privileged information with explicit question-relevant evidence, we construct GeoEvidence-6K, containing 6,750 VQA samples across seven task categories with evidence-region annotations, and develop Human Feedback-Guided Skill Refinement (HF-SR) for scalable annotation. To address context loss from tight crops and conflicting signals from imperfect teachers, RS-OPSD introduces Context-Preserving Visual Privilege (CPVP) and Correctness-Aligned Distillation (CAD). Without any additional visual search or tool calls at inference time, RS-OPSD achieves state-of-the-art (SOTA) performance on XLRS-Bench, MME-RealWorld-RS, and LRS-VQA, outperforming previous SOTA models of comparable scale by an average of 4.0 percentage points. Moreover, our 2B variant, RS-OPD-Lite, surpasses most 8B-scale models while achieving the fastest measured inference speed. Our Code, GeoEvidence-6K, and the model weights for RS-OPSD and RS-OPD-Lite are publicly available.
Figures & tables
Figure 1: Average scores on Ultra-High-Resolution remote sensing VQA benchmarks, including XLRS-Bench, MME-RealWorld-RS, and LRS-VQA.
Figure 2: Comparison of representative paradigms for UHR remote sensing VQA.
Source
Nums.
Image Size
Target Size
GSD
Country
MiniFrance ( Castillo-Navarro et al., 2022 )
110
10K
509×517
0.50m
France
HRSCD ( Daudt et al., 2019 )
40
10K
516×517
0.50m
France
SWISSIMAGE ( Federal Office of Topography swisstopo, 2024 )
300
10K
837×817
0.10m
Switzerland
GeoNRW ( Geobasis NRW, 2025 )
300
10K
997×987
0.10m
Germany
Beeldmateriaal ( Beeldmateriaal Nederland, 2025 )
300
12K
1,741×1,696
0.08m
Netherlands
Table 1: Statistics of the image sources used in GeoEvidence-6K.
Figure 3: Overview of the HF-SR pipeline, together with representative cases of GeoEvidence-6K.
Figure 4: Overview of RS-OPSD , consisting of CPVP and CAD.
Method
Pub.
Backbone
Param.
XLRS
MME
LRS
Avg.
Speed
Closed-Source Vision Language Models
GPT-4o ( Hurst et al., 2024 )
–
–
–
32.4
28.2
27.1
29.2
-
Claude 3.7 Sonnet ( 5 )
–
–
–
40.5
45.7
28.4
38.2
-
Gemini 2.5 Pro ( Comanici et al., 2025 )
–
–
–
44.1
51.7
30.9
42.2
-
Open-Source Vision Language Models
LLaVA-OV-7B ( Li et al., 2024 )
TMLR’25
LLaVA-OV
7B
43.4
53.3
27.3
41.3
2.38
Table 2: Quantitative comparison results. Green , yellow , and blue denote the first-, second-, and third-best results. Speed is the average inference time across the three benchmarks (s/sample).
Core Design
Label
Benchmarks
OPSD
CPVP
CAD
XLRS
MME-RW-RS
LRS-VQA
Avg.
Qwen3-VL-8B-Instruct
Base Model
50.5
41.9
30.1
40.8
✗
✗
✗
LoRA Finetuning on GeoEvidence-6K
49.9
55.2
32.1
45.7
✓
✗
✗
Naïve OPSD in Subsec. 5.1
50.5
53.8
31.3
45.2
✓
✓
✗
OPSD + CPVP
51.3
59.0
31.2
47.2
✓
✗
✓
OPSD + CAD
51.2
58.3
31.5
47.0
Table 3: Ablation study of the core designs in RS-OPSD.
KL Coefficient
Benchmarks
Clipping Threshold
Benchmarks
cgrad
λKL
XLRS
MME-RW-RS
LRS-VQA
Avg.
cgrad
λKL
XLRS
MME-RW-RS
LRS-VQA
Avg.
1
0
51.3
59.0
31.2
47.2
5
1×10−3
53.1
61.5
33.3
49.3
1
1×10−2
51.2
62.0
32.8
48.7
20
1×10−3
51.0
60.0
34.0
48.3
1
1×10−3
53.2
60.8
33.1
49.0
100
1×10−3
51.5
59.9
33.4
48.3
Table 4: Hyperparameter studies of KL regularization and gradient clipping.
Figure 5: Training dynamics of RS-OPSD across three epochs.
Figure 6: Dataset statistics of GeoEvidence-6K. Left: word cloud of the question texts. Middle: token-length distribution of the questions, with the mean, median, and 95th percentile marked by dashed lines. Right: distribution of task categories.
Figure 7: Prompt templates for the student and privileged teacher in RS-OPSD .
Method
Perception
Reasoning
Avg.
Sub-tasks
OC
RC
OLUC
RLUC
OCC
OCL
OMS
OSR
AD
ECR
RP
RCCD
CCR
Closed-Source Vision Language Models
GPT-4o
25.0
32.0
15.0
66.0
9.5
11.3
11.7
24.6
73.0
73.0
35.0
20.0
25.0
32.4
Claude 3.7 Sonnet
27.6
22.7
17.4
68.4
30.5
29.9
63.6
27.6
64.8
78.4
34.5
27.8
32.6
40.5
Gemini 2.5 Pro
50.0
42.0
12.0
65.5
37.5
38.1
60.0
30.2
68.0
72.0
36.0
26.7
35.0
44.1
Open-Source Vision Language Models
Appendix
Table 5: Detailed quantitative results on XLRS-Bench across perception and reasoning sub-tasks. Green , yellow , and blue denote the first-, second-, and third-best results, respectively.
Method
MME-RealWorld-RS
LRS-VQA
XLRS-Bench
Total
Sub-set
Position
Color
Count
Avg.
FAIR
Bridge
STAR
Avg.
Avg.
Accuracy
Speed
Closed-Source Vision Language Models
GPT-4o
36.4
32.4
15.9
28.2
22.2
31.8
27.4
27.1
32.4
29.2
–
Claude 3.7 Sonnet
56.6
52.1
28.5
45.7
25.7
29.5
30.1
28.4
40.5
38.2
–
Gemini 2.5 Pro
60.6
60.3
34.3
51.7
28.9
31.2
32.7
30.9
44.1
42.2
–
Open-Source Vision Language Models
Appendix
Table 6: Detailed quantitative results on MME-RealWorld-RS, LRS-VQA, and XLRS-Bench. Average inference speed is measured in s/sample.
Computational Data Science and Engineering North Carolina A&T State University Greensboro - NC, USA · College of Science and Technology North Carolina A&T State University Greensboro - NC, USA