RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA
Organizations: Tsinghua University · Zhejiang University · Central University of Finance and Economics · East China Normal University · Key Laboratory of Geographic Information Science
Abstract
Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at inference time. We instead investigate whether the benefit of zoom-in visual privilege can be internalized into the model. We introduce RS-OPSD, a reliable privileged on-policy self-distillation (OPSD) framework for UHR remote sensing VQA. To provide high-quality privileged information with explicit question-relevant evidence, we construct GeoEvidence-6K, containing 6,750 VQA samples across seven task categories with evidence-region annotations, and develop Human Feedback-Guided Skill Refinement (HF-SR) for scalable annotation. To address context loss from tight crops and conflicting signals from imperfect teachers, RS-OPSD introduces Context-Preserving Visual Privilege (CPVP) and Correctness-Aligned Distillation (CAD). Without any additional visual search or tool calls at inference time, RS-OPSD achieves state-of-the-art (SOTA) performance on XLRS-Bench, MME-RealWorld-RS, and LRS-VQA, outperforming previous SOTA models of comparable scale by an average of 4.0 percentage points. Moreover, our 2B variant, RS-OPD-Lite, surpasses most 8B-scale models while achieving the fastest measured inference speed. Our Code, GeoEvidence-6K, and the model weights for RS-OPSD and RS-OPD-Lite are publicly available.
Figures & tables
| Source | Nums. | Image Size | Target Size | GSD | Country |
|---|---|---|---|---|---|
| MiniFrance ( Castillo-Navarro et al., 2022 ) | 110 | France | |||
| HRSCD ( Daudt et al., 2019 ) | 40 | France | |||
| SWISSIMAGE ( Federal Office of Topography swisstopo, 2024 ) | 300 | Switzerland | |||
| GeoNRW ( Geobasis NRW, 2025 ) | 300 | Germany | |||
| Beeldmateriaal ( Beeldmateriaal Nederland, 2025 ) | 300 | Netherlands |
| Method | Pub. | Backbone | Param. | XLRS | MME | LRS | Avg. | Speed |
|---|---|---|---|---|---|---|---|---|
| Closed-Source Vision Language Models | ||||||||
| GPT-4o ( Hurst et al., 2024 ) | – | – | – | 32.4 | 28.2 | 27.1 | 29.2 | - |
| Claude 3.7 Sonnet ( 5 ) | – | – | – | 40.5 | 45.7 | 28.4 | 38.2 | - |
| Gemini 2.5 Pro ( Comanici et al., 2025 ) | – | – | – | 44.1 | 51.7 | 30.9 | 42.2 | - |
| Open-Source Vision Language Models | ||||||||
| LLaVA-OV-7B ( Li et al., 2024 ) | TMLR’25 | LLaVA-OV | 7B | 43.4 | 53.3 | 27.3 | 41.3 | 2.38 |
| Core Design | Label | Benchmarks | |||||
|---|---|---|---|---|---|---|---|
| OPSD | CPVP | CAD | XLRS | MME-RW-RS | LRS-VQA | Avg. | |
| Qwen3-VL-8B-Instruct | Base Model | 50.5 | 41.9 | 30.1 | 40.8 | ||
| ✗ | ✗ | ✗ | LoRA Finetuning on GeoEvidence-6K | 49.9 | 55.2 | 32.1 | 45.7 |
| ✓ | ✗ | ✗ | Naïve OPSD in Subsec. 5.1 | 50.5 | 53.8 | 31.3 | 45.2 |
| ✓ | ✓ | ✗ | OPSD + CPVP | 51.3 | 59.0 | 31.2 | 47.2 |
| ✓ | ✗ | ✓ | OPSD + CAD | 51.2 | 58.3 | 31.5 | 47.0 |
| KL Coefficient | Benchmarks | Clipping Threshold | Benchmarks | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| XLRS | MME-RW-RS | LRS-VQA | Avg. | XLRS | MME-RW-RS | LRS-VQA | Avg. | ||||
| 1 | 0 | 51.3 | 59.0 | 31.2 | 47.2 | 5 | 53.1 | 61.5 | 33.3 | 49.3 | |
| 1 | 51.2 | 62.0 | 32.8 | 48.7 | 20 | 51.0 | 60.0 | 34.0 | 48.3 | ||
| 1 | 53.2 | 60.8 | 33.1 | 49.0 | 100 | 51.5 | 59.9 | 33.4 | 48.3 | ||
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Perception | Reasoning | Avg. | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Sub-tasks | OC | RC | OLUC | RLUC | OCC | OCL | OMS | OSR | AD | ECR | RP | RCCD | CCR | |
| Closed-Source Vision Language Models | ||||||||||||||
| GPT-4o | 25.0 | 32.0 | 15.0 | 66.0 | 9.5 | 11.3 | 11.7 | 24.6 | 73.0 | 73.0 | 35.0 | 20.0 | 25.0 | 32.4 |
| Claude 3.7 Sonnet | 27.6 | 22.7 | 17.4 | 68.4 | 30.5 | 29.9 | 63.6 | 27.6 | 64.8 | 78.4 | 34.5 | 27.8 | 32.6 | 40.5 |
| Gemini 2.5 Pro | 50.0 | 42.0 | 12.0 | 65.5 | 37.5 | 38.1 | 60.0 | 30.2 | 68.0 | 72.0 | 36.0 | 26.7 | 35.0 | 44.1 |
| Open-Source Vision Language Models | ||||||||||||||
| Method | MME-RealWorld-RS | LRS-VQA | XLRS-Bench | Total | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Sub-set | Position | Color | Count | Avg. | FAIR | Bridge | STAR | Avg. | Avg. | Accuracy | Speed |
| Closed-Source Vision Language Models | |||||||||||
| GPT-4o | 36.4 | 32.4 | 15.9 | 28.2 | 22.2 | 31.8 | 27.4 | 27.1 | 32.4 | 29.2 | – |
| Claude 3.7 Sonnet | 56.6 | 52.1 | 28.5 | 45.7 | 25.7 | 29.5 | 30.1 | 28.4 | 40.5 | 38.2 | – |
| Gemini 2.5 Pro | 60.6 | 60.3 | 34.3 | 51.7 | 28.9 | 31.2 | 32.7 | 30.9 | 44.1 | 42.2 | – |
| Open-Source Vision Language Models | |||||||||||