cs.CVSep 29, 2026

RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA

Authors: Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Sihang Zhao, Chun Yuan, Jing Li

Organizations: Tsinghua University · Zhejiang University · Central University of Finance and Economics · East China Normal University · Key Laboratory of Geographic Information Science

Abstract

Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at inference time. We instead investigate whether the benefit of zoom-in visual privilege can be internalized into the model. We introduce RS-OPSD, a reliable privileged on-policy self-distillation (OPSD) framework for UHR remote sensing VQA. To provide high-quality privileged information with explicit question-relevant evidence, we construct GeoEvidence-6K, containing 6,750 VQA samples across seven task categories with evidence-region annotations, and develop Human Feedback-Guided Skill Refinement (HF-SR) for scalable annotation. To address context loss from tight crops and conflicting signals from imperfect teachers, RS-OPSD introduces Context-Preserving Visual Privilege (CPVP) and Correctness-Aligned Distillation (CAD). Without any additional visual search or tool calls at inference time, RS-OPSD achieves state-of-the-art (SOTA) performance on XLRS-Bench, MME-RealWorld-RS, and LRS-VQA, outperforming previous SOTA models of comparable scale by an average of 4.0 percentage points. Moreover, our 2B variant, RS-OPD-Lite, surpasses most 8B-scale models while achieving the fastest measured inference speed. Our Code, GeoEvidence-6K, and the model weights for RS-OPSD and RS-OPD-Lite are publicly available.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A Unified Framework for Efficient Remote Sensing Visual Question Answering: Adapting Dual, Hybrid, and Encoder-Decoder Architectures

    Jun 17, 2026Timothy Agboada, Shikha Chandel, Yadav Raj Ghimire +1Remote SensingRecent Vision-Language Models

  2. One Adapter, Every Resolution: Gated Low-Rank Adaptation for Remote Sensing VLMs

    May 8, 2026Song Zhang, Yanlong Chen, Yining Chen +2Remote SensingRecent Vision-Language Models

  3. Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

    Jul 28, 2026Fengxiang Wang, Jiangnan Huang, Mingshuo Chen +8Trustworthy Earth Observation DecisionsVisual Reasoning