cs.CVSep 28, 2026

ReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question Answering

Authors: Zhen Yao, Likai Wang, Yuming Yang, Zhihao Zheng, Bo Lang, Qiuyu Tang, Jialu Sheng, Jingqi Xu, +4 more

Organizations: Lehigh University · University of Southern California · Qualcomm AI Research

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated remarkable advances in remote sensing. However, existing remote sensing multimodal reasoning benchmarks exhibit two critical limitations: they rely on (i) template-driven questions, which causes repetitive questions; and (ii) static images that fail to capture the inherent temporal nature of drone/UAV videos. This leaves systematic evaluation of remote sensing video reasoning largely unexplored. To address this gap, we introduce ReVA, a new dataset for remote sensing video question answering, designed to assess spatiotemporal, scene-centric, and reasoning-oriented capabilities of MLLMs. ReVA comprises 2,438 drone videos spanning 18 cities worldwide (580K frames) and 22K high-quality question-answer pairs across 11 challenging QA tasks. We develop a semi-automatic annotation pipeline that leverages Text LLMs and MLLMs for question-answer generation with human verification. We evaluate 23 proprietary and open-source Video LLMs on ReVA, exposing fundamental limitations of current models. These findings position ReVA as a critical benchmark toward better remote sensing video understanding and temporal reasoning capabilities for real-world deployments. Our code and dataset are available at: https://github.com/zyaocoder/ReVA

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?

    Aug 3, 2026Hongjie Zhou, Shiqin Wang, Haoyang Chen +5Remote SensingSpatiotemporal

  2. SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing

    Jul 19, 2026Kaiwen Jing, Ruixu Jia, Bingyao Li +3Video Object SegmentationDrone Detection

  3. More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

    Jul 17, 2026Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi +2Remote SensingRecent Vision-Language Models