ReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question Answering
Organizations: Lehigh University · University of Southern California · Qualcomm AI Research
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable advances in remote sensing. However, existing remote sensing multimodal reasoning benchmarks exhibit two critical limitations: they rely on (i) template-driven questions, which causes repetitive questions; and (ii) static images that fail to capture the inherent temporal nature of drone/UAV videos. This leaves systematic evaluation of remote sensing video reasoning largely unexplored. To address this gap, we introduce ReVA, a new dataset for remote sensing video question answering, designed to assess spatiotemporal, scene-centric, and reasoning-oriented capabilities of MLLMs. ReVA comprises 2,438 drone videos spanning 18 cities worldwide (580K frames) and 22K high-quality question-answer pairs across 11 challenging QA tasks. We develop a semi-automatic annotation pipeline that leverages Text LLMs and MLLMs for question-answer generation with human verification. We evaluate 23 proprietary and open-source Video LLMs on ReVA, exposing fundamental limitations of current models. These findings position ReVA as a critical benchmark toward better remote sensing video understanding and temporal reasoning capabilities for real-world deployments. Our code and dataset are available at: https://github.com/zyaocoder/ReVA
Figures & tables
| Dataset | Year | #Video | #Image | #QA | Annotate | Unique | FP | TU | SR | CR |
| question | ||||||||||
| Natural Domain VideoQA | ||||||||||
| ActivityNet-QA Yu et al. (2019) | 2019 | 5,800 | - | 58,000 | Manual | 18,897 | ✓ | ✗ | ✓ | ✗ |
| Social-IQ Zadeh et al. (2019) | 2019 | 1,250 | - | 7,500 | Manual | 5,713 | ✓ | ✗ | ✗ | ✓ |
| NExT-QA Xiao et al. (2021) | 2021 | 5,440 | - | 52,044 | Auto | 31,173 | ✓ | ✓ | ✗ | ✓ |
| WildQA Castro et al. (2022) | 2022 | 369 | - | 916 | Manual | 251 | ✓ | ✓ | ✓ | ✗ |
| Split | Video | QA pairs |
| Train | 1,045 | 15,773 |
| Val | 379 | 2,000 |
| Test | 1,014 | 4,000 |
| Total | 2,438 | 21,773 |
| Model | Size | Factual | Temporal | Spatial | Causal | Overall | |||||||
| Perception | Understanding | Reasoning | Reasoning | ||||||||||
| GU | OR | CD | TG | TP | GR | SL | PV | CA | CO | HR | |||
| Based on Proprietary MLLMs | |||||||||||||
| LLoVi | - | 86.67 | 65.15 | 55.00 | 49.06 | 68.33 | 63.75 | 76.76 | 55.21 | 88.89 | 90.00 | 85.62 | 64.73 |
| VideoTree | - | 75.00 | 46.67 | 49.60 | 45.63 | 63.33 | 48.50 | 64.12 | 50.63 | 88.89 | 83.00 | 86.88 | 56.20 |
| VideoAgent | - | 74.44 | 43.33 | 47.20 | 39.69 | 59.44 | 46.25 | 62.94 | 52.71 | 85.56 | 79.00 | 85.00 | 53.62 |
| Variant | Global | Object | Factual | Temporal | Spatial | Causal | Overall | |||||||
| Perception | Understanding | Reasoning | Reasoning | |||||||||||
| GU | OR | CD | TG | TP | GR | SL | PV | CA | CO | HR | ||||
| I | 94.44 | 72.42 | 64.40 | 64.06 | 63.89 | 69.00 | 78.24 | 80.83 | 94.44 | 96.00 | 83.75 | 73.50 | ||
| II | ✓ | 95.00 | 74.85 | 69.80 | 65.62 | 65.28 | 70.50 | 78.24 | 81.25 | 92.78 | 93.00 | 87.50 | 75.18 | |
| III | ✓ | 95.56 | 76.67 | 67.20 | 67.18 | 65.66 | 71.50 | 80.94 | 81.50 | 94.44 | 96.00 | 91.25 | 76.12 | |
| IV | ✓ | ✓ | 96.12 | 78.03 | 70.85 | 76.21 | 76.40 | 74.25 | 83.27 | 83.74 | 95.32 | 97.65 | 90.74 | 80.04 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Size | Factual | Temporal | Spatial | Causal | Overall | |||||||
| Perception | Understanding | Reasoning | Reasoning | ||||||||||
| GU | OR | CD | TG | TP | GR | SL | PV | CA | CO | HR | |||
| Based on Proprietary MLLMs | |||||||||||||
| VideoAgent | - | 76.67 | 41.21 | 43.80 | 32.92 | 53.89 | 38.19 | 52.66 | 64.85 | 86.67 | 91.84 | 86.08 | 51.53 |
| LLoVi | - | 91.11 | 63.03 | 48.80 | 47.50 | 66.11 | 64.50 | 71.18 | 72.50 | 93.33 | 92.00 | 86.25 | 65.30 |
| VideoTree | - | 84.44 | 49.70 | 51.20 | 49.38 | 71.11 | 64.00 | 66.47 | 64.58 | 91.11 | 90.00 | 86.25 | 62.30 |
| Category / Subcategory | Text-only | First frame | Middle frame | Full |
| Factual Perception | 29.55 | 38.96 | 39.33 | 75.08 |
| General Understanding | 35.00 | 22.78 | 24.44 | 97.78 |
| Object and Land Cover Recognition | 30.00 | 39.85 | 40.30 | 74.85 |
| Change Detection | 27.00 | 43.60 | 43.40 | 67.20 |
| Temporal Understanding | 27.00 | 37.70 | 37.80 | 64.40 |
| Temporal Grounding | 25.94 | 32.19 | 32.50 | 58.44 |
| Category | Initial candidates | Final total | Accept | Minor correction | Reject | Review |
| Factual | 6314 | 5773 | 4834 (76.6%) | 610 (9.7%) | 541 (8.6%) | 329 (5.2%) |
| Temporal | 5887 | 5072 | 3613 (61.4%) | 1058 (18.0%) | 815 (13.8%) | 401 (6.8%) |
| Spatial | 9849 | 9010 | 7389 (75.0%) | 1301 (13.2%) | 839 (8.5%) | 320 (3.2%) |
| Causal | 2113 | 1918 | 1797 (85.0%) | 72 (3.4%) | 195 (9.2%) | 49 (2.3%) |
| Overall | 24163 | 21773 | 17633 (73.0%) | 3041 (12.6%) | 2390 (9.9%) | 1099 (4.5%) |
| Categories | Count | Required revision | Ratio |
| Change detection | 1879 | 774 | 41.2% |
| Temporal grounding | 3142 | 1247 | 39.7% |
| Trend and pattern analysis | 2745 | 1027 | 37.4% |
| Viewpoint and perspective analysis | 2820 | 857 | 30.4% |
| Geometric relation reasoning | 4194 | 1254 | 29.9% |
| Object and land-cover recognition | 3014 | 588 | 19.5% |
| Option | Initial generation | Final release |
| A | 15345 (70.5%) | 5,443 (25.0%) |
| B | 3730 (17.1%) | 5,443 (25.0%) |
| C | 1806 (8.3%) | 5,443 (25.0%) |
| D | 802 (3.7%) | 5,444 (25.0%) |
| E | 90 (0.4%) | 0 (0.0%) |
| Category | Pairwise agreement | Majority-vote agreement |
| Factual perception | 1512 (84.4%) | 90.6% |
| Temporal understanding | 1378 (86.3%) | 92.1% |
| Spatial & viewpoint reasoning | 852 (82.2%) | 89.2% |
| Causal reasoning | 929 (79.0%) | 87.2% |
| Overall | 4671 (83.4%) | 90.1% |
| Category | QA task | Question example |
| Factual Perception | Activity Recognition; | |
| Object Existence Identification; | ||
| General Understanding | Overall Scene Description; | |
| Object Counting; | ||
| Object Category Classification; | ||
| Object and Land Cover Recognition | Object Spatial Distribution |