Multimodal Large Language Models (MLLMs) have demonstrated remarkable advances in remote sensing. However, existing remote sensing multimodal reasoning benchmarks exhibit two critical limitations: they rely on (i) template-driven questions, which causes repetitive questions; and (ii) static images that fail to capture the inherent temporal nature of drone/UAV videos. This leaves systematic evaluation of remote sensing video reasoning largely unexplored. To address this gap, we introduce ReVA, a new dataset for remote sensing video question answering, designed to assess spatiotemporal, scene-centric, and reasoning-oriented capabilities of MLLMs. ReVA comprises 2,438 drone videos spanning 18 cities worldwide (580K frames) and 22K high-quality question-answer pairs across 11 challenging QA tasks. We develop a semi-automatic annotation pipeline that leverages Text LLMs and MLLMs for question-answer generation with human verification. We evaluate 23 proprietary and open-source Video LLMs on ReVA, exposing fundamental limitations of current models. These findings position ReVA as a critical benchmark toward better remote sensing video understanding and temporal reasoning capabilities for real-world deployments. Our code and dataset are available at: https://github.com/zyaocoder/ReVA
Figures & tables
Figure 1: Overview of ReVA. ReVA contains 22K question-answer pairs under 4 major categories and 11 QA tasks. Word clouds are presented, highlighting different attributes of various categories.
Dataset
Year
#Video
#Image
#QA
Annotate
Unique
FP
TU
SR
CR
question
Natural Domain VideoQA
ActivityNet-QA Yu et al. (2019)
2019
5,800
-
58,000
Manual
18,897
✓
✗
✓
✗
Social-IQ Zadeh et al. (2019)
2019
1,250
-
7,500
Manual
5,713
✓
✗
✗
✓
NExT-QA Xiao et al. (2021)
2021
5,440
-
52,044
Auto
31,173
✓
✓
✗
✓
WildQA Castro et al. (2022)
2022
369
-
916
Manual
251
✓
✓
✓
✗
Table 1: Dataset Comparison. ReVA is a real-world, scene-centric remote sensing VideoQA dataset comprising 2.4K videos with 22K human-annotated QA pairs and 17K unique questions.
Figure 2: Overview of the five-stage QA generation workflow. Stage 1 : Long drone video filtering and clipping; Stage 2 : Comprehensive video caption generation; Stage 3 : Scene-relevant keywords and questions with rationales generation; Stage 4 : Multi-choice answers with answer rationales generation; Stage 5 : Human verification.
Figure 3: Examples of ReVA. Correct answers are marked in green .
Figure 4: QA uniqueness analysis. ReVA achieves the highest (a) average token length (question on bottom and answer on top) and (b) 2-gram diversities, reflecting superior lexical diversity.
Figure 5: Question distribution by categories and tasks.
Split
Video
QA pairs
Train
1,045
15,773
Val
379
2,000
Test
1,014
4,000
Total
2,438
21,773
Table 2: ReVA Statistics.
Figure 6: Overview of ReMoSense Framework and Motion-Aware Alignment Module. The Motion-Aware Alignment Module captures inter-frame correspondence for global motion alignment. It then captures salient object motion iteratively for long-range temporal consistency.
Model
Size
Factual
Temporal
Spatial
Causal
Overall
Perception
Understanding
Reasoning
Reasoning
GU
OR
CD
TG
TP
GR
SL
PV
CA
CO
HR
Based on Proprietary MLLMs
LLoVi
-
86.67
65.15
55.00
49.06
68.33
63.75
76.76
55.21
88.89
90.00
85.62
64.73
VideoTree
-
75.00
46.67
49.60
45.63
63.33
48.50
64.12
50.63
88.89
83.00
86.88
56.20
VideoAgent
-
74.44
43.33
47.20
39.69
59.44
46.25
62.94
52.71
85.56
79.00
85.00
53.62
Table 3: Accuracy (%) on ReVA test set. The best-performing results are presented in bold , while the second-best results are underlined .
Variant
Global
Object
Factual
Temporal
Spatial
Causal
Overall
Perception
Understanding
Reasoning
Reasoning
GU
OR
CD
TG
TP
GR
SL
PV
CA
CO
HR
I
94.44
72.42
64.40
64.06
63.89
69.00
78.24
80.83
94.44
96.00
83.75
73.50
II
✓
95.00
74.85
69.80
65.62
65.28
70.50
78.24
81.25
92.78
93.00
87.50
75.18
III
✓
95.56
76.67
67.20
67.18
65.66
71.50
80.94
81.50
94.44
96.00
91.25
76.12
IV
✓
✓
96.12
78.03
70.85
76.21
76.40
74.25
83.27
83.74
95.32
97.65
90.74
80.04
Table 4: Ablation study for various design choices. The best-performing results are presented in bold , while the second-best results are underlined .
Figure 7: Per-task comparisons with SOTA methods on ReVA test set.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Size
Factual
Temporal
Spatial
Causal
Overall
Perception
Understanding
Reasoning
Reasoning
GU
OR
CD
TG
TP
GR
SL
PV
CA
CO
HR
Based on Proprietary MLLMs
VideoAgent
-
76.67
41.21
43.80
32.92
53.89
38.19
52.66
64.85
86.67
91.84
86.08
51.53
LLoVi
-
91.11
63.03
48.80
47.50
66.11
64.50
71.18
72.50
93.33
92.00
86.25
65.30
VideoTree
-
84.44
49.70
51.20
49.38
71.11
64.00
66.47
64.58
91.11
90.00
86.25
62.30
Appendix
Table 5: Accuracy (%) on ReVA validation set. For each block, the best-performing results are presented in bold , while the second-best results are underlined .
Figure 8: Per-task comparisons with SOTA methods on ReVA validation set.
Category / Subcategory
Text-only
First frame
Middle frame
Full
Factual Perception
29.55
38.96
39.33
75.08
General Understanding
35.00
22.78
24.44
97.78
Object and Land Cover Recognition
30.00
39.85
40.30
74.85
Change Detection
27.00
43.60
43.40
67.20
Temporal Understanding
27.00
37.70
37.80
64.40
Temporal Grounding
25.94
32.19
32.50
58.44
Appendix
Table 6: Ablation studies of different inputs of ReMoSense on ReVA.
Figure 9: Visualization of answer predictions on ReVA. The correct answers and correct predictions are highlighted in green , and incorrect predictions are in red .
Category
Initial candidates
Final total
Accept
Minor correction
Reject
Review
Factual
6314
5773
4834 (76.6%)
610 (9.7%)
541 (8.6%)
329 (5.2%)
Temporal
5887
5072
3613 (61.4%)
1058 (18.0%)
815 (13.8%)
401 (6.8%)
Spatial
9849
9010
7389 (75.0%)
1301 (13.2%)
839 (8.5%)
320 (3.2%)
Causal
2113
1918
1797 (85.0%)
72 (3.4%)
195 (9.2%)
49 (2.3%)
Overall
24163
21773
17633 (73.0%)
3041 (12.6%)
2390 (9.9%)
1099 (4.5%)
Appendix
Table 7: Statistics of per-category annotations.
Categories
Count
Required revision
Ratio
Change detection
1879
774
41.2%
Temporal grounding
3142
1247
39.7%
Trend and pattern analysis
2745
1027
37.4%
Viewpoint and perspective analysis
2820
857
30.4%
Geometric relation reasoning
4194
1254
29.9%
Object and land-cover recognition
3014
588
19.5%
Appendix
Table 8: Task-wise number and ratio of initial QA pairs that require revision.
Option
Initial generation
Final release
A
15345 (70.5%)
5,443 (25.0%)
B
3730 (17.1%)
5,443 (25.0%)
C
1806 (8.3%)
5,443 (25.0%)
D
802 (3.7%)
5,444 (25.0%)
E
90 (0.4%)
0 (0.0%)
Appendix
Table 9: Correct answer distribution before and after answer-option shuffling.
Category
Pairwise agreement
Majority-vote agreement
Factual perception
1512 (84.4%)
90.6%
Temporal understanding
1378 (86.3%)
92.1%
Spatial & viewpoint reasoning
852 (82.2%)
89.2%
Causal reasoning
929 (79.0%)
87.2%
Overall
4671 (83.4%)
90.1%
Appendix
Table 10: Agreement Results Across Different Categories
Figure 10: Web-based QA review platform for human verification of generated QA pairs.
Category
QA task
Question example
Factual Perception
Activity Recognition;
Object Existence Identification;
General Understanding
Overall Scene Description;
Object Counting;
Object Category Classification;
Object and Land Cover Recognition
Object Spatial Distribution
Appendix
Table 11: ReVA Question Type Details. Each category has 2-3 QA tasks with representative question examples.
Figure 15: Examples of each category. From top to bottom: (a) Factual Perception; (b) Temporal Understanding; (c) Spatial Reasoning; (d) Causal Reasoning.
Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene changes that cannot be captured by isolated images. Existing models primarily target single images or discrete temporal observations spanning a long time range. However, a unified evaluation setting for assessing vision-language models on continuous remote-sensing video understanding remains lacking. We introduce RSVideo-10K, a remote-sensing video dataset comprising 10,773 instances, 1.47 million frames, and 17.02 hours of footage, containing both unmanned aerial vehicles and satellite platforms. Its fixed evaluation benchmark, RSVideo-Bench, contains 2,731 test instances and evaluates two complementary aspects of remote-sensing video understanding: L1 Perception and L2 Reasoning, spanning seven capability groups and 17 tasks. Evaluations show that current vision-language models still struggle to recover small local evidence, track short-lived states, and use scene-constrained spatial relations. Based on this analysis, we further propose RSVideo, a reinforcement learning framework for small-target spatiotemporal focusing that selects question-relevant regions across frames and suppresses redundant background tokens. RSVideo achieves a maximum absolute improvement of 9.01% with InternVL3.5-14B and attains the highest accuracy of 40.63% with Qwen3.6-27B across 26 open-source vision-language backbones. Codes will be available at https://github.com/HongjieZhou0329/RSVideo.
Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved remote sensing (RS) multimodal understanding. Language-conditioned segmentation is crucial for fine-grained target understanding in Unmanned Aerial Vehicle (UAV) videos. However, this task remains challenging due to the prevalence of small, visually ambiguous targets and dynamic aerial perspectives. In this paper, we propose SkyVLaM, a multimodal large language model for UAV video understanding. SkyVLaM constructs sparse tokens directly from patch-level video representations through a temporal basis perceiver, regularizes the sparse basis to encourage complementary temporal cues, and adaptively selects a temporally coherent dense segment for high-resolution inspection. The resulting sparse and dense tokens are jointly processed by a large language model for query-conditioned segmentation. We further build SkyVid, consisting of SkyVid-VGCG and SkyVid-RVOS for video grounded conversation generation and referring video object segmentation, respectively. SkyVid contains 101 videos, 33.6K frames, and 1.53M pixel-level object instances. Experiments show that SkyVLaM provides a more effective allocation of the visual token budget and improves language-conditioned video segmentation in UAV scenarios.
Kaiwen Jing, Ruixu Jia, Bingyao Li +3
Beijing University of Posts and Telecommunications, Beijing, China · Peking University, Beijing, China · Beijing Wuzi University, Beijing, China
Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusion mechanisms. In this work, we challenge the necessity of such architectural specialization. We show that a generally capable vision-language model can achieve competitive or state-of-the-art performance at challenging remote sensing benchmarks, provided that it is trained at sufficient scale across diverse data and tasks. Our model uses a single language policy that can either answer directly in text or invoke a localization tool for segmentation and grounding. To train this heterogeneous behaviour, we employ a multi-task reinforcement learning framework with adaptive task rewards covering multiple-choice VQA, free-form VQA, captioning, detection, and segmentation across a large variety of input types. Our approach achieves competitive results across a broad set of benchmarks, including high-resolution, multi-temporal, multi-modal and multi-view tasks. Further, as training data scales, our experiments show consistent improvements across most tasks both in and out of distribution, which correlate with per-task data diversity. These findings suggest that, for remote sensing VLMs, data scale is sufficient even without architectural novelty.
Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi +2