Organizations: Department of Computer Science and Technology, BNRist, Tsinghua University · School of Electrical & Electronic Engineering, University College Dublin · School of Electronics Engineering and Computer Science, Peking University · Department of Computer Science and Technology, Zhejiang University of Technology
Open-world aerial object-goal search is a foundational yet challenging task, requiring aerial agents to autonomously explore large-scale, unstructured three-dimensional environments and reach target objects specified by semantic descriptions or reference images, rather than following route-specific instructions. However, research in this task remains at a nascent stage and relies on small, environment-specific benchmarks with heterogeneous action spaces and data formats. These limitations hinder large-scale training and cross-benchmark evaluation, constraining the scalability and generalizability of aerial agents. To address this problem, we propose AerialDojo-200K, a large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task. Specifically, we construct 42 simulation scenes spanning four scene families and 21 scene types, including 18 urban, 12 natural, six infrastructure, and six disaster scenes. To ensure data quality, 12 annotators spent two months manually annotating 109 landmarks, 2099 target objects, and 2099 object anchors across these scenes. We further construct 205,732 task instances, comprising over 100K semantic-goal and over 100K image-goal instances across Base, Standard, and Long-Horizon settings. Each task instance includes a collision-free reference trajectory and corresponding multi-view video recordings. We also develop a unified evaluation framework with a scene partition comprising 21 in-distribution scenes and 21 out-of-distribution scenes. Finally, our evaluation of five open-source and four closed-source multimodal large language models reveals that there is still a long way to go toward achieving general-purpose aerial agents. All can be found at https://fengtt42.github.io/AerialDojo/.
Figures & tables
Figure 1: The overview of AerialDojo-200K.
Benchmark
Year
Task
Action
Goal Type
Goal Specification
Ntask
Fscene
Nscene
AerialVLN
2023
VLN
4-DoF
Location
Movement instruction
25.3K
U
25
CityNav
2025
VLN
4-DoF
Location
Movement instruction
32.6K
U
34
OpenUAV
2025
VLN
6-DoF
Location
Movement instruction
12.1K
U,N
22
OpenFly
2026
VLN
4-DoF
Location
Movement instruction
103K
U
18
UAV-ON
2025
ObjNav
4-DoF
Instance
Semantic
11K
U,N
14
CityAVOS
2026
OGS
4-DoF
Instance
Semantic and Image
2.4K
U
6
Table 1: Comparison of aerial benchmarks. Ntask : the number of task instances; Fscene : Scene families; Nscene : the number of scenes. VLN: vision-and-language navigation; ObjNav: object-goal navigation; OGS: object-goal search. DoF: degrees of fredom. U: urban; N: natural; I: infrastructure; D: disaster.
Figure 2: Standard operating procedure for constructing AerialDojo-200K.
Family
Types
Scenes
LM
Objects
Anchors
Base
Standard
LH
ALL
Path (km)
Urban
9
18
57
1,030
1,030
42,232
55,840
26,992
125,064
2,607.253
Natural
6
12
21
569
569
21,668
15,188
12,976
49,832
1,008.083
Infrastructure
3
6
20
280
280
8,934
7,800
3,080
19,814
357.067
Disaster
3
6
11
220
220
7,140
3,596
286
11,022
142.910
Total
21
42
109
2,099
2,099
79,974
82,424
43,334
205,732
4,115.313
Table 2: Dataset statistics. LM: landmarks; Base, Standard, and LH denote task-instance counts for the Base, Standard, and Long-Horizon setting, respectively, including both semantic-goal and image-goal instances; ALL is their sum. Path reports the cumulative planned path length in kilometers, counted once per underlying navigation task.
Figure 3: Simulator and dataset statistics of AerialDojo-200K.
Setting
L(τref)
Action budget
Base
[5,30)
90
Standard
[30,60)
180
Long-Horizon
[60,100)
300
Table 3: Task settings.
Figure 4: The unified benchmark framework of AerialDojo-200K.
Model
3 m (primary)
5 m
SR ↑
OSR ↑
SPL ↑
DTS ↓
CR ↓
SR ↑
OSR ↑
Open-source models
Qwen3-VL-4B
0.71 ± 0.81
1.50 ± 1.33
0.51 ± 0.60
16.24 ± 1.85
22.17 ± 14.32
3.85 ± 2.92
7.99 ± 3.93
InternVL3.5-8B
0.16 ± 0.30
3.41 ± 2.36
0.11 ± 0.22
25.42 ± 11.63
62.06 ± 16.89
1.54 ± 1.74
12.04 ± 4.47
Ministral 3 8B
0.17 ± 0.31
1.28 ± 1.48
0.09 ± 0.21
16.23 ± 3.43
51.63 ± 24.41
1.64 ± 1.18
11.37 ± 5.20
Phi-4-multimodal
0.13 ± 0.35
0.60 ± 1.01
0.09 ± 0.25
23.52 ± 7.47
37.96 ± 16.67
1.65 ± 1.79
9.22 ± 6.17
Table 4: Base-task performance across four scene families. SR, OSR, SPL, and CR are percentages; DTS is in meters. Arrows indicate the preferred direction; bold denotes the best available value.
Model
3 m (primary)
5 m
SR ↑
OSR ↑
SPL ↑
DTS ↓
CR ↓
SR ↑
OSR ↑
GPT-5.6 Sol
0.26 ± 0.72
0.26 ± 0.72
0.20 ± 0.57
37.66 ± 5.47
36.57 ± 16.79
2.30 ± 2.02
2.30 ± 2.02
Claude Opus 5
0.49 ± 0.91
1.32 ± 2.33
0.40 ± 0.73
36.49 ± 8.89
56.30 ± 21.31
1.29 ± 1.91
2.61 ± 3.84
Table 5: Standard-task performance across four scene families. SR, OSR, SPL, and CR are percentages; DTS is in meters. Arrows is the preferred direction; bold denotes the best available value.
Figure 5: Base-task performance across four scene families at the 3 m success threshold. Each curve represents one model. From the center to the outer edge, SR and SPL range from 0 to 5%, OSR from 0 to 15%, DTS from 40 to 0 m, and CR from 100 to 0%; outward therefore indicates better performance. Ring labels show normalized radial values.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Prompt for MLLM evaluation in AerialDojo-200K.
Figure 7: Task composition and reference lengths. (a) Goal-conditioned instance counts by family and setting. (b) Train/Test counts in the ID and OOD scene groups. Counts in (a,b) include both goals. (c) Mean reference length, weighted by navigation-task count within each family.
Scenes
Task subset
Base
Standard
LH
ALL
ID
Train
40,640
48,398
24,604
113,642
ID
Test
10,240
12,060
6,152
28,452
OOD
Train
5,764
4,452
2,496
12,712
OOD
Test
23,330
17,514
10,082
50,926
Total
All
79,974
82,424
43,334
205,732
Appendix
Table 6: Task-instance counts by scene group, local Train/Test subset, and difficulty. Both goal modalities are included. ID/OOD identifies the scene group; Train/Test identifies the task subset within that group.
Scene
LM
Obj./Anc.
Base
Standard
LH
ALL
Path (km)
Urban
Mall_0
6
123/123
12,412
19,652
8,150
40,214
839.245
Mall_1
5
102/102
5,390
7,144
6,536
19,070
459.569
Stadium_0
5
64/64
4,344
8,032
146
12,522
213.410
Stadium_1
1
41/41
1,234
2,010
1,156
4,400
97.323
Park_0
2
14/14
346
90
6
442
5.077
Appendix
Table 7: Scene inventory and search task statistics in AerialDojo-200K. For each scene, LM reports the number of landmarks, and Obj./Anc. reports the numbers of annotated objects and anchors, respectively. Base, Standard, and LH report the numbers of search task instances in the Base, Standard, and Long-Horizon settings; ALL is their sum. Task instance counts include both goal modalities. Path (km) reports the total length of reference trajectories, with each trajectory counted once. Scene suffixes _0 and _1 indicate in-distribution (ID) and out-of-distribution (OOD) scenes, respectively.
Despite the rapid progress in data-driven 3D vision, aerial geometric 3D vision remains a formidable challenge due to the severe scarcity of large-scale, high-fidelity training data. Existing benchmarks, predominantly biased toward ground-level or object-centric views, do not account for complex viewpoint transformations and diverse environmental conditions in UAV-based sensing. To bridge this critical gap, we propose AirZoo, a unified large-scale dataset and benchmark for grounding aerial geometric 3D vision. AirZoo possesses three appealing properties: 1) Scalable Generation Pipeline: Leveraging freely available, world-scale photogrammetric 3D meshes, it renders vast outdoor environments with customizable UAV flight trajectories and configurable weather/illumination. 2) Comprehensive Scene Diversity: It provides the most extensive coverage of region types to date (spanning 378 regions across 22 countries), systematically encompassing both highly structured urban landscapes and complex unstructured natural environments. 3) Rich Geometric Annotations: Each frame provides synchronized, pixel-level metric depth and precise 6-DoF geo-referenced poses, essential for geometry-aware learning. Through three rigorous evaluation tracks -- aerial image retrieval, cross-view matching, and multi-view 3D reconstruction -- we demonstrate that AirZoo serves as a powerful pre-training engine. Extensive experiments on both public and newly collected real-world benchmarks reveal that fine-tuning on AirZoo yields substantial performance gains for SoTA models (e.g., MegaLoc, RoMa, VGGT, and Depth Anything 3), establishing a new performance upper bound for aerial spatial intelligence.
Xiaoya Cheng, Rouwan Wu, Xinyi Liu +6
National University of Defense Technology, Changsha, China · National Key Laboratory of Advanced Guidance and Control Technology, Changsha, China · Singapore University of Technology and Design, Singapore
The rapid advancement of Multimodal Large Language Models (MLLMs) has empowered Unmanned Aerial Vehicle (UAV) with exceptional capabilities in spatial reasoning, semantic understanding, and complex decision-making, making them inherently suited for UAV Search and Rescue (SAR). However, existing UAV SAR research is dominated by traditional vision and path-planning methods and lacks a comprehensive and unified benchmark for embodied agents. To bridge this gap, we first propose the novel task of \textbf{Embodied Search and Rescue (ESAR)}, which requires aerial agents to autonomously explore complex environments, identify rescue clues, and reason about victim locations to execute informed decision-making. Additionally, we present \textbf{ESARBench}, the first comprehensive benchmark designed to evaluate MLLM-driven UAV agents in highly realistic SAR scenarios. Leveraging Unreal Engine 5 and AirSim, we construct four high-fidelity, large-scale open environments mapped directly from real-world Geographic Information System (GIS) data to ensure photorealistic landscapes. To rigorously simulate actual rescue operations, our benchmark incorporates dynamic variables including weather conditions, time of day, and stochastic clue placement. Furthermore, we create a dataset of 600 tasks modeled after real-world rescue cases and propose a robust set of evaluation metrics. We evaluate diverse baselines, ranging from traditional heuristics to advanced ground and aerial MLLM-based ObjectNav agents. Experimental results highlight the challenges in ESAR, revealing critical bottlenecks in spatial memory, aerial adaptation, and the trade-off between search efficiency and flight safety. We hope ESARBench serves as a valuable resource to advance research on Embodied Search and Rescue domain. Source code and project page: https://4amgodvzx.github.io/ESAR.github.io.
Air-Ground Object Search (AGOS) in urban environments is a challenging embodied task, which requires an Unmanned Aerial Vehicle (UAV) and an Unmanned Ground Vehicle (UGV) to jointly search for and verify a specified target vehicle from multi-view visual references. To study this underexplored problem, we introduce AGOS-Bench, the first dedicated benchmark for evaluating whether general-purpose Vision-Language Models (VLMs) can integrate aerial discoveries and ground-level verification through UAV-UGV cooperation. We further provide AGOS-Dataset as the companion resource of exemplary trajectories constructed by an automatic pipeline. It consists of 7.7k episodes for searching objects of diverse categories and attributes, spanning three difficulty levels. To address the AGOS task, we propose AGOS-Agent, a training-free and tool-augmented approach. The agentic method relieves VLMs from complex and dynamic coordination via a deliberate search-handoff-verify cooperation protocol, only demanding VLMs for scene understanding and decision-making. Extensive experiments on nine VLMs show that AGOS-Agent improves overall success rate for eight of the nine evaluated backbones while reducing decision steps for all nine. On the hard split, the SR and SPL of Gemini-3.6-Flash increase from 8.6% to 55.7% and from 7.6% to 44.0%, respectively.
Boao Yu, Zimo Chen, Junreng Rao +4
National University of Defense Technology National Key Laboratory of Digital Intelligent Modeling and Simulation