SearchWorld: Spatial Value-Grounded Imagination for UAV Object Search via World Models
Organizations: National Key Laboratory of Digital Intelligent Modeling and Simulation, National University of Defense Technology · College of Systems Engineering, National University of Defense Technology · Aerospace Information Research Institute, Chinese Academy of Sciences · BNRist, Tsinghua University · Southeast University
Abstract
Autonomous unmanned aerial vehicle (UAV) object search involves a closed loop of perception, decision-making, and action under partial observability. Urban environments pose several challenges: large search areas and narrow egocentric views limit coverage, dense 3D geometry constrains safe motion, and open-world instructions require identifying a specific target among distractors. Many existing methods mitigate partial observability through explicit maps or memory representations, yet remain largely reactive, reasoning over past observations without explicitly predicting future states. World models enable prospective reasoning through imagined rollouts. However, image-generating world models can incur high inference latency, while spatially grounded planning remains challenging for latent world models. We propose SearchWorld, a recurrent state-space world model that connects explicit spatial memory with value-guided imagination. The model maintains BEV exploration and obstacle memory and decodes a task-aware spatial value layer to guide search. A cognition-action network uses this learned spatial value prior to improve the policy through imagined rollouts, without training a separate scalar critic. Training progresses from world-model learning to expert imitation and imagination-based exploration refinement. On UAV-ON, SearchWorld improves the success rate to 23.8% (19.5% for the strongest published agent) and raises oracle success to 35.5%, while remaining robust on unseen scenes (19.9% success rate). By grounding imagination in explicit spatial representations, SearchWorld enables UAV agents to plan prospectively rather than react.
Figures & tables
| Method | Small | Medium | Large | Total | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SR | OSR | SPL | SR | OSR | SPL | SR | OSR | SPL | SR | OSR | SPL | |
| Random | 4.14% | 7.80% | 2.80% | 3.33% | 8.10% | 3.05% | 2.48% | 8.07% | 1.62% | 3.70% | 8.00% | 2.66% |
| CLIP-H | 2.86% | 8.43% | 1.51% | 10.95% | 16.67% | 7.17% | 13.04% | 19.25% | 10.53% | 6.20% | 11.90% | 4.15% |
| AOA-V | 2.86% | 25.44% | 0.54% | 5.71% | 27.62% | 1.73% | 7.45% | 27.95% | 1.038% | 4.20% | 26.30% | 0.87% |
| AOA-F | 4.45% | 16.38% | 1.61% | 10.48% | 17.62% | 6.36% | 14.29% | 21.74% | 10.66% | 7.30% | 17.50% | 4.06% |
| WMNav | 5.64% | 20.82% | 1.77% | 9.04% | 23.91% | 3.30% | 16.20% | 28.29% | 6.56% | 8.04% | 22.67% | 2.86% |
| Configuration | SR | OSR | SPL |
|---|---|---|---|
| SearchWorld (Ours, full) | 23.80% | 35.52% | 8.03% |
| (Q1) World-model architecture | |||
| direct imitation (BC, no world model) | 14.33% | 25.68% | 7.20% |
| X-Mobility-style (RSSM + imitation only) | 12.91% | 19.88% | 6.95% |
| Dreamer-style (scalar critic in imagination) | 19.07% | 24.99% | 7.64% |
| (Q2) Explicit spatial-cognition memory | |||
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Parameter | Value |
| BEV grid | grid size | |
| cell resolution | ||
| region coverage | ||
| grid centering | search-region center | |
| Exploration layer | sensing range | |
| FOV width |