cs.ROOct 7, 2026

MultiFly: A Real-World Multimodal Aerial Dataset with Annotation-Efficient Label Transfer and Cross-Modal Semantic Consistency

Authors: Markus Gross, Andreas Greiner, Taehyoung Kim, Sivasubiramaniam Subbiah, Tomaž Cotič, Sai Bharadwaj Matha, Conrad Christoph, Oussema Dhaouadi, +7 more

Organizations: Autonomous Aerial Systems, Fraunhofer Institute IVI · Computer Vision Group, Technical University of Munich · Munich Center for Machine Learning (MCML) · Computer Vision for Digital Twins, University of Cambridge · Institute of Innovative Mobility, Univ. of Applied Sciences Ingolstadt

Abstract

We introduce MultiFly, a real-world, low-altitude UAV dataset for semantic perception across RGB, thermal, LiDAR, and radar modalities. MultiFly provides 17,272 synchronized samples from four suburban scenes with frame-wise annotations for 15 semantic classes, together with calibration and GNSS-RTK/IMU measurements. To avoid costly and inconsistent modality-specific annotation, we propagate labels from only 115 manually annotated RGB images through shared geometric representations to all four modalities. This approach generates semantic labels for 17,157 additional RGB images, 17,272 thermal images, 840M LiDAR points, and 3.4M radar points. Transferred annotations achieve 89.93% average agreement with held-out manual annotations, and 90.94% average semantic consistency across all six modality pairs. We further establish semantic segmentation benchmarks for all four modalities, revealing distinct architectural behavior for dense LiDAR and sparse radar data. Taken together, MultiFly provides a scalable foundation for multimodal aerial perception and, to the best of our knowledge, the first public real-world low-altitude aerial benchmark that combines consistent frame-wise semantic annotations for RGB, thermal, LiDAR, and radar. Data at https://github.com/markus-42/multifly.

Figures & tables

Explore similar work

CardsList
  1. SegFly: A Dataset and 2D-3D-2D Paradigm for Aerial RGB-Thermal Semantic Segmentation at Scale

    Mar 18, 2026Markus Gross, Sai Bharadhwaj Matha, Rui Song +4Remote Sensing Image SegmentationCross-Modal Image Registration

  2. GAAT: Geometry-Aware Alignment Transformer for Multimodal UAV Perception

    Aug 28, 2026Jingpu Yang, Debin Tang, Yilin Sun +4Cross-Modal AlignmentMultimodal Pretraining

  3. PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation

    Apr 17, 2026Shuyan Ke, Yifan Mei, Changli Wu +4Image SegmentationRemote Sensing Image Segmentation