Micro unmanned aerial vehicles (micro-UAVs) are small enough to reach confined spaces that larger robots cannot access, but too small to carry the sensing and computing power required for autonomous flight. We move the localization stack entirely off the aerial platform onto a quadruped robot with a 7-degree-of-freedom (DOF) arm, which supplies the micro-UAV (27 g bare, 42 g with fiducial markers) its full 6-DOF pose. A camera at the arm's end-effector detects AprilTag fiducial markers on the drone and composes that observation with the quadruped's own self-localization to place the drone in a shared map frame, so the ground robot localizes its partner, rather than only tracking it relative to the camera. The arm acts as an actively-controlled observer, repositioning to keep the drone in view as both robots move; the drone carries only an inertial measurement unit and fuses the external pose to fly commanded setpoints. In lab flights the external pose is accurate to 12-16 mm, enough to fly the drone autonomously within 2-5 cm of motion-capture-fed control. Having the quadruped actively follow the drone reduces the tracking error from 11.0 cm to 6.9 cm by holding the camera in the close range, where the markers are most accurate.
Figures & tables
Figure 1: The integrated hardware. The Boston Dynamics Spot carries the Spot Arm (7 DOF including the gripper)—whose end-effector camera supplies the AprilTag observations of the drone that drive the perception path—and the Jetson AGX Orin on its payload rails. The Crazyflie 2.1+ with its four-AprilTag chassis is the aerial platform the arm camera observes.
Figure 2: System architecture. A perception path (the quadruped’s sensors and the arm-mounted camera through the quadruped’s SLAM, the apriltag_relocalizer , and AprilTag drone-pose estimation) and a command path (trajectory commander through the micro-UAV control stack and radio link to the drone) run concurrently in a single ROS 2 domain; control consumes the composed Tmap→drone and the drone publishes no pose of its own.
Figure 3: Experiment 1: AprilTag vs. MOCAP position and orientation error by camera resolution (raw data and Horn/Markley aligned).
Figure 4: Experiment 2: MOCAP vs AprilTag-controlled flight. (a) Commanded trajectory (dashed) and per-waypoint median flown position with IQR; AprilTag overshoots the y = ±0.5 turn-arounds. (b) Age of the control pose reaching the drone: AprilTag is a median 72 ms stale vs 2 ms for MOCAP, explaining the overshoot.
Figure 5: Experiment 3: Planned trajectory tracking error by quadruped motion using 640 × 480 resolution. (a) tracking error boxplots per quadruped motion condition; (b) tracking error over time per quadruped motion condition
Nano unmanned aerial vehicles (nano-UAVs) can navigate confined spaces that larger robots cannot, but their payload capacity severely limits the sensors and compute available for self-localization in global navigation satellite system (GNSS)-denied environments. We present a vision-based system that localizes a nano-UAV from a quadruped robot with an arm-mounted camera. The quadruped tracks the drone, estimates its position in its own coordinate system using segmentation masks and depth, and transmits that position over a real-time radio link. To ensure continuous tracking, we developed a perception-aware nonlinear Model Predictive Controller (NMPC) that dynamically adjusts the quadruped's body and arm to maximize the drone's visibility, treating observation reliability as a primary control objective. Relying solely on this external estimate, the nano-UAV executed predefined trajectories from takeoff to landing with a median 3D localization error of 59 mm. The system demonstrates robustness against short visual occlusions by smoothly transitioning to onboard inertial flight when line-of-sight is temporarily lost. Ultimately, this framework allows a quadruped to offload the localization burden of a nano-UAV, enabling inspection of complex spaces that neither robot could navigate alone.
Alejandro Lorite Mora, Dimitrios Arapis, Andrés Faíña
IT University of Copenhagen, Denmark · Novo Nordisk A/S, Hillerød, Denmark
Continuous 6-DoF pose estimation is essential for autonomous UAV operations. Yet, existing visual odometry and SLAM methods accumulate drift and yield only relative, up-to-scale trajectories. Single-frame geo-localization, in turn, discards temporal continuity and remains too slow for real-time use. We present OrthoTrack, a training-free system that estimates continuous 6-DoF UAV trajectories using only publicly available orthophotos and surface models as a map prior. OrthoTrack matches keyframes against the orthophoto and lifts correspondences to metric 3D via the surface model. It then propagates these map-anchored correspondences to intermediate frames with optical flow, producing absolute, metrically scaled poses at every frame without GPS or post-hoc alignment. We also introduce the MovingDrone Dataset, a large-scale benchmark pairing photorealistic UAV sequences with dense 6-DoF ground truth and co-registered multi-modal geodata including multi-temporal orthophotos. On MovingDrone and real-world benchmarks, OrthoTrack runs in real time on a single GPU. It outperforms all baselines by a large margin, even those receiving oracle scale and alignment. By relying on publicly available geodata, OrthoTrack enables deployment to new regions without site-specific adaptation.
Oussema Dhaouadi, Zuria Bauer, Johannes Michael Meier +3
ETH Zurich · TU Munich · University of Cambridge +2
Accurate localization of unmanned aerial vehicles (UAVs) is essential for applications such as structural health monitoring, especially in environments where Global Positioning System (GPS) signals are denied or unreliable, like indoor spaces, tunnels, urban canyons, or areas beneath large structures. To address this challenge, we propose Cross-Fusion, a novel method for real-time UAV localization that integrates data from a 3D Light Detection and Ranging (LiDAR) and a monocular camera. A key contribution is its cross-session fusion strategy, which integrates visual and geometric information collected from multiple agents during routine baseline surveys to improve localization consistency and map completeness. The system employs LiDAR-based odometry for motion tracking and image-based feature matching via a single red-green-blue (RGB) camera to correct drift and improve accuracy. Unlike visual-inertial systems, Cross-Fusion maintains a simple sensor setup and avoids the complexity of stereo or global shutter configurations. Experimental results demonstrate that Cross-Fusion achieves localization accuracy comparable to GPS-based methods and performs reliably in challenging feature-sparse environments.
Cong Hoang Quach, Chi Thanh Vo, Dong LT. Tran +3
1 – School of Electrical and Data Engineering, University of Technology Sydney, New South Wales, Australia · 2 – Institute of Aerospace Engineering and Technology, Duy Tan University, Da Nang, Vietnam · 5 – Department of Artificial Intelligence and Robotics, Sejong University, Seoul, South Korea +2