Nano unmanned aerial vehicles (nano-UAVs) can navigate confined spaces that larger robots cannot, but their payload capacity severely limits the sensors and compute available for self-localization in global navigation satellite system (GNSS)-denied environments. We present a vision-based system that localizes a nano-UAV from a quadruped robot with an arm-mounted camera. The quadruped tracks the drone, estimates its position in its own coordinate system using segmentation masks and depth, and transmits that position over a real-time radio link. To ensure continuous tracking, we developed a perception-aware nonlinear Model Predictive Controller (NMPC) that dynamically adjusts the quadruped's body and arm to maximize the drone's visibility, treating observation reliability as a primary control objective. Relying solely on this external estimate, the nano-UAV executed predefined trajectories from takeoff to landing with a median 3D localization error of 59 mm. The system demonstrates robustness against short visual occlusions by smoothly transitioning to onboard inertial flight when line-of-sight is temporarily lost. Ultimately, this framework allows a quadruped to offload the localization burden of a nano-UAV, enabling inspection of complex spaces that neither robot could navigate alone.
Figures & tables
Fig. 1: The quadruped localizes the nano-UAV with no marker on the drone and no localization sensing aboard it. It computes the drone’s position with its arm camera, in its own map, and sends it over radio. Overlays are projected into the footage from the trial’s own recording: the yellow cross is the markerless estimate the drone flew on, the blue circle is the motion-capture ground truth used for evaluation only. Insets: the quadruped’s camera view and a 3× magnification of the drone.
Fig. 2: The deployed system. Solid path: one position per frame, expressed in the quadruped’s own map frame. Shaded stages are the three integrity gates of the deployed gripper-camera path. A failed gate publishes nothing and the drone’s filter coasts on its IMU, a dropout instead of a wrong position. Section IV-A reports their aggregate effect. Dashed: the accepted position also drives the perception-aware controller (Section III-E ), which commands the arm and body so the drone stays observable.
Fig. 3: Live accuracy of the deployed estimator (gripper camera) against motion capture over 17363 positions from eight closed-loop flights. Each point is the median 3D error of the positions recorded within a 0.2 m interval of distance, the band around the curve their interquartile range, and intervals holding fewer than 80 positions are not drawn. The gray span marks the distance range all three conditions cover, and the inset box gives each condition’s median and interquartile range over that shared range.
Fig. 4: Localization accuracy on the earlier stereo camera, comparing the fiducial baseline against three markerless configurations. Each point is the median 3D error of the positions recorded within a distance interval, and the band around a curve their interquartile range. The bare raw-depth series is medians only. Turning the gates on removes the bare airframe from the plot entirely, safely converting fatal false depths into dropouts ( 4 to 0 wrong publications over 5539 held-out frames).
Controller
n
Off-axis
Range
Body
Hand
med / P95 ( ∘ )
(m)
(m 2 )
(m 2 )
Reactive
6
10.7 / 13.7
0.86
1.49
1.52
NMPC, low-motion
3
0.70 / 7.5
1.25
0.42
0.79
NMPC, deployed
3
0.54 / 2.7
1.25
1.72
1.61
TABLE I: Perception-aware control on the quadruped.
Fig. 5: The closed-loop flights of Section IV-C . (a) One flight from above, with the commanded trajectory (dashed, waypoints W1 to W10) drawn over the translucent traces: the drone by motion capture (blue), the hand-camera estimate (orange), the quadruped’s body heading (red), and the hand carrying the camera (purple). The loop begins and ends over the landing pad (W0 & W10). (b) Error against progress along the trajectory, defined as the instantaneous distance from the commanded position to the drone’s true motion-capture position (mean ± one standard deviation). Shaded vertical bands indicate the 1.0 s hold at each waypoint. Gray traces represent the baseline control error from nine flights flown entirely on motion capture. Orange traces represent markerless flights, grouped by their commanded heights (top strip), which change dynamically at each waypoint.
Drone height
n
Config.
Range
Look-up
Avail.
3D error (mm)
Reached
(m)
(m)
med ( ∘ )
(%)
med
P95
0.6 to 0.95
3
S
1.2
2 to 4
99.3 to 99.8
58 to 60
121 to 124
10/10
0.95 to 1.65
2
C
1.2
31 to 33
98.5 to 99.6
50 to 59
131 to 143
10/10
2.0
1
C
1.3
58
63.4
84
146
9/10
2.0
2
L
1.8
44 to 45
77.7 to 84.5
87 to 89
150 to 153
10/10
2.5
2
C
-
-
-
-
-
0/10
TABLE II: Flown envelope on a clear scene.
Fig. 6: Where a blind drone’s drift goes. Each trace is one of 25 human occlusions, separated by whether the drone was hovering (top, solid medians) or flying the trajectory (bottom, dashed medians). Drift is resolved on the camera’s axes at the instant of loss: motion along the optical axis affects range, cross-frame motion affects tracking angle. The along-axis median reaches 0.22 m while cross-frame drift stays under 0.08 m. Medians terminate when fewer than five occlusions remain active.
Fig. 7: Two flights behind a step ladder, against time from the first waypoint command, with top arrows indicating the direction of travel for each leg. Error is measured against motion capture (blue) and the drone’s onboard filter (green), and the gap between them is the unobservable drift while blind. Shaded regions mark when the camera lost the drone: dark where a ladder rail crossed the line of sight, light where the line stayed clear and the detector instead failed against background clutter. Red marks the safety hand-back to the reference feed.
Micro unmanned aerial vehicles (micro-UAVs) are small enough to reach confined spaces that larger robots cannot access, but too small to carry the sensing and computing power required for autonomous flight. We move the localization stack entirely off the aerial platform onto a quadruped robot with a 7-degree-of-freedom (DOF) arm, which supplies the micro-UAV (27 g bare, 42 g with fiducial markers) its full 6-DOF pose. A camera at the arm's end-effector detects AprilTag fiducial markers on the drone and composes that observation with the quadruped's own self-localization to place the drone in a shared map frame, so the ground robot localizes its partner, rather than only tracking it relative to the camera. The arm acts as an actively-controlled observer, repositioning to keep the drone in view as both robots move; the drone carries only an inertial measurement unit and fuses the external pose to fly commanded setpoints. In lab flights the external pose is accurate to 12-16 mm, enough to fly the drone autonomously within 2-5 cm of motion-capture-fed control. Having the quadruped actively follow the drone reduces the tracking error from 11.0 cm to 6.9 cm by holding the camera in the close range, where the markers are most accurate.
Alejandro Lorite Mora, Andrés Faíña
IT University of Copenhagen, Rued Langgaards Vej 7, Copenhagen 2300, Denmark · Helix Lab, Campus Kalundborg 3, 4400 Kalundborg, Denmark · Novo Nordisk A/S, Hallas Alle 1, 4400 Kalundborg, Denmark
Cooperation in multi-UAV systems requires reliable relative perception so that follower vehicles can maintain formation and continue their mission safely even when absolute positioning sensors degrade or fail. This paper presents a vision-based cooperative formation framework running on a follower UAV that uses a front-facing RGB-D camera to detect, track, and localize a leader UAV in real-time. A lightweight YOLO-based detector is trained on a dedicated drone dataset and deployed onboard to predict leader bounding boxes, which are then fused with depth information via a pinhole camera model to estimate the leader's relative pose. These estimates provide a leader-follower position controller and can also be used as a backup when GPS or external localization is unavailable. This framework is implemented as a set of ROS nodes and evaluated in a physics-based multi-UAV simulation built on XTDrone, with sensor noise and communication dropouts. We evaluate detection accuracy, runtime, and formation-keeping error under nominal conditions and under simulated failures of the positioning sensors. The results show that the proposed framework maintains stable leader-follower formations with reasonable computational cost and provides a practical basis for extending vision-based cooperative formation control to real-world multi-UAV systems.
AI, IoT and Robotics Lab (AIR Lab), UAVs Group, Autonomous Robotics Systems Limited, Hyderabad, India · Department of Microelectronics and VLSI Design, University of Hyderabad, Hyderabad, India · Department of Internet of Things, Ideabytes Software India Private Limited, Hyderabad, India +5
Aerial-ground cooperation requires real-time UAV--UGV relative-state information. Instead of maintaining global estimates for both robots, direct control in a UGV-attached non-inertial frame avoids reliance on global localization. Vision-based relative pose estimation with a passive marker offers a low-cost and effective solution. However, a fixed camera may lose sight of the moving UGV when the required UAV attitude conflicts with the field-of-view (FOV) constraint. To address this, we propose COPA, a robust active-perception framework for global-state-free aerial-ground cooperation. We use a single-axis gimbal to decouple the camera optical axis from the UAV pitch attitude. We derive an active-perception model that relates UAV motion, gimbal angle, and UGV motion to the target image-plane state.A Temporal Convolutional Network (TCN) predicts short-horizon UGV acceleration and angular velocity from recent motion history without global-state measurements. The model predictive control (MPC) uses these predictions to jointly optimize UAV and gimbal control. Simulations show that COPA maintains continuous target visibility, while ablation studies confirm that the TCN reduces peak errors during UGV motion transitions. Real-world experiments with UGV accelerations up to 3m/s^2 and yaw rates up to 1.0rad/s demonstrate robust tracking.
Mingxuan Zhang, Jiajun Yu, Baozhe Zhang +5
State Key Laboratory of Industrial Control Technology, Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou 310027, China · Huzhou Institute of Zhejiang University and Huzhou Key Laboratory of Autonomous Systems, Huzhou 313000, China · The Chinese University of Hong Kong, Shenzhen 518172, China