Agile multi-UAV flight requires accurate and low-latency onboard estimation of the kinematic states of neighboring UAVs for collision avoidance, motion coordination, etc. Most vision-based approaches rely on position-only measurements, inferring velocity and acceleration indirectly from displacement. We show that this introduces a fixed structural delay in the estimation of higher-order states, which limits the achievable agility. To address this, we propose to integrate tilt measurements, provided by a state-of-the-art visual detector, which inform about the thrust direction of co-planar multirotor UAVs. We benchmark four position-only and five pose-aware estimators, including a novel formulation of a linear thrust-constraining Kalman filter, on two real-world and one high-fidelity photorealistic simulated dataset over different levels of agility (3-21 m/s^2). In our setup, pose-aware estimation consistently reduces the average velocity and acceleration estimation errors by 40% and 57% across the three datasets with the proposed KF formulation outperforming the other estimators. Position-only filters exhibit a constant ~300 ms delay in acceleration step response independent of agility, whereas the tilt-constrained estimators operate near the physical response limit given by the camera frame-rate by observing the change in thrust direction before the displacement accumulates. In a closed-loop leader-follower simulated experiment with NMPC control, position-only estimation of the leader's state fails to facilitate stable hovering of the follower, while the proposed estimator enables tracking of lateral maneuvers exceeding 2g of acceleration.
Figures & tables
Fig. 1 : Onboard view from a follower UAV during closed-loop tracking at accelerations above 2g . Red boxes show raw YOLOv5-6D detections. The first frame ( t1 ) is the current image, the remaining frames ( t2 – t6 ) are a time-lapse during an aggressive lateral maneuver. The times t1 – t6 correspond to the markers in Fig. 5 .
Fig. 2 : Relation of the world frame W and the uav ’s body frame B through the position vector p and orientation matrix R . Basis vectors of these frames and the relevant acceleration components of the total uav acceleration a=aT+ad+ag are also denoted.
Dataset
Unreal (ours)
Phantom 4 [ 10 ]
Mavic 2 [ 10 ]
Number of sequences
140
68
61
Total duration
1333s
667s
668s
Number of frames
33460
16734
16755
Sampling period
40ms
40ms
40ms
Camera motion
Flying
Static
Static
Camera range ∗
1.58.0m
1.95.1m
1.95.1m
TABLE I : Comparison of the benchmarking datasets.
Unreal (simulated)
Phantom 4 (real-world)
Mavic 2 (real-world)
Estimator
Pos. ( m )
Vel. ( ms−1 )
Acc. ( ms−2 )
Pos. ( m )
Vel. ( ms−1 )
Acc. ( ms−2 )
Pos. ( m )
Vel. ( ms−1 )
Acc. ( ms−2 )
Position-only
CV-KF
0.100
2.300
–
0.131
0.755
–
0.070
0.793
–
CV-KF+BDC
0.112
0.739
–
0.131
0.415
–
0.071
0.426
–
CA-KF
0.104
0.791
4.899
0.131
0.487
2.337
0.093
0.609
2.440
CA-KF+BDC
0.099
0.787
4.989
0.131
0.415
2.039
0.078
0.488
2.102
TABLE II : men values for position, velocity, and acceleration across datasets and estimators.
Fig. 3 : Velocity and acceleration men for the different agility levels in the Unreal dataset. cv - kf is omitted due to a significantly higher error at all levels (see Table II ). Dots denote mean over 10 trajectories, lines indicate standard deviation.
Agility group
Low
Mid
High
Reference max. ( ms−2 )
3.0
6.0
9.0
12.0
15.0
18.0
21.0
Flown acc. Q2 † ( ms−2 )
1.6
3.0
4.0
5.7
7.0
7.9
10.0
Flown acc. p98 † ( ms−2 )
2.8
5.7
7.9
11.1
14.0
15.9
19.1
Flown tilt p98 † ( \SIUnitSymbolDegree )
15
32
45
62
71
86
95
Num. of sequences
20
20
20
20
20
20
20
† Q2, p98 denote the 50% and 98% percentiles.
TABLE III : Agility levels in the Unreal (simulated) dataset.
Low (3–9)
Mid (12–15)
High (18–21)
Estimator
xy
z
xy
z
xy
z
Velocity men ( ms−1 )
CA-KF+BDC
0.376
0.248
0.393
0.306
0.553
0.404
Pose-KF [ 10 ]
0.194
0.220
0.219
0.332
0.312
0.444
Z-KF+BDC
0.136
0.207
0.192
0.260
0.259
0.330
Relative improvement of Z-KF+BDC over CA-KF+BDC
TABLE IV : Per-axis men values on the Unreal dataset across the different agility groups.
Fig. 4 : Acceleration step response at 9ms−2 magnitude. Shaded regions indicate variance across five runs. Crosses represent camera frames.
aref
τgt
τCA-KF+BDC
τZ-KF+BDC
( ms−2 )
ms
frames
ms
frames
ms
frames
3
152.0
3.8
456.4
11.4
156.0
3.9
6
159.2
4.0
482.9
12.1
208.4
5.2
9
164.4
4.1
477.6
11.9
247.3
6.2
12
172.1
4.3
474.8
11.9
292.1
7.3
15
180.6
4.5
476.1
11.9
327.0
8.2
TABLE V : Step response delay over reference magnitude.
Fig. 5 : Closed-loop leader-follower tracking with zlkf and NMPC. Note that with a position-only estimator, the follower is not even capable of stable hovering (not shown). Markers t1 – t6 reference Fig. 1 . (a) Position. (b) Velocity. (c) Acceleration.
Short-horizon prediction is essential for electro-optical UAV tracking, especially when the target is small, maneuvering, or intermittently observed. Image center, line-of-sight, and range measurements provide direct constraints on target position, but their constraints on acceleration are weak. As a result, prediction can lag during aggressive maneuvers. This paper proposes an image-domain tilt constrained distributed fusion method for maneuvering UAV tracking. The method uses the apparent roll and pitch of a rotorcraft target in the image as low-level maneuver cues. A weak-prior auto-labeling pipeline first generates oriented bounding box and image-domain tilt labels from synchronized video, gimbal IMU, and UAV IMU data. A YOLO-OBB detector is then trained to provide online target position and tilt measurements. The front-end Python implementation is publicly available at github.com/ShineMinxing/PythonYOLO. In the fusion stage, the UAV state is modeled by position, velocity, and acceleration. Image-domain roll and pitch are introduced as acceleration-related pseudo-observations. For distributed tracking, one mobile gimbal camera and two fixed ground cameras are fused asynchronously. Camera attitude error states are augmented into the filter to absorb extrinsic drift and cross-camera systematic inconsistency. A Mahalanobis gate with time-since-last-valid covariance widening is used to reject false detections and handle dropouts. In simulation, adding roll/pitch observations reduces the prediction RMSE from 1.991 m to 0.821 m and decreases the cumulative prediction error by 60.75%. In real distributed experiments, a self-consistency evaluation shows an 18.10% reduction in cumulative prediction error. The results show that image-domain tilt can provide useful acceleration constraints for robust short-horizon UAV prediction.
Minxing Sun, Yao Mao
Institute of Optics and Electronics, Chinese Academy of Sciences, Chengdu, China · Institute for Infocomm Research (I2R), Agency for Science, Technology and Research, Singapore · Shenzhen Astralldynamics Technology
Aerial-ground cooperation requires real-time UAV--UGV relative-state information. Instead of maintaining global estimates for both robots, direct control in a UGV-attached non-inertial frame avoids reliance on global localization. Vision-based relative pose estimation with a passive marker offers a low-cost and effective solution. However, a fixed camera may lose sight of the moving UGV when the required UAV attitude conflicts with the field-of-view (FOV) constraint. To address this, we propose COPA, a robust active-perception framework for global-state-free aerial-ground cooperation. We use a single-axis gimbal to decouple the camera optical axis from the UAV pitch attitude. We derive an active-perception model that relates UAV motion, gimbal angle, and UGV motion to the target image-plane state.A Temporal Convolutional Network (TCN) predicts short-horizon UGV acceleration and angular velocity from recent motion history without global-state measurements. The model predictive control (MPC) uses these predictions to jointly optimize UAV and gimbal control. Simulations show that COPA maintains continuous target visibility, while ablation studies confirm that the TCN reduces peak errors during UGV motion transitions. Real-world experiments with UGV accelerations up to 3m/s^2 and yaw rates up to 1.0rad/s demonstrate robust tracking.
Mingxuan Zhang, Jiajun Yu, Baozhe Zhang +5
State Key Laboratory of Industrial Control Technology, Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou 310027, China · Huzhou Institute of Zhejiang University and Huzhou Key Laboratory of Autonomous Systems, Huzhou 313000, China · The Chinese University of Hong Kong, Shenzhen 518172, China
A precise state estimate is crucial for a tight feedback control that enables agile and near-obstacle flights of UAVs. The state-of-the-art methods fuse slow pose measurements with high-frequency inertial measurements to obtain a precise state estimate. However, the inertial measurements from the IMU onboard the UAV are degraded by vibrations from spinning propellers and the precision of the estimated state suffers. We propose a novel approach based on the preintegration of accelerations obtained from motor speeds. We show that the accelerations obtained in this manner can be used for state propagation on their own to achieve better precision without including the IMU. Further, we propose a factor composed of the preintegrated motor speeds that can be directly employed in factor graph optimization frameworks. We combine our factor with LiDAR measurements into the proposed Motor Angular Speed LiDAR Odometry (MAS-LO) algorithm for precise state estimation, which we open-source. Lastly, we evaluate the estimation precision against a state-of-the-art inertial algorithm LIO-SAM to show 28% improvement in position and 65% in velocity estimation accuracy, 14% lower measurement lag, and high robustness to wrong parameter values.
Matěj Petrlík, Filip Novák, Robert Pěnička +1
Department of Cybernetics, Faculty of Electrical Engineering, Czech Technical University in Prague, 166 36, Prague 6, Czech Republic