Interpretable Fuzzy Inference for UAV Target Tracking Using Bounding-Box Geometry
Authors: Reza Ahmari, Ahmad Mohammadi, Vahid Hemmati, Nicholas Edmond, Hossein Z. Saghazadeh, Olusola Odeyomi, Parham Kebria, Abdollah Homaifar
Organizations: Department of Computer Science, North Carolina A&T State University, 1601 E Market St, Greensboro, 27411, NC, USA. · Department of Electrical and Computer Engineering, North Carolina A&T State University, 1601 E Market St, Greensboro, 27411, NC, USA.
Vision-based guidance of unmanned aerial vehicles (UAVs) toward unmanned ground vehicles (UGVs) supports cooperative aerial--ground robotics, but reliable continuous yaw estimation from onboard vision remains challenging because of sensing uncertainty, limited computation, and the need for interpretable control. Existing deep-learning and geometric-reconstruction approaches often require large datasets, external localization, or complex modeling assumptions, reducing transparency and deployment suitability on resource-constrained platforms. We present an interpretable fuzzy-inference framework that generates continuous yaw commands from low-dimensional features extracted from YOLO boxes: target centroid location, area, and aspect ratio. No explicit geometric modeling is required. A Mamdani fuzzy system serves as an interpretable baseline using a shoulder--triangle--shoulder input partition. It is followed by a first-order Takagi--Sugeno model with three antecedent membership terms per input, whose parameters are derived from training-set quantiles, yielding a compact 27-rule structure. Evaluation uses 6{,}169 labeled samples from a VICON motion-capture environment. Across five randomized train--test splits, the Takagi--Sugeno model achieves a test-set mean absolute error of 0.140∘±0.003∘, a root mean squared error of 0.200∘±0.008∘, and a maximum absolute error of 1.254∘±0.121∘. Within-threshold accuracies are 99.676 for ±1∘ and 100.000 for both ±3∘ and ±5∘. Directional consistency between image-plane horizontal displacement and predicted yaw sign reaches 90.254. These results show that the framework is transparent, data-efficient, computationally lightweight, and suitable for real-time vision-based UAV guidance toward mobile ground targets.