Unsupervised Maneuver-Aware Acoustic Fault Detection for Autonomous Drones
Authors: Ali M Ali, Ziyi Tang, Nurdaulet Nazarbay, Hashim A. Hashim
Organizations: Department of Mechanical and Aerospace Engineering, Carleton University, Ottawa, ON, K1S-5B6, Canada · Department of Electrical and Computer Engineering at the University of Alberta, Edmonton, AB, Canada · The Hong Kong Polytechnic University, Hong Kong
This paper presents a maneuver-aware acoustic fault detection framework for autonomous drones that integrates Noise2Noise-inspired deep learning denoising with maneuver-conditioned reconstruction. A key practical constraint motivating this work is that labeled faulty-flight data are difficult and potentially unsafe to collect; the proposed framework therefore follows an unsupervised learning paradigm in which only nominal flight recordings are required for training. In flight environments, acoustic signals acquired from unmanned aerial vehicles are subject to variability arising both from environmental noise and from structured, maneuver-dependent aerodynamic effects. To address these challenges simultaneously, a two-stage learning architecture is developed. In the first stage, a Noise2Noise-inspired denoising model attenuates stochastic acoustic noise while preserving fault-relevant spectral-temporal structures, without requiring clean reference signals. In the second stage, a maneuver-Conditioned Convolutional AutoEncoder (maneuver-CCAE) is trained using maneuver-related labels including drone type and flight direction to model nominal acoustic behavior under varying operating conditions. Fault detection is subsequently performed using reconstruction error as an anomaly score. Experimental results demonstrate that the proposed maneuver-aware conditioning raises the area under the ROC curve (AUC) from \AUCaeOnly (unconditioned baseline) to \AUCfull (full model), validating the critical role of maneuver-dependent modeling. The complete framework is deployed on an NVIDIA Jetson Orin Nano Super embedded platform within a ROS2 pipeline, achieving an end-to-end fault detection latency of approximately 20ms per audio segment with a TensorRT half-precision (FP16) backend, confirming real-time viability for onboard UAV health monitoring.
Figures & tables
Fig. 1 : Overview of the proposed maneuver-aware acoustic fault detection framework. An onboard microphone feeds a Noise2Noise-inspired denoising module, whose output is processed by a maneuver-conditioned autoencoder; reconstruction error serves as the anomaly score.
Fig. 2 : Feature extraction pipeline. (a) Raw signal x(n) at 16 kHz. (b) Power spectrogram ∣X(f,τ)∣2 . (c) Log-mel spectrogram Xlog-mel(k,τ) .
Fig. 3 : Time-domain comparison of healthy and anomalous acoustic signals with their moving-average envelopes.
Fig. 4 : Acoustic analysis of a backward maneuver (Type-A drone). (a) Healthy log-mel spectrogram. (b) Anomalous log-mel spectrogram. (c) Difference map. (d) Fourier magnitude spectra: the anomalous condition shows spectral spreading and additional sideband components caused by propulsion system faults.
Stage
Input
Encoder 1
Encoder 2
Encoder 3
Encoder 4
Bottleneck
Decoder 1
Decoder 2
Output
Operation
Input
Conv+CBN
Conv+CBN
Conv+CBN
Conv+CBN
Conv
Upsample+Concat
Upsample+Concat
Conv
Channels
1
1→32→128
128→128
128→256→512
512→1024
1024
2048→1024→512
1024→512→256
256→1
Size
128×64
128×64
64×32
32×16
16×8
8×4
16×8
64×32
128×64
TABLE I : Denoising network architecture summary. Channel counts are given as input → output for each block; the feature-map size is halved at every encoder stage and doubled at every decoder stage, with skip connections concatenated from the corresponding encoder stage.
Stage
Parameters
FLOPs (per segment)
Latency (Orin Nano Super)
Log-mel extraction (librosa)
—
—
≈2 ms
N2N-Inspired Denoising Network
≈145.3M
≈69.9G FLOPs
≈15 ms
Maneuver-CCAE
≈1.04M
≈0.98G FLOPs
≈2 ms
Anomaly score + threshold
—
—
<1 ms
Total
≈ 146.4 M
≈ 70.9 G FLOPs
≈20 ms
TABLE II : Computational complexity of the proposed two-stage pipeline. Parameter counts are obtained directly from the released PyTorch models, and FLOPs are reported as 2× the number of multiply–accumulate operations (i.e. FLOPs=2×MACs ) for a single 128×63 log-mel segment. Latencies are means over 1000 TensorRT FP16 inference calls on the target NVIDIA Jetson Orin Nano Super (MAXN Super power mode) with the models pre-loaded in GPU memory.
Stage
Input
Embedding
Cond. Input
Encoder
Latent
Decoder
Output
Operation
Spectrogram
Conv layers
Concat with condition maps
Conv layers
Fully Connected
Conv layers
Conv
Channels
1
16→64
76
128→256→64→16
16
16→64→256→64→16→4
4→1
Size
128×64
64×32
64×32
4×2
1×1
128×64
128×64
TABLE III : Maneuver-Conditioned Convolutional Autoencoder architecture summary.
Fig. 5 : Proposed two-stage acoustic fault detection framework. Stage 1: encoder-decoder denoising network. Stage 2: maneuver-conditioned convolutional autoencoder. The reconstruction error serves as the anomaly score.
Platform
Directions
Train (normal)
Eval (normal+faulty)
Test (faulty)
Total
Type A (Holy Stone HS720)
6
1,800
360
480
2,640
Type B (MJX Bugs 12 EIS)
6
1,800
360
480
2,640
Type C (ZLRC SG906 Pro2)
6
1,800
360
480
2,640
Total
6
5,400
1,080
1,440
7,920
TABLE IV : Dataset composition by drone type, following the ICSV31 AI Challenge partition. Each file is a non-overlapping 2-second segment (recorded at 48 kHz, downsampled to 16 kHz). The train split contains normal segments only; the eval split contains both normal and anomalous segments spanning six fault types (three propeller, three motor); the test split contains anomalous segments spanning eight fault types (two additional) and, containing no normal segments, is not used for the reported detection metrics. All platforms are recorded across the six maneuver directions.
Predicted
Normal
Faulty
Actual
Normal
0.704
0.296
Faulty
0.093
0.907
TABLE VI : Row-normalized confusion matrix of the proposed full model on the evaluation split, obtained from the reported recall (TPR =0.907 ) and false-positive rate (FPR =0.296 ).
Category
Method
AUC
Accuracy
Precision
Recall
F1
Traditional ML (single)
IF-1 (Orig)
0.607
0.562
0.868
0.146
0.250
IF-1 (Denoised)
0.776
0.653
0.807
0.402
0.536
GMM-1 (Orig)
0.605
0.578
0.882
0.180
0.298
GMM-1 (Denoised)
0.782
0.651
0.796
0.406
0.537
OC-SVM-1 (Orig)
0.605
0.586
0.872
0.202
0.328
OC-SVM-1 (Denoised)
0.782
0.673
0.781
0.481
0.596
TABLE VII : Comparison of unsupervised anomaly detection methods. Orig = original log-mel spectrograms; Denoised = spectrograms processed by the proposed N2N-inspired stage. The suffix -1 denotes a single global model fitted to all operating conditions, and -18 denotes a bank of 18 condition-specific models (3 drone types × 6 maneuver directions). All methods share the same train/eval partition, log-mel features ( F×T=128×63 ), 95th-percentile threshold, and evaluation protocol.
Fig. 6 : Performance evaluation. (a) ROC curve (AUC =0.884 ). (b) Training loss convergence. (c) t-SNE latent clustering. (d) Reconstruction error distributions for nominal and faulty observations.
Passive acoustic sensing is an attractive modality for counter-unmanned aerial system (counter-UAS) defence: it is covert, low-cost, and effective against drones with small radar cross-sections or minimal radio emissions. We present EchoHawk, an open and fully reproducible reference pipeline that detects a drone from its rotor harmonics, estimates its blade-passing frequency, and localises it with a microphone array via classical wideband beamforming (delay-and-sum, MVDR, MUSIC) and time-delay processing (GCC-PHAT, SRP-PHAT), followed by temporal tracking. We evaluate the system on a physically transparent synthetic benchmark that pits drones against hard low-frequency harmonic confusers, such as ground vehicles, and on real recorded audio. Our central methodological contribution is a documented case of session-level data leakage in a widely used public dataset: because its recordings are pre-segmented into short clips, naive clip-level splits place adjacent slices of the same continuous recording in both training and test sets, inflating reported performance. Enforcing recording-session-grouped cross-validation reduces, for example, a random-forest baseline's detection probability at a 1% false-alarm rate from 0.796 to 0.745, yielding honest numbers. All code, figures, and a synthetic data generator are released so that every result runs without any download.
Microphones mounted on UAVs enable aerial acoustic scene analysis applications such as search-and-rescue, wildlife monitoring, and industrial inspection. However, drone rotor noise often dominates the mixture signal at SNRs well below -10 dB, making source recovery extremely challenging. Existing enhancement and source separation methods are typically designed for near-balanced mixtures and degrade substantially in drone audition settings. In this work, we propose DRONEAUDIONET, a drone noise suppression method that reframes a source separation model as a drone noise estimator. To better model drone-dominant mixtures, we introduce a learnable mask-scaling mechanism that allows mask magnitudes beyond unity, together with an additive residual correction term for improved drone estimation and source recovery. We train and evaluate our model on a publicly available drone audition dataset and test generalizability on an out-of-domain dataset with unseen drone hardware and flight modes. Results show that DRONEAUDIONET consistently improves downstream sound classification performance, with the largest gains observed for human vocal sounds. Our findings demonstrate the importance of drone-specific modeling for robust aerial acoustic perception and highlight the potential of source separation methods for real-world drone-assisted search-and-rescue.
Chitralekha Gupta, Soundarya Ramesh, Yifei Luo +1
School of Computing, National University of Singapore
Multi-rotor aerial autonomous vehicles (MAVs, more widely known as "drones") have been generating increased interest in recent years due to their growing applicability in a vast and diverse range of fields (e.g., agriculture, commercial delivery, search and rescue). The sensitivity of visual-based methods to lighting conditions and occlusions had prompted growing study of navigation reliant on other modalities, such as acoustic sensing. A major concern in using drones in scale for tasks in non-controlled environments is the potential threat of adversarial attacks over their navigational systems, exposing users to mission-critical failures, security breaches, and compromised safety outcomes that can endanger operators and bystanders. While previous work shows impressive progress in acoustic-based drone localization, prior research in adversarial attacks over drone navigation only addresses visual sensing-based systems. In this work, we aim to compensate for this gap by supplying a comprehensive analysis of the effect of PGD adversarial attacks over acoustic drone localization. We furthermore develop an algorithm for adversarial perturbation recovery, capable of markedly diminishing the affect of such attacks in our setting.
Tamir Shor, Chaim Baskin, Alex Bronstein
Department of Computer Science Technion – Israel Institute of Technology Haifa, Israel · School of Electrical and Computer Engineering Ben-Gurion University of the Negev Be’er Sheva, Israel · Technion – Israel Institute of Technology +1