Unsupervised Maneuver-Aware Acoustic Fault Detection for Autonomous Drones
Authors: Ali M Ali, Ziyi Tang, Nurdaulet Nazarbay, Hashim A. Hashim
Organizations: Department of Mechanical and Aerospace Engineering, Carleton University, Ottawa, ON, K1S-5B6, Canada · Department of Electrical and Computer Engineering at the University of Alberta, Edmonton, AB, Canada · The Hong Kong Polytechnic University, Hong Kong
This paper presents a maneuver-aware acoustic fault detection framework for autonomous drones that integrates Noise2Noise-inspired deep learning denoising with maneuver-conditioned reconstruction. A key practical constraint motivating this work is that labeled faulty-flight data are difficult and potentially unsafe to collect; the proposed framework therefore follows an unsupervised learning paradigm in which only nominal flight recordings are required for training. In flight environments, acoustic signals acquired from unmanned aerial vehicles are subject to variability arising both from environmental noise and from structured, maneuver-dependent aerodynamic effects. To address these challenges simultaneously, a two-stage learning architecture is developed. In the first stage, a Noise2Noise-inspired denoising model attenuates stochastic acoustic noise while preserving fault-relevant spectral-temporal structures, without requiring clean reference signals. In the second stage, a maneuver-Conditioned Convolutional AutoEncoder (maneuver-CCAE) is trained using maneuver-related labels including drone type and flight direction to model nominal acoustic behavior under varying operating conditions. Fault detection is subsequently performed using reconstruction error as an anomaly score. Experimental results demonstrate that the proposed maneuver-aware conditioning raises the area under the ROC curve (AUC) from \AUCaeOnly (unconditioned baseline) to \AUCfull (full model), validating the critical role of maneuver-dependent modeling. The complete framework is deployed on an NVIDIA Jetson Orin Nano Super embedded platform within a ROS2 pipeline, achieving an end-to-end fault detection latency of approximately 20ms per audio segment with a TensorRT half-precision (FP16) backend, confirming real-time viability for onboard UAV health monitoring.
Figures & tables
Fig. 1 : Overview of the proposed maneuver-aware acoustic fault detection framework. An onboard microphone feeds a Noise2Noise-inspired denoising module, whose output is processed by a maneuver-conditioned autoencoder; reconstruction error serves as the anomaly score.
Fig. 2 : Feature extraction pipeline. (a) Raw signal x(n) at 16 kHz. (b) Power spectrogram ∣X(f,τ)∣2 . (c) Log-mel spectrogram Xlog-mel(k,τ) .
Fig. 3 : Time-domain comparison of healthy and anomalous acoustic signals with their moving-average envelopes.
Fig. 4 : Acoustic analysis of a backward maneuver (Type-A drone). (a) Healthy log-mel spectrogram. (b) Anomalous log-mel spectrogram. (c) Difference map. (d) Fourier magnitude spectra: the anomalous condition shows spectral spreading and additional sideband components caused by propulsion system faults.
Stage
Input
Encoder 1
Encoder 2
Encoder 3
Encoder 4
Bottleneck
Decoder 1
Decoder 2
Output
Operation
Input
Conv+CBN
Conv+CBN
Conv+CBN
Conv+CBN
Conv
Upsample+Concat
Upsample+Concat
Conv
Channels
1
1→32→128
128→128
128→256→512
512→1024
1024
2048→1024→512
1024→512→256
256→1
Size
128×64
128×64
64×32
32×16
16×8
8×4
16×8
64×32
128×64
TABLE I : Denoising network architecture summary. Channel counts are given as input → output for each block; the feature-map size is halved at every encoder stage and doubled at every decoder stage, with skip connections concatenated from the corresponding encoder stage.
Stage
Parameters
FLOPs (per segment)
Latency (Orin Nano Super)
Log-mel extraction (librosa)
—
—
≈2 ms
N2N-Inspired Denoising Network
≈145.3M
≈69.9G FLOPs
≈15 ms
Maneuver-CCAE
≈1.04M
≈0.98G FLOPs
≈2 ms
Anomaly score + threshold
—
—
<1 ms
Total
≈ 146.4 M
≈ 70.9 G FLOPs
≈20 ms
TABLE II : Computational complexity of the proposed two-stage pipeline. Parameter counts are obtained directly from the released PyTorch models, and FLOPs are reported as 2× the number of multiply–accumulate operations (i.e. FLOPs=2×MACs ) for a single 128×63 log-mel segment. Latencies are means over 1000 TensorRT FP16 inference calls on the target NVIDIA Jetson Orin Nano Super (MAXN Super power mode) with the models pre-loaded in GPU memory.
Stage
Input
Embedding
Cond. Input
Encoder
Latent
Decoder
Output
Operation
Spectrogram
Conv layers
Concat with condition maps
Conv layers
Fully Connected
Conv layers
Conv
Channels
1
16→64
76
128→256→64→16
16
16→64→256→64→16→4
4→1
Size
128×64
64×32
64×32
4×2
1×1
128×64
128×64
TABLE III : Maneuver-Conditioned Convolutional Autoencoder architecture summary.
Fig. 5 : Proposed two-stage acoustic fault detection framework. Stage 1: encoder-decoder denoising network. Stage 2: maneuver-conditioned convolutional autoencoder. The reconstruction error serves as the anomaly score.
Platform
Directions
Train (normal)
Eval (normal+faulty)
Test (faulty)
Total
Type A (Holy Stone HS720)
6
1,800
360
480
2,640
Type B (MJX Bugs 12 EIS)
6
1,800
360
480
2,640
Type C (ZLRC SG906 Pro2)
6
1,800
360
480
2,640
Total
6
5,400
1,080
1,440
7,920
TABLE IV : Dataset composition by drone type, following the ICSV31 AI Challenge partition. Each file is a non-overlapping 2-second segment (recorded at 48 kHz, downsampled to 16 kHz). The train split contains normal segments only; the eval split contains both normal and anomalous segments spanning six fault types (three propeller, three motor); the test split contains anomalous segments spanning eight fault types (two additional) and, containing no normal segments, is not used for the reported detection metrics. All platforms are recorded across the six maneuver directions.
Predicted
Normal
Faulty
Actual
Normal
0.704
0.296
Faulty
0.093
0.907
TABLE VI : Row-normalized confusion matrix of the proposed full model on the evaluation split, obtained from the reported recall (TPR =0.907 ) and false-positive rate (FPR =0.296 ).
Category
Method
AUC
Accuracy
Precision
Recall
F1
Traditional ML (single)
IF-1 (Orig)
0.607
0.562
0.868
0.146
0.250
IF-1 (Denoised)
0.776
0.653
0.807
0.402
0.536
GMM-1 (Orig)
0.605
0.578
0.882
0.180
0.298
GMM-1 (Denoised)
0.782
0.651
0.796
0.406
0.537
OC-SVM-1 (Orig)
0.605
0.586
0.872
0.202
0.328
OC-SVM-1 (Denoised)
0.782
0.673
0.781
0.481
0.596
TABLE VII : Comparison of unsupervised anomaly detection methods. Orig = original log-mel spectrograms; Denoised = spectrograms processed by the proposed N2N-inspired stage. The suffix -1 denotes a single global model fitted to all operating conditions, and -18 denotes a bank of 18 condition-specific models (3 drone types × 6 maneuver directions). All methods share the same train/eval partition, log-mel features ( F×T=128×63 ), 95th-percentile threshold, and evaluation protocol.
Fig. 6 : Performance evaluation. (a) ROC curve (AUC =0.884 ). (b) Training loss convergence. (c) t-SNE latent clustering. (d) Reconstruction error distributions for nominal and faulty observations.
Department of Computer Science Technion – Israel Institute of Technology Haifa, Israel · School of Electrical and Computer Engineering Ben-Gurion University of the Negev Be’er Sheva, Israel · Technion – Israel Institute of Technology +1