We propose a compact underwater acoustic classification framework combining multi-representation feature engineering, temporal statistical pooling, and compact convolutional architectures designed for acoustic time-frequency and cochlear representations. We investigate multiple conventional and auditory-inspired representations and first evaluate lightweight classifiers and Conventional Neural Networks (CNNs) on ShipsEar dataset. On the provided split, a two-layer CNN achieves a macro F1 of 0.9918, while a Radial Basis Function Support Vector Machine (RBF-SVM) reaches 0.9883. However, source-recording provenance cannot be reconstructed, preventing verification of recording-independent generalisation. We therefore evaluate on DeepShip dataset using recording-level partitioning before segmentation. Under this protocol, a 157K-parameter compact CNN achieves a test macro F1 of 0.7226, while an 11.17M-parameter ResNet18 provides no improvement in validation performance under the matched setting. These results demonstrate the importance of representation-aware feature and model design, together with rigorous recording-level evaluation, for classification performance and deployability in compact underwater acoustic systems.
Figures & tables
Figure 1: Example ShipsEar waveform and corresponding conventional and auditory-inspired acoustic representations used in this study.
Model
Pooling
Input dim.
Params
Test F1
RBF-SVM
None
60,800
–
0.9641
RBF-SVM
Mean+std
256
–
0.9883
2-layer CNN
None
302,080
1.668M
0.8987
2-layer CNN
Mean+std
5,120
183K
0.9698
Table 1: Effect of temporal statistical pooling on ShipsEar STFT classification.
Feature
Model
Params
Acc.
Macro F1
STFT
2L-CNN
183,237
0.9910
0.9918
STFT
RBF-SVM
N/A
0.9888
0.9883
CQT
2L-CNN
208,037
0.9820
0.9839
Mel
2L-CNN
208,837
0.9775
0.9789
LI
RBF-SVM
N/A
0.9707
0.9712
MFCC
RBF-SVM
N/A
0.9663
0.9677
Table 2: Selected ShipsEar classification results under the available fixed 5 second split.
Feature
Model
Params
Val. F1
Test
Mel
4L-CNN
157K
0.6519
0.7226 F1
STFT
4L-CNN
157K
0.6351
0.7129 F1
MFCC
4L-CNN
149K
0.6162
0.6735 F1
CQT
4L-CNN
173K
0.6408
0.6902 F1
BM
4L-CNN
154K
0.6766
0.6694 F1
IHC
4L-CNN
154K
0.6562
0.6336 F1
Table 3: Selected DeepShip classification results and published comparisons.
Figure 2: Class-wise mean normalized Grad-CAM relevance as a function of frequency for correct DeepShip test predictions using the four-layer STFT CNN. The frequency profiles are obtained by averaging the Grad-CAM maps over time and across predictions within each class.
Department of Electrical and Computer Engineering, Texas A&M University, College Station, Texas, 77843, USA · Massachusetts Institute of Technology Lincoln Laboratory, Lexington, Massachusetts, 02421, USA