eess.SPOct 7, 2026

Robust Prediction of Internal Wave-Affected Multi-Scale Sound Speed Distribution Using Lightweight Kolmogorov-Arnold Networks with Hybrid Basis Functions

Authors: Wei Huang, Junpeng Lu, Tianhe Xu, Hao Zhang, Feng Yin

Organizations: Faculty of Information Science and Engineering, Ocean University of China, Qingdao, Shandong 266404, China · School of Space Science and Technology, Shandong University at Weihai, Weihai, Shandong 264200, China · School of Artificial Intelligence, The Chinese University of Hong Kong (Shenzhen), Shenzhen, Guangdong 518172, China

Abstract

The underwater sound speed distribution directly governs acoustic propagation paths, rendering it critically important for underwater acoustic communication and target localization. Conventional sound speed profile (SSP) prediction methods provide a good way to estimate the underwater sound speed distribution without on-site data measurement, thus breaking through the coverage area constraints of sonar observation equipment and making the model universal in most marine areas. However, underwater sound speed exhibits multi-scale variations, such as diurnal, quarterly, and intermittent fluctuations caused by ocean processes such as internal waves. This makes it difficult for the fixed structure models in existing methods to have good generalization ability for multi-scale sound speed distribution patterns. To tackle this problem, we proposed a lightweight hybrid basis-function empowered Kolmogorov-Arnold network (LHBF-KAN) model for multi-scale sound speed prediction. We aim to construct a multi-branch representation layer in which different basis functions respond to distinct temporal patterns, from slowly varying background trends to rapid fluctuations induced by dynamic ocean processes, allowing the model to naturally accommodate the inherently multi-scale evolution of sound speed at different depths. To prevent the multi-branch structure from increasing model size, a pruning strategy is further introduced to suppress branches with consistently low contribution during training, yielding a compact architecture, suitable for deployment on resource constrained underwater platforms.

Figures & tables

Explore similar work

Sep 24, 2026cs.SD

Towards Deployable Underwater Vessel Classification

We propose a compact underwater acoustic classification framework combining multi-representation feature engineering, temporal statistical pooling, and compact convolutional architectures designed for acoustic time-frequency and cochlear representations. We investigate multiple conventional and auditory-inspired representations and first evaluate lightweight classifiers and Conventional Neural Networks (CNNs) on ShipsEar dataset. On the provided split, a two-layer CNN achieves a macro F1 of 0.9918, while a Radial Basis Function Support Vector Machine (RBF-SVM) reaches 0.9883. However, source-recording provenance cannot be reconstructed, preventing verification of recording-independent generalisation. We therefore evaluate on DeepShip dataset using recording-level partitioning before segmentation. Under this protocol, a 157K-parameter compact CNN achieves a test macro F1 of 0.7226, while an 11.17M-parameter ResNet18 provides no improvement in validation performance under the matched setting. These results demonstrate the importance of representation-aware feature and model design, together with rigorous recording-level evaluation, for classification performance and deployability in compact underwater acoustic systems.
May 29, 2026eess.SP

Neural Radiated-Noise Fields for Unmanned Underwater Vehicle Noise Spectrum Prediction in Three-Dimensional Scenes

Radiated noise in unmanned underwater vehicles (UUVs) is an important indicator for characterizing acoustic signatures and evaluating platform performance. To address the strong dependence of traditional physics-based modeling and numerical simulation methods on target structural information and environmental boundary conditions, and their inability to achieve continuous spatial spectrum-response modeling in three-dimensional scenes, this paper proposes a neural radiated-noise field (NRNF). An NRNF represents the UUV radiated-noise spectrum as a continuous function of the three-dimensional UUV position, the three-dimensional hydrophone position, the UUV yaw angle, and the frequency, enabling query-based prediction at arbitrary spatial locations. The proposed method employs sinusoidal encoding for position and frequency, and introduces a learnable three-dimensional scene feature grid to explicitly represent environmental structure and propagation effects. A spectrum-prediction dataset is constructed from lake trials, and the proposed model is evaluated under three settings: horizontal extrapolation, depth extrapolation, and cross-run generalization. Results show that the NRNF achieves an average prediction error of 3.5 dB in the 50 to 5000 Hz band. Horizontal extrapolation is easiest, depth extrapolation is the most challenging, and cross-run generalization is of intermediate difficulty. Further ablation results demonstrate that the scene feature grid significantly improves the prediction stability and spatial generalization of the model.
May 6, 2026cs.SD

Hearing the Ocean: Bio-inspired Gammatone-CNN framework for Robust Underwater Acoustic Target Classification

This study presents a bio inspired signal processing framework for robust Underwater Acoustic Target Recognition (UATR). The latest state of the art methods often fail to resolve dense low frequency harmonic structures in vessel propulsion signals under high noise conditions, which is addressed by the proposed framework using a biologically inspired Gammatone filter bank that emulates the cochlea nonlinear frequency selectivity. By distributing filters according to the Equivalent Rectangular Bandwidth (ERB) scale, the framework achieves a high fidelity representation of engine radiated tonals while effectively suppressing isotropic ambient interference. The resulting Cochleagram features are processed by a lightweight, custom designed Convolutional Neural Network (CNN) that leverages large receptive fields to integrate spectral-temporal continuities. Experimental results on the VTUAD dataset demonstrate a state of the art classification accuracy of 98.41%, outperforming Continuous Wavelet Transform and Mel Frequency Cepstral Coefficients baselines by 3.5% and 7.7% respectively. Furthermore, the framework achieves an inference latency of only 0.77 ms and a 0.971 Cohen Kappa score, validating its efficacy for real time deployment on autonomous, low-power sonar hardware.