Reliability-aware short-term roll prediction for unmanned surface vehicles via multi-task learning and adaptive centralization
Authors: Kaizhen Li, Xi Zhou, Zihao Wang, Dan Zhang, Jianjian Liu, Xiaowei Li
Organizations: School of Mechatronic Engineering and Automation, Shanghai University, Shanghai, 200444, China · Shanghai Artificial Intelligence Laboratory, Shanghai, 200032, China · Institute of Artificial Intelligence, Shanghai University, Shanghai, 200444, China
Reliable roll prediction of unmanned surface vehicles (USVs) is essential for ensuring navi?gational safety and enhancing autonomous decision-making. While existing studies primarily focus on improving prediction accuracy, the quantification of prediction reliability remains insufficiently addressed. To bridge this gap, this paper proposes a reliability-aware prediction paradigm that integrates confidence assessment into the predictive pipeline. The architecture utilizes a multi-task learning structure where a shared feature extraction backbone feeds into dual heads: a regression head for precise roll prediction and a quantification head for confidence scoring. This configuration provides accurate prediction and corresponding confidence for risk?sensitive downstream tasks. In addition, an adaptive centralization strategy tailored for short?term real-time roll prediction is introduced to improve model generalization under varying operational conditions. Experiments conducted on a real-sea dataset demonstrate that the proposed method effectively quantifies the reliability of prediction results and maintains superior generalization under varying conditions, offering significant potential for practical engineering applications.
Figures & tables
Figure 1 : Overall structure
Figure 2 : Adaptive centralization
Figure 3 : Structure of Bi-LSTM
Figure 4 : Consistency effect
Dimensions
Symbol
Value
Length between perpendiculars(m)
Lpp
7.5
Waterline breadth(m)
BWL
2.6
Depth(m)
D
1.61
Draft(m)
T
0.54
Displacement( m3 )
∇
3.62
Longitudinal center of gravity (m)
LGG
2.45
Table 1 : USV dimension parameters
Figure 5 : Analysis of USV roll motion: (a) Time-series overview and (b) statistical distribution.
Figure 6 : Distribution drift
Dataset
Sailing Condition
Sea State
Wave encounter
Sample Number
Mean(deg)
Max(deg)
Min(deg)
#0
Initial Condition
4
Beam Sea
28701
3.151
14.660
-13.550
#1
Altered Condition
4
Oblique Sea
6651
2.201
13.730
-8.621
#2
Altered Condition
3
Oblique Sea
8401
0.547
9.590
-7.340
Table 2 : Multi-condition dataset description
Figure 7 : Sliding window method
Figure 8 : label smoothing
Module
Parameter
Value
Adaptive module
Kernel size of moving_average
3
Hidden dimension of linear
[128,128]
Feature extract
Bi-lstm layers
3
Hidden dimension of Bi-lstm
[128,128]
Activation function
ReLU
Dropout rate
0.3
Table 3 : Model paremeters configuration
Figure 9 : Comparison between No-standardization, Z-Score and proposed AdaC
Figure 10 : Comparison among advanced normalization methods, including RevIN, Dish-ts, San, and the proposed AdaC.
Dataset
Metric
No-standardization
Z-Score
RevIN
Dish-ts
San
AdaC (ours)
#1
RMSE
2.470 ± 0.05
2.383 ± 0.03
2.110 ± 0.005
2.477 ± 0.08
2.323 ± 0.06
2.065 ± 0.02
MAE
1.703 ± 0.03
1.646 ± 0.03
1.448 ± 0.005
1.705 ± 0.05
1.607 ± 0.04
1.411 ± 0.02
#2
RMSE
2.720 ± 0.09
2.401 ± 0.07
1.377 ± 0.03
2.448 ± 0.22
2.196 ± 0.16
1.335 ± 0.02
MAE
1.952 ± 0.07
1.711 ± 0.06
0.936 ± 0.02
1.709 ± 0.16
1.537 ± 0.12
0.919 ± 0.01
Table 4 : RMSE and MAE ( mean ± std) of different preprocessing strategies under altered conditions
Figure 11 : Prediction error analysis under different forecasting horizons.
Figure 12 : Ablation study on different backbone feature extractors
Figure 13 : Confidence assessment over each step of prediction
Model
80%
90%
95%
ND
0.686±0.28
0.800±0.24
0.871±0.21
TD
0.636±0.28
0.746±0.25
0.827±0.22
HQ
0.815±0.24
0.895±0.19
0.940±0.14
Table 5 : PICP and AWD ( mean ± std) under different confidence levels
Figure 14 : The comparison of confidence interval prediction at different confidence levels
Figure 15 : Comparison of prediction interval coverage probability (PICP) for ND, TD, and the proposed HQ method under different nominal confidence levels (80%, 90%, and 95%).
Figure 16 : Comparison of average width deviation (AWD) for ND, TD, and the proposed HQ method under different nominal confidence levels (80%, 90%, and 95%).
Figure 17 : Parity plot (predicted vs. ground truth) with the identity line (y=x); point transparency indicates the predicted confidence (higher opacity denotes higher confidence)
Figure 18 : The relationship of Error and Confidence
Figure 19 : Ablation results with and without the proposed consistency loss.
In/Output
20-5
20-10
20-15
20-20
Without Cons
0.327
1.104
1.831
2.214
With Cons
0.340
1.143
1.840
2.175
Rate(%)
3.97
3.53
0.49
-1.76
Table 6 : Variation of RMSE with and without the consistency loss
Model
RMSE
MAE
p -value
Significant
p -value
Significant
No-standardization
0.0006
Yes
0.0013
Yes
Z-Score
0.0016
Yes
0.0015
Yes
RevIN
0.0413
Yes
0.2529
No
Dish-ts
0.0093
Yes
0.0083
Yes
San
0.0097
Yes
0.0103
Yes
Table 7 : Significance analysis of AdaC against different normalization strategies.
Confidence level
Comparison
PICP
AWD
p -value
Significant
p -value
Significant
80%
ND
3.78×10−21
Yes
1.34×10−21
Yes
TD
1.81×10−32
Yes
2.85×10−31
Yes
90%
ND
1.17×10−17
Yes
3.10×10−17
Yes
TD
1.23×10−30
Yes
2.25×10−28
Yes
95%
ND
2.07×10−14
Yes
4.91×10−14
Yes
Table 8 : Significance analysis of HQ against ND and TD under different confidence levels.
School of Information Engineering, Henan University of Science and Technology, Luoyang 471023, China · Shenzhen Research Institute of Big Data, Shenzhen 518172, China · Chinese University of Hong Kong, Shenzhen, Shenzhen 518172, China +5
Institute of Optics and Electronics, Chinese Academy of Sciences, Chengdu, China · Institute for Infocomm Research (I2R), Agency for Science, Technology and Research, Singapore · Shenzhen Astralldynamics Technology