Reliability-aware short-term roll prediction for unmanned surface vehicles via multi-task learning and adaptive centralization
Authors: Kaizhen Li, Xi Zhou, Zihao Wang, Dan Zhang, Jianjian Liu, Xiaowei Li
Organizations: School of Mechatronic Engineering and Automation, Shanghai University, Shanghai, 200444, China · Shanghai Artificial Intelligence Laboratory, Shanghai, 200032, China · Institute of Artificial Intelligence, Shanghai University, Shanghai, 200444, China
Reliable roll prediction of unmanned surface vehicles (USVs) is essential for ensuring navi?gational safety and enhancing autonomous decision-making. While existing studies primarily focus on improving prediction accuracy, the quantification of prediction reliability remains insufficiently addressed. To bridge this gap, this paper proposes a reliability-aware prediction paradigm that integrates confidence assessment into the predictive pipeline. The architecture utilizes a multi-task learning structure where a shared feature extraction backbone feeds into dual heads: a regression head for precise roll prediction and a quantification head for confidence scoring. This configuration provides accurate prediction and corresponding confidence for risk?sensitive downstream tasks. In addition, an adaptive centralization strategy tailored for short?term real-time roll prediction is introduced to improve model generalization under varying operational conditions. Experiments conducted on a real-sea dataset demonstrate that the proposed method effectively quantifies the reliability of prediction results and maintains superior generalization under varying conditions, offering significant potential for practical engineering applications.
Figures & tables
Figure 1 : Overall structure
Figure 2 : Adaptive centralization
Figure 3 : Structure of Bi-LSTM
Figure 4 : Consistency effect
Dimensions
Symbol
Value
Length between perpendiculars(m)
Lpp
7.5
Waterline breadth(m)
BWL
2.6
Depth(m)
D
1.61
Draft(m)
T
0.54
Displacement( m3 )
∇
3.62
Longitudinal center of gravity (m)
LGG
2.45
Table 1 : USV dimension parameters
Figure 5 : Analysis of USV roll motion: (a) Time-series overview and (b) statistical distribution.
Figure 6 : Distribution drift
Dataset
Sailing Condition
Sea State
Wave encounter
Sample Number
Mean(deg)
Max(deg)
Min(deg)
#0
Initial Condition
4
Beam Sea
28701
3.151
14.660
-13.550
#1
Altered Condition
4
Oblique Sea
6651
2.201
13.730
-8.621
#2
Altered Condition
3
Oblique Sea
8401
0.547
9.590
-7.340
Table 2 : Multi-condition dataset description
Figure 7 : Sliding window method
Figure 8 : label smoothing
Module
Parameter
Value
Adaptive module
Kernel size of moving_average
3
Hidden dimension of linear
[128,128]
Feature extract
Bi-lstm layers
3
Hidden dimension of Bi-lstm
[128,128]
Activation function
ReLU
Dropout rate
0.3
Table 3 : Model paremeters configuration
Figure 9 : Comparison between No-standardization, Z-Score and proposed AdaC
Figure 10 : Comparison among advanced normalization methods, including RevIN, Dish-ts, San, and the proposed AdaC.
Dataset
Metric
No-standardization
Z-Score
RevIN
Dish-ts
San
AdaC (ours)
#1
RMSE
2.470 ± 0.05
2.383 ± 0.03
2.110 ± 0.005
2.477 ± 0.08
2.323 ± 0.06
2.065 ± 0.02
MAE
1.703 ± 0.03
1.646 ± 0.03
1.448 ± 0.005
1.705 ± 0.05
1.607 ± 0.04
1.411 ± 0.02
#2
RMSE
2.720 ± 0.09
2.401 ± 0.07
1.377 ± 0.03
2.448 ± 0.22
2.196 ± 0.16
1.335 ± 0.02
MAE
1.952 ± 0.07
1.711 ± 0.06
0.936 ± 0.02
1.709 ± 0.16
1.537 ± 0.12
0.919 ± 0.01
Table 4 : RMSE and MAE ( mean ± std) of different preprocessing strategies under altered conditions
Figure 11 : Prediction error analysis under different forecasting horizons.
Figure 12 : Ablation study on different backbone feature extractors
Figure 13 : Confidence assessment over each step of prediction
Model
80%
90%
95%
ND
0.686±0.28
0.800±0.24
0.871±0.21
TD
0.636±0.28
0.746±0.25
0.827±0.22
HQ
0.815±0.24
0.895±0.19
0.940±0.14
Table 5 : PICP and AWD ( mean ± std) under different confidence levels
Figure 14 : The comparison of confidence interval prediction at different confidence levels
Figure 15 : Comparison of prediction interval coverage probability (PICP) for ND, TD, and the proposed HQ method under different nominal confidence levels (80%, 90%, and 95%).
Figure 16 : Comparison of average width deviation (AWD) for ND, TD, and the proposed HQ method under different nominal confidence levels (80%, 90%, and 95%).
Figure 17 : Parity plot (predicted vs. ground truth) with the identity line (y=x); point transparency indicates the predicted confidence (higher opacity denotes higher confidence)
Figure 18 : The relationship of Error and Confidence
Figure 19 : Ablation results with and without the proposed consistency loss.
In/Output
20-5
20-10
20-15
20-20
Without Cons
0.327
1.104
1.831
2.214
With Cons
0.340
1.143
1.840
2.175
Rate(%)
3.97
3.53
0.49
-1.76
Table 6 : Variation of RMSE with and without the consistency loss
Model
RMSE
MAE
p -value
Significant
p -value
Significant
No-standardization
0.0006
Yes
0.0013
Yes
Z-Score
0.0016
Yes
0.0015
Yes
RevIN
0.0413
Yes
0.2529
No
Dish-ts
0.0093
Yes
0.0083
Yes
San
0.0097
Yes
0.0103
Yes
Table 7 : Significance analysis of AdaC against different normalization strategies.
Confidence level
Comparison
PICP
AWD
p -value
Significant
p -value
Significant
80%
ND
3.78×10−21
Yes
1.34×10−21
Yes
TD
1.81×10−32
Yes
2.85×10−31
Yes
90%
ND
1.17×10−17
Yes
3.10×10−17
Yes
TD
1.23×10−30
Yes
2.25×10−28
Yes
95%
ND
2.07×10−14
Yes
4.91×10−14
Yes
Table 8 : Significance analysis of HQ against ND and TD under different confidence levels.
Safe navigation for Unmanned Surface Vehicles (USVs) under the International Regulations for Preventing Collisions at Sea (COLREGs) remains challenging in dynamic maritime environments, especially when perception uncertainty is miscalibrated. Errors in state estimation can produce unreliable belief states that mislead value learning, while logic based on discrete traffic rules can cause abrupt action corrections. To address these challenges, we integrate Credibility-Weighted Value Learning (CWVL) with Covariance- and Recovery-Aware Control Barrier Function Quadratic Programming (CoReCBF-QP). CWVL derives a dynamic trust factor from the discrepancy between the covariance estimated by the filter and empirical error statistics. This factor modulates the critic's heteroscedastic loss and limits overfitting to miscalibrated observations. CoReCBF expands the collision geometry according to uncertainty and incorporates terms for braking and turning recovery. The resulting hyperbolic safety boundary preserves feasible avoidance velocities and supplies the QP safety constraint. A continuous COLREGs-aware reference in the objective promotes starboard maneuvers in Rule 14 head-on and Rule 15 give-way crossing encounters. Simulations show improved robustness in collision avoidance and COLREGs event compliance, achieving an 82.0% success rate with ten target ships beyond the training range.
Yuhang Zhang, Shuqi Chai, Yukang Zhang +5
School of Information Engineering, Henan University of Science and Technology, Luoyang 471023, China · Shenzhen Research Institute of Big Data, Shenzhen 518172, China · Chinese University of Hong Kong, Shenzhen, Shenzhen 518172, China +5
To address the challenges of behavioral multimodality, limited semantic utilization, and long-term error accumulation in vessel trajectory prediction, this paper proposes M3-Former, a multimodal trajectory prediction framework enhanced by large language models (LLMs). The proposed framework incorporates vessel static attributes and navigational intent as semantic priors for long-term trajectory modeling. Specifically, a unified multimodal representation space is constructed, in which static semantic information is encoded by a pre-trained LLM and aligned with dynamic trajectory features through self-attention. To jointly capture global route planning and local motion variations, a dual-granularity Mixture-of-Experts (MoE) architecture is introduced, where sequence-level experts model global navigation trends and token-level experts refine fine-grained maneuvering behaviors. In addition, a Steering-Weighted Cross-Entropy loss is designed to alleviate the long-tail distribution of sparse turning samples and improve prediction accuracy in critical maneuvering scenarios. Experiments on a real-world Danish AIS dataset demonstrate that M\textsuperscript{3}-Former consistently outperforms state-of-the-art baselines across prediction horizons from 1 to 4 hours. In the 4-hour prediction task, the proposed method reduces Average Displacement Error (ADE) and Final Displacement Error (FDE) by 4.4% and 5.1%, respectively, compared with the strongest baseline. Qualitative and ablation analyses further verify that semantic fusion effectively reduces long-term trajectory drift, while the dual-granularity MoE improves robustness in complex waterways and route-branching scenarios. The proposed framework establishes a semantic-guided hierarchical prediction paradigm, in which high-level navigational intent and local motion dynamics are jointly modeled for robust long-term vessel trajectory forecasting.
Wenzhe Jin, Haina Tang
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
Short-horizon prediction is essential for electro-optical UAV tracking, especially when the target is small, maneuvering, or intermittently observed. Image center, line-of-sight, and range measurements provide direct constraints on target position, but their constraints on acceleration are weak. As a result, prediction can lag during aggressive maneuvers. This paper proposes an image-domain tilt constrained distributed fusion method for maneuvering UAV tracking. The method uses the apparent roll and pitch of a rotorcraft target in the image as low-level maneuver cues. A weak-prior auto-labeling pipeline first generates oriented bounding box and image-domain tilt labels from synchronized video, gimbal IMU, and UAV IMU data. A YOLO-OBB detector is then trained to provide online target position and tilt measurements. The front-end Python implementation is publicly available at github.com/ShineMinxing/PythonYOLO. In the fusion stage, the UAV state is modeled by position, velocity, and acceleration. Image-domain roll and pitch are introduced as acceleration-related pseudo-observations. For distributed tracking, one mobile gimbal camera and two fixed ground cameras are fused asynchronously. Camera attitude error states are augmented into the filter to absorb extrinsic drift and cross-camera systematic inconsistency. A Mahalanobis gate with time-since-last-valid covariance widening is used to reject false detections and handle dropouts. In simulation, adding roll/pitch observations reduces the prediction RMSE from 1.991 m to 0.821 m and decreases the cumulative prediction error by 60.75%. In real distributed experiments, a self-consistency evaluation shows an 18.10% reduction in cumulative prediction error. The results show that image-domain tilt can provide useful acceleration constraints for robust short-horizon UAV prediction.
Minxing Sun, Yao Mao
Institute of Optics and Electronics, Chinese Academy of Sciences, Chengdu, China · Institute for Infocomm Research (I2R), Agency for Science, Technology and Research, Singapore · Shenzhen Astralldynamics Technology