In predictive maintenance, early and late RUL prediction errors carry asymmetric consequences, yet hyperparameter optimization typically targets a single accuracy metric that treats both directions equally. This study treats the optimization objective itself as a design variable. Five architectures (MLP, LSTM, XGBoost, TCN, and Transformer) are evaluated under three regimes: single-objective maximization of R2, single-objective minimization of the NASA scoring function, and a multi-objective formulation that jointly optimizes both criteria. The multi-objective search employs NSGA-II with Entropy-CRITIC weighting for Pareto selection. Seventy-five model-dataset-strategy combinations are assessed on the NASA C-MAPSS turbofan and BackBlaze hard-disk drive benchmarks. On C-MAPSS, all strategies achieve comparable accuracy (R2≈0.89), yet multi-objective optimization reduces directional imbalance by approximately 33%, improving calibration of early versus late predictions. Model rankings prove configuration-dependent, with simpler architectures frequently outperforming deeper temporal models. On BackBlaze, the objectives shift from complementary to conflicting, producing divergent Entropy-CRITIC weights and a substantial generalization gap (best R2≈0.34). These results demonstrate that the optimization objective materially shapes prognostic behavior and that multi-objective search provides a practical mechanism for calibrating prediction timeliness in RUL modeling.
Figures & tables
Study
Dataset
Prediction algorithm(s)
Optimization algorithm
Optimization strategy
Evaluation metrics
Haris et al. [ 10 ]
Supercapacitor dataset (simulated)
Deep Belief Network
Bayesian Optimization and Hyperband (BOHB)
Minimize RMSE
RMSE, MAE, R 2
Li & He [ 5 ]
C-MAPSS
DCNN
Bayesian Optimization
Minimize the Deep Learning Error (DLE)
RMSE, Score
Kong et al. [ 11 ]
PCoE
CNN + LSTM
Bayesian Optimization
Minimize RMSE
MAE, MAPE, RMSE
Elsherif et al. [ 48 ]
C-MAPSS
CAELSTM
Tree-structured Parzen Estimator (TPE)
Minimize RMSE
RMSE, MAE, NASA Score
Yang et al. [ 49 ]
NASA battery dataset, CALCE
GBLS Booster
Tree-structured Parzen Estimator (TPE)
Minimize MAPE
MAPE, RMSE, R 2 , Relative Error (RE)
Kara [ 50 ]
NASA battery dataset
CNN-LSTM
PSO
Minimize RMSE
MAE, MAPE, RMSE, AE
Table 1: Key related research.
FD001
FD002
FD003
FD004
BackBlaze
Train / Val / Test units
85 / 15 / 100
221 / 39 / 259
85 / 15 / 100
212 / 37 / 248
1,956 / — / 488
Train observations
20,631
53,759
24,720
61,249
111,971
Test observations
13,096
33,991
16,596
41,214
28,308
Mean cycles/unit (train)
206.3
206.8
247.2
246.0
57.6
Mean train RUL
86.8
86.9
93.1
93.0
30.0
% at RUL cap (train)
39.4%
39.5%
49.4%
49.2%
10.9%
Table 2: Dataset summary.
Figure 1: The methodology of the study.
Dataset
Strategy
R 2
RMSE
NASA
α -20%
DI
C-MAPSS FD001
Single-R 2
0.890
13.19
255.8
70.0%
6.20
Single-NASA
0.889
13.32
248.2
73.6%
4.20
Multi-objective
0.878
13.87
305.8
66.2%
4.80
C-MAPSS FD002
Single-R 2
0.903
13.34
795.6
65.6%
9.15
Single-NASA
0.889
14.28
960.4
60.8%
13.01
Multi-objective
0.889
14.21
928.1
65.3%
6.76
Table 3: Comparative analysis across the HPO strategies (averaged over the five algorithms).
Dataset
Algorithm
R 2
RMSE
NASA
α -20%
DI
FD001
XGBoost
0.916
11.57
190.3
75.0%
1.33
MLP
0.899
12.73
249.1
72.0%
5.00
LSTM
0.891
13.23
250.8
72.0%
5.33
Transformer
0.868
14.54
348.7
69.7%
7.67
TCN
0.855
15.24
310.9
61.0%
6.00
FD002
XGBoost
0.911
12.80
649.4
68.6%
4.70
Table 4: Comparative analysis across algorithms (averaged over the three optimization strategies).
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Algorithm
Hyperparameter
Search range
Type
MLP
Sequence length
{20, 30, 40, 50}
Categorical
Hidden layers
[2, 5]
Integer
First layer size
{128, 256, 512}
Categorical
Layer shrink factor
{1.0, 0.75, 0.5}
Categorical
Dropout rate
[0.1, 0.5], step 0.1
Float
L2 regularization
[ 10−5 , 10−3 ]
Log-float
Appendix
Table 5: Hyperparameter search spaces for each algorithm.