In predictive maintenance, early and late RUL prediction errors carry asymmetric consequences, yet hyperparameter optimization typically targets a single accuracy metric that treats both directions equally. This study treats the optimization objective itself as a design variable. Five architectures (MLP, LSTM, XGBoost, TCN, and Transformer) are evaluated under three regimes: single-objective maximization of R2, single-objective minimization of the NASA scoring function, and a multi-objective formulation that jointly optimizes both criteria. The multi-objective search employs NSGA-II with Entropy-CRITIC weighting for Pareto selection. Seventy-five model-dataset-strategy combinations are assessed on the NASA C-MAPSS turbofan and BackBlaze hard-disk drive benchmarks. On C-MAPSS, all strategies achieve comparable accuracy (R2≈0.89), yet multi-objective optimization reduces directional imbalance by approximately 33%, improving calibration of early versus late predictions. Model rankings prove configuration-dependent, with simpler architectures frequently outperforming deeper temporal models. On BackBlaze, the objectives shift from complementary to conflicting, producing divergent Entropy-CRITIC weights and a substantial generalization gap (best R2≈0.34). These results demonstrate that the optimization objective materially shapes prognostic behavior and that multi-objective search provides a practical mechanism for calibrating prediction timeliness in RUL modeling.
Figures & tables
Study
Dataset
Prediction algorithm(s)
Optimization algorithm
Optimization strategy
Evaluation metrics
Haris et al. [ 10 ]
Supercapacitor dataset (simulated)
Deep Belief Network
Bayesian Optimization and Hyperband (BOHB)
Minimize RMSE
RMSE, MAE, R 2
Li & He [ 5 ]
C-MAPSS
DCNN
Bayesian Optimization
Minimize the Deep Learning Error (DLE)
RMSE, Score
Kong et al. [ 11 ]
PCoE
CNN + LSTM
Bayesian Optimization
Minimize RMSE
MAE, MAPE, RMSE
Elsherif et al. [ 48 ]
C-MAPSS
CAELSTM
Tree-structured Parzen Estimator (TPE)
Minimize RMSE
RMSE, MAE, NASA Score
Yang et al. [ 49 ]
NASA battery dataset, CALCE
GBLS Booster
Tree-structured Parzen Estimator (TPE)
Minimize MAPE
MAPE, RMSE, R 2 , Relative Error (RE)
Kara [ 50 ]
NASA battery dataset
CNN-LSTM
PSO
Minimize RMSE
MAE, MAPE, RMSE, AE
Table 1: Key related research.
FD001
FD002
FD003
FD004
BackBlaze
Train / Val / Test units
85 / 15 / 100
221 / 39 / 259
85 / 15 / 100
212 / 37 / 248
1,956 / — / 488
Train observations
20,631
53,759
24,720
61,249
111,971
Test observations
13,096
33,991
16,596
41,214
28,308
Mean cycles/unit (train)
206.3
206.8
247.2
246.0
57.6
Mean train RUL
86.8
86.9
93.1
93.0
30.0
% at RUL cap (train)
39.4%
39.5%
49.4%
49.2%
10.9%
Table 2: Dataset summary.
Figure 1: The methodology of the study.
Dataset
Strategy
R 2
RMSE
NASA
α -20%
DI
C-MAPSS FD001
Single-R 2
0.890
13.19
255.8
70.0%
6.20
Single-NASA
0.889
13.32
248.2
73.6%
4.20
Multi-objective
0.878
13.87
305.8
66.2%
4.80
C-MAPSS FD002
Single-R 2
0.903
13.34
795.6
65.6%
9.15
Single-NASA
0.889
14.28
960.4
60.8%
13.01
Multi-objective
0.889
14.21
928.1
65.3%
6.76
Table 3: Comparative analysis across the HPO strategies (averaged over the five algorithms).
Dataset
Algorithm
R 2
RMSE
NASA
α -20%
DI
FD001
XGBoost
0.916
11.57
190.3
75.0%
1.33
MLP
0.899
12.73
249.1
72.0%
5.00
LSTM
0.891
13.23
250.8
72.0%
5.33
Transformer
0.868
14.54
348.7
69.7%
7.67
TCN
0.855
15.24
310.9
61.0%
6.00
FD002
XGBoost
0.911
12.80
649.4
68.6%
4.70
Table 4: Comparative analysis across algorithms (averaged over the three optimization strategies).
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Algorithm
Hyperparameter
Search range
Type
MLP
Sequence length
{20, 30, 40, 50}
Categorical
Hidden layers
[2, 5]
Integer
First layer size
{128, 256, 512}
Categorical
Layer shrink factor
{1.0, 0.75, 0.5}
Categorical
Dropout rate
[0.1, 0.5], step 0.1
Float
L2 regularization
[ 10−5 , 10−3 ]
Log-float
Appendix
Table 5: Hyperparameter search spaces for each algorithm.
Remaining Useful Life (RUL) estimation is a critical component of Prognostics and Health Management (PHM), enabling proactive maintenance scheduling and reducing unplanned failures in industrial equipment. This paper presents a comparative study of machine learning approaches for RUL estimation on the NASA C-MAPSS turbofan engine dataset: classical baselines (Ridge Regression, Polynomial Ridge, and XGBoost), a 1D Convolutional Neural Network (CNN), and a Long Short-Term Memory (LSTM) network. All models are evaluated on the FD001 and FD003 subsets under an identical preprocessing pipeline to ensure a fair comparison. Among raw-sequence models, the LSTM achieves RMSE of 14.93 and 14.20 on FD001 and FD003 respectively, outperforming the deep LSTM reported by Zheng et al.~\cite{paper} (RMSE 16.14 and 16.18) despite using a simpler single-layer architecture. The 1D CNN achieves RMSE of 16.97 on FD001 and 15.68 on FD003, demonstrating competitive performance on FD003 while producing more conservative RUL predictions on FD001. Ridge Regression is evaluated on raw and engineered features, while other classical models use only engineered inputs. XGBoost achieves an RMSE of 13.36 on FD003, highlighting the competitiveness of nonlinear modeling.
Astitva Goel, Samarth Galchar, Sumit Kanu
Khoury College of Computer Science, Northeastern University
This study presents a novel hybrid prognostic framework for uncertainty-aware Remaining Useful Life (RUL) estimation in turbofan engines using the NASA C-MAPSS dataset. The framework employs a state-aware strategy that bifurcates the engines operational lifespan into "healthy" and "degraded" regimes. An LSTM-based autoencoder, trained strictly on nominal data (RUL > 150 cycles), monitors reconstruction error to act as a robust state classifier. For the healthy regime, a Conditional Weibull Survival Analysis is used for Mean Residual Life estimation. For the degraded regime, a Probabilistic Neural Network with Monte Carlo Dropout captures both aleatoric and epistemic uncertainties. Rather than using rigid binary labels, a calibrated sigmoid function converts the autoencoders output into continuous state probabilities, dynamically weighting the final ensemble prediction. The primary strength of this framework is its generation of physically consistent uncertainty bands, yielding high-confidence predictions near end-of-life while accurately reflecting the inherent variance of early operation, providing a robust tool for risk-informed maintenance.
Xabier Belaunzaran, Antonio Nappa, Arkaitz Artetxe +1
Fundación Vicomtech · Basque Research and Technology Alliance (BRTA) · Mikeletegi 57, Donostia-San Sebastian, 20009, Spain +2
Accurate prediction of Remaining Useful Life (RUL) in aero-engines is vital for predictive maintenance, improved operational reliability, and reduced lifecycle costs. While deep learning approaches have demonstrated strong potential in this area, most existing methods focus primarily on model architecture design and treat input features uniformly, often neglecting the influence of data preprocessing. In this work, we propose a novel preprocessing pipeline that enhances RUL prediction by improving data quality and temporal representation before model training. Our approach leverages complete temporal sequences and generates RUL estimates at each timestep, enabling the model to capture fine-grained degradation dynamics and deliver continuous prognostic insights throughout the engine's operational life. To validate the effectiveness of the proposed pipeline, we conduct experiments on the NASA C-MAPSS dataset. Comparative evaluations against a suite of state-of-the-art neural models including CNN, RNN, LSTM, DCNN, TCN, BiGRU-TSAM, AGCNN, and ATCN, demonstrate that our approach consistently achieves superior accuracy and robustness in aero-engine RUL prediction. These results highlight the critical role of preprocessing in maximizing the effectiveness of neural prognostic models.
Florent Imbert, Tosin Adewumi, Hui Han
Machine Learning Group · Lulea University of Technology · Lulea, Sweden