Stream-Based Active Learning with Cooperative Neural Networks for Data-Efficient Partial Inverse Design: An Automotive Glass Run Channel Case Study
Authors: Agung Nugraha, Hyerin Kwon, Heungjun Im, Gian Antariksa, Jihwan Lee
Organizations: Department of Industrial and Data Engineering, Major in Industrial Data Science and Engineering, Pukyong National University, Busan, 48513, Republic of Korea · DRB Co., Ltd., Busan, 46329, Republic of Korea
Inverse design in engineering often runs into a simple problem. Each labeled training sample must be produced through expensive simulation, so building a large dataset is slow and costly. This study addresses that problem for partial inverse design, where only some design variables are specified and the rest must be inferred to reach a target performance value. We propose CoNN-AL, a framework for data-efficient partial inverse design that adds stream-based active learning to the Cooperative Neural Network with Denoising Autoencoder (CoNN-DAE). The model estimates predictive uncertainty through Monte Carlo dropout and uses it to decide, in real time, which incoming candidate samples are worth labeling, so the limited labeling budget is spent on the most informative designs. We validate the framework on a real-world automotive glass run channel dataset of more than 900,000 unique simulated designs. With only 20,000 actively selected labels, about 2.3% of the training pool, CoNN-AL reaches R-squared values of 0.967 to 0.982 across all missing-variable levels, approaching the upper-bound models trained on far more data. It reaches R-squared of at least 0.95 with 30 to 40% fewer labels than random sampling at the more difficult missing-variable levels and, at the most challenging level, is the only strategy in this study to reach R-squared of 0.98. Together with this work, we publicly release the dataset to support future research on data-driven design.
Figures & tables
Figure 1: Architecture of the Cooperative Neural Network with Denoising Autoencoder (CoNN-DAE), adapted from Nugraha et al. [ 9 ] . The surrogate model (top) maps complete designs to the performance output. The imputation model (bottom) reconstructs missing variables via an autoencoder, jointly optimized through reconstruction and performance losses.
Figure 2: Stream-based active learning loop integrated with CoNN-DAE. Unlabeled samples stream in continuously, and the model estimates predictive uncertainty via MC-Dropout. High-uncertainty samples are queried from the oracle (FEA simulation) and added to the labeled pool, while low-uncertainty samples are discarded.
Figure 3: (a) Illustration of the GRC sealing profile location on the vehicle window frame. (b) The GRC profile during finite element analysis in MSC Marc.
Group
Variable
Unit
Lip Related Data
Seal gap
mm
Lip tip radius
mm
Lip radius at glass flat
mm
Lip radius at top flat
mm
Notch Related Data
Notch lip thickness
mm
Notch radius
mm
Table 1: Groups and units of the GRC design variables and output. Units refer to the original measurements; the shared dataset provides the design variables in standardized form.
Metric
Definition
Coefficient of Determination ( R2 )
1−∑i=1m(yi−yˉ)2∑i=1m(yi−y^i)2
Mean Squared Error (MSE)
m1∑i=1m(yi−y^i)2
Mean Absolute Error (MAE)
m1∑i=1m∣yi−y^i∣
Root Mean Squared Error (RMSE)
m1∑i=1m(yi−y^i)2
Symmetric Mean Absolute Percentage Error (sMAPE)
m2i=1∑m∣yi∣+∣y^i∣∣yi−y^i∣
Table 2: Evaluation metrics and their definitions.
Figure 4: Radar chart comparison of upper-bound performance metrics (MSE, MAE, sMAPE, R2 ) between 100,000 and 886,215 training sizes across MM levels. Performance is largely saturated at 100,000 samples.
Figure 5: Ground truth vs. prediction scatter plots for CoNN-DAE trained on 100,000 samples (MM = 1 to MM = 6). Points align closely with the identity line across all MM levels.
Figure 6: Ground truth vs. prediction scatter plots for CoNN-DAE trained on 886,215 samples (MM = 1 to MM = 6). The pattern closely matches the 100,000 model, confirming performance saturation.
Figure 7: Learning curves for random sampling baseline across all MM levels. Performance metrics (MSE, MAE, RMSE, sMAPE, R2 ) are shown as a function of training size (10,000–100,000). Insets report each metric at the 40,000-sample checkpoint (dashed line). The bottom-right panel shows the number of training samples required to reach R2≥0.95 at each MM level. A performance plateau is reached around 40,000 samples.
Figure 8: Active learning convergence curves (MSE, MAE, RMSE, sMAPE, R2 ) over 100 rounds, with insets reporting each metric at round 18 (20,000 labeled samples), marked by the dashed line. The bottom-right panel shows the number of labeled samples, and the corresponding round, required to reach R2≥0.95 at each MM level. Performance is consistent across all MM levels at this early checkpoint.
Figure 9: Ground truth vs. prediction scatter plots at active learning round 18 (20,000 labels) for MM = 1 to MM = 6. Near upper-bound prediction accuracy is achieved with approximately 2.3% of the full labeling budget.
Figure 10: R2 comparison at three active learning checkpoints (round 18, round 74, round 100) across MM levels. Round 74 (76,000 labeled samples) represents the optimal stopping point.
Figure 11: Performance comparison of active learning, random sampling, and the upper-bound models (100,000 and 886,215) at 20,000 labeled samples (MM = 6). Active learning outperforms random sampling and approaches the upper-bound.
Figure 12: Performance comparison at the most efficient active learning checkpoint (76,000 labels) vs. random sampling (60,000) and upper-bound (100,000, 886,215) under MM = 6. Active learning achieves lower MSE and slightly higher R2 than the upper-bound models, while MAE is higher.
R2 Threshold
Random Sampling
Active Learning
Efficiency Ratio
0.90
10,000
8,000
1.3 ×
0.92
10,000
10,000
1.0 ×
0.94
20,000
12,000
1.7 ×
0.95
20,000
14,000
1.4 ×
0.96
30,000
15,000
2.0 ×
0.97
40,000
20,000
2.0 ×
Table 3: Label efficiency comparison at R2 thresholds (MM = 6), showing the number of labeled samples required by random sampling and active learning to reach each threshold.
Universidade Federal de Santa Catarina, Florianópolis, SC, Brazil · Universidade Católica de Pelotas, Pelotas, RS, Brazil · Université de Lorraine, CNRS, CRAN, Vandoeuvre-lès-Nancy, France