Predictive Inorganic Synthesis based on Machine Learning using Small Data sets: a case study of Hydrodynamic Diameter-controlled Cu Nanoparticles
Authors: Brent Motmans, Digvijay Ghogare, Thijs G. I. van Wijk, Joren Van Herck, Saba Heidarian, Pieter De Meyer, Berend Smit, An Hardy, +1 more
Organizations: Hasselt University, Institute for Materials Research (IUMAT), Quantum & Artificial inTelligence design Of Materials (QuATOMs), Martelarenlaan 42, B-3500 Hasselt, Belgium · Hasselt University, Institute for Materials Research (IUMAT), Hybrid Materials Design (HyMaD), Martelarenlaan 42, B-3500 Hasselt, Belgium · imec, IUMAT, Wetenschapspark 1, B-3590 Diepenbeek, Belgium · Energyville, IUMAT, Thor Park 8320, B-3600 Genk, Belgium · Hasselt University, Institute for Materials Research (IUMAT), Design and Synthesis of Inorganic Materials (DESINe), Martelarenlaan 42, B-3500 Hasselt, Belgium · Electrochemistry Excellence Centre (ELEC), Materials & Chemistry Unit, Flemish Institute for Technological Research (VITO), Boeretang 200, Mol, 2400 Belgium · Laboratory of Molecular Simulation (LSMO), Institut des Sciences et Ing´enierie Chimiques, ´Ecole Polytechnique F´ed´erale de Lausanne (EPFL), Rue de l’Industrie 17, CH-1951 Sion, Switzerland
Cu NPs have a broad applicability, yet their synthesis is sensitive to subtle changes in reaction parameters. This sensitivity, combined with the time- and resource-intensive nature of experimental optimization, poses a major challenge in achieving reproducible and size-controlled synthesis. While ML shows promise in materials research, its application is often limited by scarcity of large high-quality experimental data sets. This study explores ML to predict the DLS-derived hydrodynamic diameter of Cu NPs using a small data set of 25 syntheses. Latin Hypercube Sampling is used to efficiently cover the parameter space while creating the experimental data set. Ensemble regression models successfully predict hydrodynamic diameters with good predictive performance given the limited dataset. Since quantitative regression requires a unique DLS-derived hydrodynamic diameter, the regression model is restricted to mono-modal DLS distributions, while a complementary classification model identifies synthesis conditions for which quantitative prediction is applicable. Using equivalent out-of-sample validation, the ML and DoE models showed comparable generalization. The final ensemble model achieved an R2=0.74 compared to 0.60 for the DoE model, while retaining the complete synthesis parameter space, making it better suited for synthesis guidance. Additionally, classification models using both random forests and LLMs are evaluated to distinguish between large and small particles. These classification models exhibited only modest predictive performance, indicating that this small dataset is insufficient to fully exploit the capabilities of complex LLMs. Overall, this study demonstrates that carefully curated small data sets, paired with robust classical ML, can effectively support the synthesis of Cu NPs and highlights that for lab-scale studies, complex models like LLMs may offer limited benefits.
Figures & tables
Figure 1: Visualization of the workflow followed to create the data set and model the syntheses.
Figure 2: The computational modeling starts with selecting the best performing data set. Next, the optimal value for the regularization strength α is selected. Then, 100 model instances are trained on random subsets of the data set. From these an ensemble model is created through averaging over all instances. Finally, through an iterative application of feature engineering and selection next generation models are created.
Polynomial degree
RMSE [nm]
MAE [nm]
R2
Data set A
3
63.46
50.50
0.42
Data set B
3
59.44
46.06
0.35
Data set C
3
34.09
24.28
0.73
Table 1: Performance metrics of the best-performing ensemble model for the different data sets.
Training
OOB
Final ensemble model
# Features
RMSE [nm]
MAE [nm]
RMSE [nm]
MAE [nm]
RMSE [nm]
MAE [nm]
R2
1st generation
19
22.17
16.05
86.52
68.51
32.64
22.53
0.75
2nd generation
7
29.40
20.47
51.85
41.36
32.61
22.28
0.75
3rd generation
11
27.78
19.79
63.66
52.18
33.07
23.35
0.74
4th generation
6
30.57
22.39
49.90
40.92
33.03
23.81
0.74
DoE
2
−−
−−
48.22a
40.46a
40.91
33.54
0.60
Table 2: Performance metrics of different regression model generations.
Figure 3: a) Plot of the experimentally measured hydrodynamic diameter and their standard deviation in ascending order (black) with the corresponding predicted hydrodynamic diameter of the fourth generation ensemble model (red); b) Parity plot of the actual, experimental hydrodynamic diameter against the corresponding predicted hydrodynamic diameter using the fourth generation ensemble model.
Figure 4: a) Plot of the experimental measured hydrodynamic diameters and their standard deviation (black) and the statistical model (red); b) Parity plot of the actual, experimental hydrodynamic diameter against the corresponding predicted hydrodynamic diameter using the statistical model.
Sample
Targeted [nm]
Experimental [nm]
26
115
70.78\penalty(±4.68 )
27
115
117.90\penalty(±32.06 )
28
220
235.60\penalty(±34.57 )
29
220
128.60\penalty(±49.54 )
30
275
165.10\penalty(±41.04 )
31
275
189.60\penalty(±59.50 )
Table 3: Comparison of targeted hydrodynamic diameters, experimentally obtained hydrodynamic diameters and their σ from predicted syntheses routes with the fourth generation ensemble model.
Accuracy
κ
GPT-J
0.48
−0.04
LLAMA 3.1
0.64
0.28
Random forest
0.56
0.11
Table 4: Performance metrics of binary classification models.
Figure 5: a) Confusion matrix of the SVM classifier for mono- and multi-modal DLS distributions; b) Visualization of the SVM ensemble predictions in the synthesis parameter space. The color scale represents the predicted probability of a multi-modal DLS distribution; c) 3D uncertainty map showing how uncertainty is concentrated around the models decision boundary, while large portions of the parameter space show almost zero uncertainty.
Sample
Theoretical Cu conc. a [mM]
Effective Cu prec. b [mM]
TMAH c [ μ l]
Timed [min]
Temp. e [ ∘ C]
1
5.00
5.06
36.0
8
196
2
4.00
4.21
29.0
4
188
3
3.00
2.91
21.8
6
198
4
2.00
2.00
14.6
7
184
5
9.00
8.82
65.4
4
180
6
3.00
3.11
21.8
1
194
Table 1: Sampled synthesis parameters used for the 25 experimental syntheses. (a) Theoretical concentration of Cu(OAc) 2 precursor in 5 ml ethylene glycol in mM; (b) Effective Cu concentration of Cu(OAc) 2 used; (c) Volume of 2.75 M TMAH in H 2 O added to 5 ml ethylene glycol in μ l; (d) Time in minutes reaction mixture is kept on defined temperature; (e) Temperature in ∘ C of reaction mixture for defined reaction time.
Sample
Effective Cu conc. a [mM]
TMAH b [ μ l]
Timec [min]
Temp. d [ ∘ C]
Size e [nm]
1
5.06
36.0
8
196
252.50 ( ± 59.37)
2
4.21
29.0
4
188
228.40 ( ± 41.47)
3
2.91
21.8
6
198
148.80 ( ± 25.56)
4
2.00
14.6
7
184
99.90 ( ± 14.07)
5
8.82
65.4
4
180
183.30 ( ± 25.42)
6
3.11
21.8
1
194
160.70 ( ± 27.62)
Table 2: Data set A containing 25 datapoints with particle size, the smallest size detected by DLS. (a) Effective concentration of Cu(OAc) 2 precursor in 5 ml ethylene glycol in mM; (b) Volume of 2.75 M TMAH in H 2 O added to 5 ml ethylene glycol in μ l; (c) Time in minutes reaction mixture is kept on defined temperature; (d) Temperature in ∘ C of reaction mixture for defined reaction time. (e) Hydrodynamic diameter of the NPs and the standard deviation ( σ ) in nm as measured using DLS.
Sample
Effective Cu conc. a [mM]
TMAH b [ μ l]
Timec [min]
Temp. d [ ∘ C]
Size e [nm]
1
5.06
36.0
8
196
252.50 ( ± 59.37)
2
4.21
29.0
4
188
228.40 ( ± 41.47)
3
2.91
21.8
6
198
148.80 ( ± 25.56)
4
2.00
14.6
7
184
99.90 ( ± 14.07)
5
8.82
65.4
4
180
183.30 ( ± 25.42)
6
3.11
21.8
1
194
160.70 ( ± 27.62)
Table 3: Data set B containing 25 datapoints with particle size, the size with the highest percentage contribution in DLS. (a) Effective Cu concentration of Cu(OAc) 2 used in 5 ml ethylene glycol; (b) Volume of 2.75 M TMAH in H 2 O added to 5 ml ethylene glycol in μ l; (c) Time in minutes reaction mixture is kept on defined temperature; (d) Temperature in ∘ C of reaction mixture for defined reaction time; (e) Hydrodynamic diameter of the NPs and the σ in nm as measured using DLS.
Sample
Effective Cu conc. a [mM]
TMAH b [ μ l]
Timec [min]
Temp. d [ ∘ C]
Size e [nm]
1
5.06
36.0
8
196
252.50 ( ± 59.37)
2
4.21
29.0
4
188
228.40 ( ± 41.47)
3
2.91
21.8
6
198
148.80 ( ± 25.56)
4
2.00
14.6
7
184
99.90 ( ± 14.07)
5
8.82
65.4
4
180
183.30 ( ± 25.42)
6
3.11
21.8
1
194
160.70 ( ± 27.62)
Table 4: Data set C containing 18 datapoints with monomodal DLS particle-size distribution. (a) Effective Cu concentration of Cu(OAc) 2 used in 5 ml ethylene glycol in mM; (b) Volume of 2.75 M TMAH in H 2 O added to 5 ml ethylene glycol in μ l; (c) Time in minutes reaction mixture is kept on defined temperature; (d) Temperature in ∘ C of reaction mixture for defined reaction time; (e) Hydrodynamic diameter of the NPs and the σ in nm as measured using DLS.
Figure 1: Hyperparameter optimization for size prediction with data set C. LASSO regularized polynomials up to order eight are compared, and MAE of the final ensemble models is shown for α =1.0 (red), α =0.1 (orange), α =0.01 (green), and α =0.001 (blue).
Figure 2: Overview of feature importance, sorted by increasing importance in the third degree polynomial, in the polynomial models with degree two (red) and degree three (black), and trained on data set C with α =0.1. x_0 = Cu precursor concentration, x_1 = Reaction temperature and x_2 = Reaction time.
Feature
Average EI [%]
Minimum EI [%]
Maximum EI [%]
[Cu]
84.5
10.0
100
T
99.6
97.0
100
time
82.2
54.0
100
[Cu]2
82.3
53.0
100
[Cu]×T
80.0
43.0
100
[Cu]×time
99.1
95.0
100
Table 6: Summary of the dominance importance analysis performed on the third generation ensemble model.
Feature
xˉ
σ
T
1.88×102
8.07×100
[Cu]×time
4.42×10−2
3.72×10−2
[Cu]3
1.75×10−7
2.19×10−7
e[Cu]
1.00×100
2.32×10−3
time[Cu]
8.90×10−4
9.27×10−4
ln([Cu]×time)
−3.49×100
9.20×101
Table 7: Mean and standard deviation ( σ ) used to standardize the different features in the 4th generation ensemble model.
T
[Cu]×time
[Cu]3
e[Cu]
time[Cu]
ln([Cu]×time)
MAE train
MAE OOB
RMSE train
RMSE OOB
1
1
1
1
0
1
24.45
37.69
32.47
45.57
1
0
1
1
0
1
26.21
36.86
34.73
44.36
1
0
1
1
0
0
26.66
34.43
35.33
42.58
Table 8: Summary of the three best performing models resulting from the dominance importance analysis performed on the fourth generation ensemble model. 1 indicating the feature is included in the model, and 0 indicating the feature is not included.
T
[Cu]×time
[Cu]3
e[Cu]
time[Cu]
ln([Cu]×time)
MAE train
MAE OOB
RMSE train
RMSE OOB
1
0
1
0
1
1
33.22
50.70
40.70
59.00
0
0
1
0
1
1
32.93
48.63
43.14
59.22
0
1
1
0
1
1
28.53
49.08
37.71
60.06
1
1
1
0
1
1
29.56
52.38
36.88
61.11
0
1
0
0
1
1
42.05
59.45
50.77
71.62
1
0
0
0
1
1
46.58
64.07
55.77
72.66
Table 9: Because the engineered e[Cu] feature exhibits only limited variation over the investigated concentration range, dominance importance analysis is performed to assess whether its removal affects the predictive performance of the ensemble model. 1 indicating the feature is included in the model, and 0 indicating the feature is not included. None of the models excluding e[Cu] outperformed the original fourth generation ensemble model.
Figure 3: a) DLS results showing particle hydrodynamic diameter and standard deviations for all detected peaks per sample. Each point represents a population of a distinct size (Peak 1 in blue, Peak 2 in green, and Peak 3 in orange), with error bars indicating the corresponding standard deviation. Samples are sorted in ascending order based on the size of Peak 1. Peaks above 5000 nm were excluded due to their negligible intensity ( < 1.5%); b) Distribution of the smallest detected particle size per sample as measured by DLS. The red line indicates the median particle hydrodynamic diameter.
Figure 4: Selected area electron diffraction pattern of Cu0 nanoparticles from sample 22, recorded at an acceleration voltage of 120 kV. The presence of concentric diffraction rings indicates a polycrystalline material. The rings are indexed to the face-centred cubic (FCC) crystal lattice of metallic copper, with corresponding (111), (200) and (220) planes. The measured d-spacings for these planes are 2.08 Å , 1.83 Å , and 1.29 Å , respectively, in good agreement with standard reference values.
Figure 5: Particle-hydrodynamic diameter distribution resulting from DLS of experimental synthesis a)Sample_1; b)Sample_2; c)Sample_3; d)Sample_4; e)Sample_5; f)Sample_6; g)Sample_7; h)Sample_8; i)Sample_9; j)Sample_10.
Figure 6: Particle-hydrodynamic diameter distribution resulting from DLS of experimental synthesis a)Sample_11; b)Sample_12; c)Sample_13; d)Sample_14; e)Sample_15; f)Sample_16; g)Sample_17; h)Sample_18; i)Sample_19; j)Sample_20.
Figure 7: Particle-hydrodynamic diameter distribution resulting from DLS of experimental synthesis a)Sample_21; b)Sample_22; c)Sample_23; d)Sample_24; e)Sample_25.
Figure 8: Hyperparameter tuning of the SVM classifier. Heatmap showing the mean cross-validated accuracy as a function of the regularization parameter C and kernel parameter γ . a) Shows the full parameter grid, where the row corresponding to C =0.01 yields nearly constant accuracy values across all γ , indicating degenerate model behavior. b) Shows a slightly adjusted parameter grid with C =0.011 , demonstrating that this effect is not intrinsic to the model but arises from algorithmic instability associated with the limited dataset size. This highlights the importance of careful interpretation of hyperparameter optimization results in small-data regimes. Notably, the selected optimal hyperparameters (C =1.0 , γ=1.0 ) lie outside this regime, and therefore this behavior does not affect the final model selection.
Figure 9: Confusion matrix for an ensemble SVM model with hyperparameters C =0.1 and γ=0.01 yielding an apparent LOOCV accuracy of 1.0 . The ensemble model predicts only the mono-modal class, indicating a collapse to a single-class prediction despite the high LOOCV accuracy.
Figure 10: Learning curve analysis of predictions for the binary class of the size of nanoparticles. Models fine-tuned with 10 (blue), 25 (green), 50 (orange) and 100 (red) epochs were validated. GPT-J was used as the base model.
Figure 13: a) Plot of the experimentally measured hydrodynamic diameter and their standard deviation in ascending order (black) with the corresponding predicted particle size of the first generation ensemble model (red); b) Parity plot of the actual, experimental size against the corresponding predicted hydrodynamic diameter using the first generation ensemble model.
Figure 14: a) Plot of the experimentally measured hydrodynamic diameter and their standard deviation in ascending order (black) with the corresponding predicted particle size of the second generation ensemble model (red); b) Parity plot of the actual, experimental hydrodynamic diameter against the corresponding predicted size using the second generation ensemble model.
Figure 15: a) Plot of the experimentally measured hydrodynamic diameter and their standard deviation in ascending order (black) with the corresponding predicted particle size of the third generation ensemble model (red); b) Parity plot of the actual, experimental hydrodynamic diameter against the corresponding predicted size using the third generation ensemble model.
Figure 16: Comparison of the DLS-derived hydrodynamic diameters obtained from the original syntheses (black) and independent repeat syntheses (red) for samples 7, 20, and 24. Error bars indicate the standard deviation of the DLS measurements.
Sample
Cu conc. a [mM]
TMAH b [ μ l]
Timec [min]
Temp. d [ ∘ C]
Targeted size e [nm]
Obtained size f [nm]
26
8.82
64.0
18
175
115
70.78 ( ± 4.68)
27
2.30
16.8
4
195
115
117.90 ( ± 32.06)
28
7.51
54.6
16
175
220
235.60 ( ± 34.57)
29
3.91
27.8
7
189
220
128.60 ( ± 49.45)
30
5.51
40.0
15
175
275
165.10 ( ± 41.04)
31
6.01
43.6
4
179
275
189.60 ( ± 59.50)
Table 10: Synthesis suggestions made by fourth generation model. (a) Effective concentration of Cu(OAc) 2 precursor in 5 ml ethylene glycol in mM; (b) Volume of 2.75 M TMAH in H 2 O added to 5 ml ethylene glycol in μ l; (c) Time in minutes reaction mixture is kept on defined temperature; (d) Temperature in ∘ C of reaction mixture for defined reaction time; (e) Targeted hydrodynamic diameter in nm of the Cu NPs measured using DLS in ethanol; (f) Obtained hydrodynamic diameter in nm of the Cu NPs and the σ measured using DLS in ethanol.
Targeted size a [nm]
Cu conc. b [mM]
115
2.01
115
10.10
220
3.94
220
8.16
275
No real solution
275
No real solution
Table 11: Synthesis suggestions made by DoE model. (a) Targeted hydrodynamic diameter in nm of the Cu NPs measured using DLS in ethanol; (b) Effective concentration of Cu(OAc) 2 precursor in 5 ml ethylene glycol in mM.
Department of Chemistry, Washington University, St. Louis, MO 63130, USA. · Department of Chemistry, Fudan University, Shanghai 200438, China. · Department of Computer and Information Science, University of Pennsylvania, PA 19104, USA. +2