cond-mat.mtrl-sciDec 18, 2025

Predictive Inorganic Synthesis based on Machine Learning using Small Data sets: a case study of Hydrodynamic Diameter-controlled Cu Nanoparticles

Authors: Brent Motmans, Digvijay Ghogare, Thijs G. I. van Wijk, Joren Van Herck, Saba Heidarian, Pieter De Meyer, Berend Smit, An Hardy, +1 more

Organizations: Hasselt University, Institute for Materials Research (IUMAT), Quantum & Artificial inTelligence design Of Materials (QuATOMs), Martelarenlaan 42, B-3500 Hasselt, Belgium · Hasselt University, Institute for Materials Research (IUMAT), Hybrid Materials Design (HyMaD), Martelarenlaan 42, B-3500 Hasselt, Belgium · imec, IUMAT, Wetenschapspark 1, B-3590 Diepenbeek, Belgium · Energyville, IUMAT, Thor Park 8320, B-3600 Genk, Belgium · Hasselt University, Institute for Materials Research (IUMAT), Design and Synthesis of Inorganic Materials (DESINe), Martelarenlaan 42, B-3500 Hasselt, Belgium · Electrochemistry Excellence Centre (ELEC), Materials & Chemistry Unit, Flemish Institute for Technological Research (VITO), Boeretang 200, Mol, 2400 Belgium · Laboratory of Molecular Simulation (LSMO), Institut des Sciences et Ing´enierie Chimiques, ´Ecole Polytechnique F´ed´erale de Lausanne (EPFL), Rue de l’Industrie 17, CH-1951 Sion, Switzerland

Abstract

Cu NPs have a broad applicability, yet their synthesis is sensitive to subtle changes in reaction parameters. This sensitivity, combined with the time- and resource-intensive nature of experimental optimization, poses a major challenge in achieving reproducible and size-controlled synthesis. While ML shows promise in materials research, its application is often limited by scarcity of large high-quality experimental data sets. This study explores ML to predict the DLS-derived hydrodynamic diameter of Cu NPs using a small data set of 25 syntheses. Latin Hypercube Sampling is used to efficiently cover the parameter space while creating the experimental data set. Ensemble regression models successfully predict hydrodynamic diameters with good predictive performance given the limited dataset. Since quantitative regression requires a unique DLS-derived hydrodynamic diameter, the regression model is restricted to mono-modal DLS distributions, while a complementary classification model identifies synthesis conditions for which quantitative prediction is applicable. Using equivalent out-of-sample validation, the ML and DoE models showed comparable generalization. The final ensemble model achieved an R2=0.74 compared to 0.60 for the DoE model, while retaining the complete synthesis parameter space, making it better suited for synthesis guidance. Additionally, classification models using both random forests and LLMs are evaluated to distinguish between large and small particles. These classification models exhibited only modest predictive performance, indicating that this small dataset is insufficient to fully exploit the capabilities of complex LLMs. Overall, this study demonstrates that carefully curated small data sets, paired with robust classical ML, can effectively support the synthesis of Cu NPs and highlights that for lab-scale studies, complex models like LLMs may offer limited benefits.

Figures & tables

Explore similar work

CardsList
  1. Coupling Language Models with Physics-based Simulation for Synthesis of Inorganic Materials

    May 29, 2026Edward W. Staley, Tom Arbaugh, Michael Pekala +6Materials ScienceLarge Language Model-Guided Scientific Discovery

  2. Predicting Scale-Up of Metal-Organic Framework Syntheses with Large Language Models

    Apr 21, 2026Peter Walther, Hongrui Sheng, Xinxin Liu +7AI-Assisted Scientific ResearchMaterials Science

  3. A Statistical Approach to Estimating Sample Size of Machine Learning Models

    Sep 10, 2026Dat Phan-Trong, Sunil Gupta, Svetha VenkateshStatistical Power Analysis