cond-mat.mtrl-sciDec 18, 2025

Predictive Inorganic Synthesis based on Machine Learning using Small Data sets: a case study of Hydrodynamic Diameter-controlled Cu Nanoparticles

Authors: Brent Motmans, Digvijay Ghogare, Thijs G. I. van Wijk, Joren Van Herck, Saba Heidarian, Pieter De Meyer, Berend Smit, An Hardy, +1 more

Organizations: Hasselt University, Institute for Materials Research (IUMAT), Quantum & Artificial inTelligence design Of Materials (QuATOMs), Martelarenlaan 42, B-3500 Hasselt, Belgium · Hasselt University, Institute for Materials Research (IUMAT), Hybrid Materials Design (HyMaD), Martelarenlaan 42, B-3500 Hasselt, Belgium · imec, IUMAT, Wetenschapspark 1, B-3590 Diepenbeek, Belgium · Energyville, IUMAT, Thor Park 8320, B-3600 Genk, Belgium · Hasselt University, Institute for Materials Research (IUMAT), Design and Synthesis of Inorganic Materials (DESINe), Martelarenlaan 42, B-3500 Hasselt, Belgium · Electrochemistry Excellence Centre (ELEC), Materials & Chemistry Unit, Flemish Institute for Technological Research (VITO), Boeretang 200, Mol, 2400 Belgium · Laboratory of Molecular Simulation (LSMO), Institut des Sciences et Ing´enierie Chimiques, ´Ecole Polytechnique F´ed´erale de Lausanne (EPFL), Rue de l’Industrie 17, CH-1951 Sion, Switzerland

Abstract

Cu NPs have a broad applicability, yet their synthesis is sensitive to subtle changes in reaction parameters. This sensitivity, combined with the time- and resource-intensive nature of experimental optimization, poses a major challenge in achieving reproducible and size-controlled synthesis. While ML shows promise in materials research, its application is often limited by scarcity of large high-quality experimental data sets. This study explores ML to predict the DLS-derived hydrodynamic diameter of Cu NPs using a small data set of 25 syntheses. Latin Hypercube Sampling is used to efficiently cover the parameter space while creating the experimental data set. Ensemble regression models successfully predict hydrodynamic diameters with good predictive performance given the limited dataset. Since quantitative regression requires a unique DLS-derived hydrodynamic diameter, the regression model is restricted to mono-modal DLS distributions, while a complementary classification model identifies synthesis conditions for which quantitative prediction is applicable. Using equivalent out-of-sample validation, the ML and DoE models showed comparable generalization. The final ensemble model achieved an R2=0.74 compared to 0.60 for the DoE model, while retaining the complete synthesis parameter space, making it better suited for synthesis guidance. Additionally, classification models using both random forests and LLMs are evaluated to distinguish between large and small particles. These classification models exhibited only modest predictive performance, indicating that this small dataset is insufficient to fully exploit the capabilities of complex LLMs. Overall, this study demonstrates that carefully curated small data sets, paired with robust classical ML, can effectively support the synthesis of Cu NPs and highlights that for lab-scale studies, complex models like LLMs may offer limited benefits.

Figures & tables

Explore similar work

May 29, 2026cs.AI

Coupling Language Models with Physics-based Simulation for Synthesis of Inorganic Materials

Modern generative machine learning (ML) models can propose novel inorganic crystalline materials with targeted properties; however, synthesis planning of these materials remains difficult due to the complexity of the associated physical processes and limited availability of computational tools. We introduce a novel hybrid framework to evaluate Large Language Models (LLMs) in inorganic synthesis planning by combining thermodynamic databases with simplified kinetics models to approximate realistic synthesis conditions. As a case study, we focus on the niobium-oxygen system, which features multiple industrially relevant oxide phases with well-characterized data. In computational simulations, we compare LLM-generated synthesis routes with classical path-planning algorithms, showing that the implicit priors in LLMs can yield more viable strategies. In our evaluation setting, classical search methods serve primarily as a foil rather than a direct competitor. This illustrates the relative complexity of the problem and highlights where the LLM's implicit priors add value.
Apr 21, 2026cond-mat.mtrl-sci

Predicting Scale-Up of Metal-Organic Framework Syntheses with Large Language Models

Scalable synthesis remains the gate between MOF discovery and industrial deployment, as scale-up know-how is fragmented across disparate reports. We introduce ScaleMOF, a literature-mined dataset and a positive-unlabeled learning strategy that fine-tunes large language models. Achieving 93.5% accuracy, this proof-of-concept serves as a literature-grounded ranking tool prioritizing plausible scale-up candidates.
Sep 10, 2026cs.LG

A Statistical Approach to Estimating Sample Size of Machine Learning Models

Sample size determination for machine learning (ML) prediction models is challenging because conventional power analysis typically requires the predictor-outcome relationship and effect structure to be specified a priori. Nonlinear ML models learn complex prediction surfaces that do not admit straightforward analytical power calculations. We propose a framework that approximates nonlinear ML models with localized linear representations and estimates sample size requirements by evaluating statistical power across these local regions.