cs.LGAug 5, 2026

CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications

Authors: Brendan SmithSusana Lopez-MorenoEric Dolores-CuencaSangil KimJose L. Mendoza-CortesNijamudheen Abdulrahiman

Organizations: Kernfield Labs, London, United Kingdom. · Department of Mathematics, Pusan National University, Republic of Korea. · Humanoid Olfactory Display Center, Pusan National University, Republic of Korea. · Industrial Mathematics Center, Pusan National University, Republic of Korea. · Yonsei University, Republic of Korea. · Department of Chemical Engineering & Materials Science, Michigan State University, United States. · Department of Physics & Astronomy, Michigan State University, East Lansing, Michigan 48824, United States.

Abstract

CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications. CheMLFlow targets a common bottleneck in scientific machine learning development, where researchers often need to assemble data acquisition, curation, representation, model training, validation, screening, interpretation, and reporting into a reproducible pipeline, even when their primary research contribution concerns only one stage. CheMLFlow provides modular workflow components, ready-to-run reference pipelines, standardized artifacts, and evaluation outputs that reduce orchestration overhead and support benchmarking across methods and datasets. The platform is designed to be extensible, reproducible, and automation friendly, with pluggable representations and models, deterministic splits, explicit run artifacts, batch execution, and report generation. As scientific software increasingly moves toward agent assisted experimentation, CheMLFlow's configuration driven workflows and structured outputs also provide a practical interface for coding agents to help users construct experiments, inspect results, and summarize findings under human supervision. This article describes the system architecture, core workflows, and benchmarks that reach literature performance for quantum mechanical, physicochemical and bioactivity property prediction, and use cases involving time series datasets demonstrating applications beyond molecular chemistry datasets.

Explore similar work

CardsList