LLM-Assisted Generation of Transparent, Open-Source Multiphysics Models of Electrochemical Devices
Authors: Sebastian Castro, Maya F. Schuchert, Spencer A. McCluskey, Eric W. Lees, Justin C. Bui
Organizations: Department of Chemical and Biomolecular Engineering, New York University, Brooklyn, NY, USA · Department of Chemical and Biological Engineering, University of British Columbia, Vancouver, BC, Canada
Multiphysics continuum models are powerful tools for studying electrochemical devices, enabling in silico reactor design and resolution of local pH, potential, and concentration fields that govern device performance but are difficult to measure experimentally. However, constructing such models requires substantial numerical expertise or reliance on proprietary software. Here, we show that frontier large language model agents can remove this implementation burden while keeping the underlying physics under researcher control. Using one-dimensional electrochemical CO2 reduction to CO in a porous gas diffusion electrode as a test case, we develop a machine-readable, human-specified modeling harness containing governing equations, parameters, numerical methods, logical build stages, and human-verifiable checkpoints. From this specification, the agent reproducibly constructs complete multiphysics models in open-source Julia. Independently built models, including fully autonomous agent-built models, agree with an equivalent COMSOL implementation to within 0.7% of the peak CO partial current density, and with one another to within 0.04%. Systematically planted errors demonstrate the importance of explicit specifications for reproducibility and reveal the agent's capabilities and limitations in debugging model physics. This framework establishes a more transparent approach to multiphysics modeling in which physical descriptions and governing equations, rather than specialized code, become the primary inputs for computational model development.
Figures & tables
Fig. 1: (a) Schematic of the agentic multiphysics modeling workflow employed here with a human-in-the-loop. (b) Schematic of the CO 2 reduction porous electrode model case study, along with the 5-stage implementation approach in Julia, with the primary physics additions/changes included in each stage. Elyte, electrolyte; CL, catalyst layer; GDL, gas diffusion layer; DOFs, degrees of freedom; BC, boundary condition; HER, hydrogen evolution reaction.
Fig. 2: Stage-by-stage build progression and internal fields of final model. CO (a) linear-scale polarization, (b) log-scale polarization and (c) faradaic efficiency after each build stage and the fit. Converged Stage-5 profiles at five applied potentials: across the catalyst layer, (d) pH, (e) CO 2 , (f) HCO 3 – and (g) CO 3 2– concentrations; across the gas diffusion layer and catalyst layer, (h) CO 2 and (i) CO partial pressures. Markers represent experimental data. Voltages vs the standard hydrogen electrode (SHE).
Fig. 3: Agreement of three independently built Julia models with the COMSOL model. (a) CO and H 2 partial current densities, (b) faradaic efficiencies and (c) CO 2 (aq) concentration across the catalyst layer at five applied potentials, with the deviation of each from COMSOL below it: (d) currents, (e) FE CO , in percentage points, and (f) CO 2 (aq). Current and FE deviations are evaluated at the COMSOL sweep potentials. Markers represent experimental data.
Fig. 4: Solver performance of the Julia implementation. (a) Solve time for each implementation's full sweep (89 potentials in Julia, 61 in COMSOL), both run on the same machine at matched core budget (32 cores/threads, ≈ 20,000 DOFs). (b) Measured cost of Jacobian and linear-solver choices at 20,038 DOFs. Every configuration returns the same polarization curve to the precision written, so bars differ only in cost. Ratios are relative to the configuration specified in the guide (UMFPACK with pattern reuse and forward-mode automatic differentiation). LU, lower–upper factorization; AD, automatic differentiation.
Fig. 5: How does the multiphysics agent handle errors? Errors encountered naturally during model development and planted systematically into a mature guide, their respective symptoms that the human can see and report to the agent, and the agent’s response to resolve the issue. The planted error study was conducted with Claude Opus 4.8. Agents were instructed not to consult outside resources, including other model builds, unless stated otherwise (see Supplementary Notes 6 and 8). Elyte, electrolyte; CL, catalyst layer; BC, boundary condition; α BC , boundary-condition blending factor; k MT , mass-transfer coefficient; PAC, pseudo-arclength continuation; FSG, Fuller–Schettler–Giddings; Φ L , liquid-phase potential; cond(J), condition number of the Jacobian.
Fig. 6: Autonomous agent builds without a human-in-the-loop. Comparison of Claude Opus 5 and Sonnet 5 build replicate (a) linear-scale and (b) log-scale polarization curves to COMSOL and experimental data. (c) Deviation of each build's i CO from COMSOL at the COMSOL sweep potentials.
Developing constitutive models that capture how materials deform under load traditionally requires years of specialized expertise in continuum mechanics, machine learning, and scientific programming. Large language models (LLMs) have recently been shown to lower this barrier by generating constitutive models on demand, but existing single-agent pipelines lack systematic checks that the resulting models respect fundamental physical laws. To close this gap, we introduce the first multi-agent LLM-driven approach for constitutive model generation: a Creator agent proposes a model tailored to the data, while an Inspector agent critically audits each proposal against nine physical constraints and returns it for refinement whenever a violation is detected. We demonstrate this concept with constitutive artificial neural networks (CANNs) and benchmark it on brain tissue and rubber as isotropic materials, and on porcine skin tissue as a transversely isotropic material with a preferred fiber direction, using two different LLM backbones (Claude Opus 4.7 and Kimi K2.5). Whether a generated model satisfies the physical constraints is assessed numerically, by probing each constraint across a broad sample of deformation states, rotations, and perturbation directions. Adding the Inspector raises the share of exported models that pass all these checks from 90% to 95% for Opus and from 47% to 60% for Kimi. In addition, the generated models are on par with or even surpass expert-designed models in accuracy, extrapolate reliably beyond the training data, and generalize remarkably well to unseen loading paths. Separating generation from inspection thus turns LLM-driven constitutive modeling into a substantially more trustworthy process. The paradigm is deliberately technique-agnostic...
Marius Tacke, Matthias Busch, Kian Abdolazizi +4
1Helmholtz-Zentrum Hereon, Geesthacht, Germany · 2Hamburg University of Technology, Hamburg, Germany · 3RWTH Aachen University, Aachen, Germany +2
Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results in the context of scientific claims. To address this question, we present AutoMat, a benchmark for evaluating LLM-based agents' ability to reproduce claims from computational materials science. AutoMat poses three interrelated challenges: recovering underspecified computational procedures, navigating specialized toolchains, and determining whether the resulting evidence supports a claim. By working closely with subject matter experts, we curate a set of claims from real materials science papers to test whether coding agents can recover and execute the end-to-end workflow needed to support (or undermine) such claims. We then evaluate multiple representative coding agent settings across several foundation models. Our results show that current LLM-based agents obtain low overall success rates on AutoMat, with the best-performing setting achieving a success rate of only 53%. Error analysis further reveals that agents perform worst when workflows must be reconstructed from paper text alone and that they fail primarily due to incomplete procedures, methodological deviations, and execution fragility. Taken together, these findings position AutoMat as both a benchmark for computational scientific reproducibility and a tool for diagnosing the current limitations of agentic systems in AI-for-science settings.
Building personalized cardiac electrophysiology (EP) digital twins requires identifying the appropriate model structure for each patient, not merely fitting parameters. Traditional methods rely on experts to manually prescribe hybrid physics-neural architectures, which requires deep domain expertise and does not transfer across patients. Recent works have applied large language models (LLMs) to generate or act as hybrid models. However, despite their promising generalization capacity, these LLM-based methods lack the structural priors needed for stable cardiac simulations. Hence, we propose LEADS, a framework that formulates cardiac EP domain knowledge as a structured action space and utilizes an LLM agent to discover hybrid models. The agent follows an iterative reasoning-and-action loop to select, combine, and refine hybrid models, whilst gradient descent handles parameter fitting. The proposed LEADS designs every candidate model towards physically grounded, interpretable, and numerically stable, while allowing open-ended architectural discovery. We validate LEADS on synthetic data with three ground-truth reaction models and on real cardiac EP data, demonstrating that it outperforms both human-designed hybrid models and other LLM-based hybrid modeling.
Ziqi Zhou, Yubo Ye, Sumeet Atul Vadhavka +2
Rochester Institute of Technology, Rochester, NY, 14623, USA