MOOSEnger is a modeling and simulation AI agent framework for the Multiphysics Object-Oriented Simulation Environment (MOOSE) ecosystem, built around a simulation-aware harness that combines an interchangeable reasoning model with grounded domain knowledge, revised simulation artifacts, MOOSE-specific validation, and executable solver feedback. This surrounding system addresses a central limitation of one-shot large language model generation: small syntax, schema, reference, or solver-configuration errors can prevent a plausible input from executing, while successful execution alone does not establish scientific correctness. MOOSEnger's simulation-aware harness integrates MOOSE knowledge retrieval, Hierarchical Input Text (HIT)-aware parsing, syntax metadata, language-server diagnostics, revision-controlled authoring, and local or MCP-backed validation and execution in a generate-check-repair-run workflow that binds evidence to each input revision and guides bounded repair before acceptance. Across 200 prompts spanning eight simulation families, the MOOSEnger harness increases executable success from 10/200 (5%) to 179/200 (89.5%) with GPT 5.2 API and from 0/200 to 153/200 (76.5%) with Gemma 4 31B. A complementary ten-case Method of Manufactured Solutions benchmark moves beyond executability: all ten generated inputs satisfy the semantic-alignment criterion, and eight execute successfully while meeting the prescribed single-mesh numerical-accuracy criterion. These results show that executable reliability depends on the complete agent system rather than on the reasoning model alone, and that simulation-aware harnessing provides a path toward physics-informed verification and future full application-level and engineering verification and validation implementation.
Figures & tables
Design requirement
MOOSEnger mechanism
Domain grounding
Retrieval over documentation and examples, MOOSE-native object and parameter metadata, and task-local artifacts
Artifact identity
Revisioned drafts, immutable validation candidates, and a single accepted input
Repair authority
The model proposes revisions, and the harness binds evidence to the inspected revision and controls candidate acceptance and promotion
Deterministic validation
HIT structure, block and object schemas, and parameter, cross-link, symbol, reference, and language-server checks
Executable evidence
Authoritative validation by the target MOOSE application, bounded probes, mesh generation, and full simulation execution
Recovery control
Structured diagnostics, failure-progress checks, finite repair budgets, and controlled replanning, escalation, or termination
Table 1 : Design requirements and corresponding MOOSEnger mechanisms.
Figure 1 : Core-plus-domain architecture. The reasoning model is replaceable, while the MOOSE plugin and core runtime retain the domain evidence, artifact lifecycle, validation policy, and execution gates. Model substitution therefore changes how a candidate is produced, not the artifact-level conditions used to accept it.
Figure 2 : MOOSEnger knowledge-grounding workflow. MOOSE documentation and examples, machine-readable syntax metadata, community or project knowledge, and application-specific knowledge packs are synchronized and transformed into provenance-aware knowledge resources. The grounding layer combines vector retrieval with a deterministic MOOSE object registry, promoted domain knowledge, and optional typed relationships. At runtime, relevant evidence is merged with source attribution to support input authoring and repair, while MOOSE validation and execution remain the authority for determining whether the generated simulation artifact is acceptable.
Figure 3 : Artifact-governed validation and repair lifecycle. Each mutable draft is frozen as an immutable candidate before the configured pre-promotion gates are evaluated. Failed candidates produce revision-bound evidence for bounded repair, controlled replanning, or termination, while only a passing candidate may be promoted. Requested execution and application-specific checks provide additional evidence and determine whether the accepted revision can support a task-level completion claim.
Assurance gate
Evidence examined
Supported claim
Structural validity
HIT structure; object and parameter schemas; symbol, reference, and language-server findings
Internally consistent under the implemented syntax, schema, and reference checks.
Requirement alignment
Explicit task constraints and their assessed representations in the candidate
Represents the assessed user requirements within the declared evaluation scope.
Configuration validity
Authoritative validation by the target MOOSE application
Accepted by the target executable at the configuration stage.
Bounded execution
Harness-controlled execution within prescribed limits
Initializes and advances within the configured probe limits.
Requested operation
Exit status, logs, outputs, and artifact manifest for the requested mesh, probe, or simulation
Successfully completes the requested operation using the accepted revision.
Application-specific verification
Manufactured-solution results or other explicitly defined application evidence
Satisfies the specified numerical or V&V criterion within its defined scope.
Table 2 : Evidence gates and the assurance claim supported by each gate.
Surface
Primary use
Capability
Command-line interface and evaluation harness
Interactive and batch workflows
Direct access to MOOSE question answering, input authoring, validation, and execution.
MCP server
Tool-enabled editors and agents
Callable MOOSE validation, mesh generation, and execution services, including local or remote execution.
A2A server
Other software agents
Discoverable delegation of complete MOOSE tasks with streamed progress and optional session continuity.
Claude Code plugin
Assistant and integrated development environment users
Interactive MOOSE authoring, revision, validation, and execution within a persistent conversation.
Table 3 : MOOSEnger access surfaces. Each surface uses the same artifact-governed harness and acceptance gates.
Problem family
Prompt coverage
Diffusion
Steady and transient diffusion on 1D and 2D domains, with standard boundary conditions and lightweight solution outputs.
Transient heat conduction
Time-dependent conduction with explicit material properties, time-varying boundary conditions, and transient execution requirements.
Solid mechanics
Small-strain elasticity in one, two, and three dimensions, with displacement or traction constraints and stress or reaction outputs.
Porous flow
Darcy and pressure-diffusion problems with permeability, viscosity, porosity, mixed boundary conditions, and selected verification quantities.
Navier–Stokes
Steady and transient incompressible-flow problems requiring coupled velocity–pressure formulations and appropriate solver configurations.
Phase field
Allen–Cahn and Cahn–Hilliard interface-evolution or coarsening problems with varied initial conditions, boundary conditions, and free-energy outputs.
Table 4 : Problem families in the 200-prompt MOOSE authoring benchmark. Each family contains 25 prompts.
Problem
Final absolute L2 error
Physics
Align.
Steady diffusion (Dirichlet)
EL2(u)=1.6448×10−4
Pass
Pass
Steady diffusion (Neumann)
—execution failed
Fail
Pass
Transient diffusion (Dirichlet)
EL2(u)=2.3225×10−2
Pass
Pass
Transient diffusion (Neumann)
EL2(u)=1.3523×10−3
Pass
Pass
Allen–Cahn phase field
EL2(η)=1.6958×10−2
Pass
Pass
Single-phase Darcy pressure
EL2(p)=1.6435×10−4
Pass
Pass
Table 6 : Results of the ten-case MMS evaluation performed by MOOSEnger using Gemma 4 31B . Physics success requires successful MOOSE execution and an accepted L2 error below the prescribed threshold. A dash indicates that no authoritative error was reported because MOOSE exited with a nonzero status.
Figure 4 : MOOSEnger evidence trail for the VTB hexagonal duct-bowing case. The figure shows the original problem specification, knowledge grounding, input drafting, validation diagnostics, repair, and the resulting clean MOOSE input artifact. The final simulation comparison is shown separately in Figure 5 so that the agent workflow and numerical outcome can be inspected independently.
Figure 5 : Final result comparison for the International Atomic Energy Agency hexagonal duct-bowing demonstration. (a) Result obtained from the MOOSEnger-generated and repaired input and (b) the published reference model. The workflow summary reports the agent effort required to obtain the accepted simulation input.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Governing model and manufactured solution
Boundary and initial conditions
1. Steady diffusion—Dirichlet
The governing equation is −∇2u=f,D=1. The manufactured solution is u∗=sin(πx)sin(πy), which gives f=2π2sin(πx)sin(πy).
BC: u=0on ∂Ω. IC: None.
2. Steady diffusion—Neumann
The governing equation is −∇2u=f,D=1, with the domain-average constraint ⟨u⟩Ω=2. The manufactured solution is u∗=2+cos(πx)cos(πy), which gives f=2π2cos(πx)cos(πy).
BC: ∂nu=∇u⋅n=0 on ∂Ω . The mean constraint ⟨u⟩Ω=2 removes the constant nullspace. No Dirichlet condition is applied. IC: None.
3. Transient diffusion—Dirichlet
The governing equation is ∂t∂u−∇2u=s, with D=1,tf=1. The manufactured solution is u∗=(1+t)sin(πx)sin(πy), which gives s=[1+2π2(1+t)]sin(πx)sin(πy).
BC: u=0on ∂Ω. IC: u(x,y,0)=sin(πx)sin(πy).
Appendix
Table 7 : The ten MMS benchmark problems. Source terms are obtained by substituting the manufactured solution into the corresponding governing equation.
Execution-based evaluation of LLM-generated code implicitly treats successful execution as a proxy for correctness. In scientific simulation, this proxy is insufficient: a generated input file can run, mesh, and converge while encoding governing equations that differ from the user's intent. We call this mismatch between intended physics and generated code the comprehension-generation gap. We instantiate this in MOOSE, where Kernel and BC objects map compositionally to weak-form residual terms, enabling deterministic reconstruction of the encoded PDE and comparison against an intended contract. We formalize this comparison as the Intent Fidelity Score (IFS), a structural metric covering governing terms, BCs, ICs, coefficients, and time scheme. Building on IFS, we develop a PDE-grounded refinement loop that uses deterministic violation reports to correct generated code iteratively. We evaluate on MooseBench, a 220-case multiphysics benchmark with PDE-level ground truth released with this work. On this benchmark, our method consistently improves mean IFS over direct generation, with gains concentrated on hard cases. On the subset where direct generation falls below IFS 0.7, refinement adds +0.22 to +0.41 absolute IFS. In the deployment audit, execution-only repair improves execution success while leaving 39-40% of all 220 cases runnable but still solving the wrong physics across the three main deployment-audit models, exposing executability and intent fidelity as separable failure modes. Static proof-of-concept experiments on four PDE-oriented DSLs (UFL/FEniCS, FreeFEM, FiPy, and Devito) suggest that the reconstruction-and-comparison pattern extends beyond MOOSE. These findings reinforce that executable simulation code should be verified against the mathematical structure it is intended to encode, not accepted on execution alone.
Zhenghan Song, Yulong Liu, Cheng Wan +4
Cornell University · Columbia University · Harvard University & Nanyang Technological University
The promise of AI-driven scientific discovery hinges on whether AI agents can autonomously design and execute the computational workflows that underpin modern science. Molecular dynamics (MD) simulation presents a natural test bed to stress-test this claim; it requires translating physical intuition into syntactically and semantically correct input scripts, reasoning about initial and boundary conditions, diagnosing numerically unstable trajectories, and interpreting outputs against known physical behavior and laws. We introduce MDGYM, a benchmark of 169 expert-curated MD simulations spanning LAMMPS and GROMACS, two widely used MD packages, across three increasing difficulty levels. We evaluate three agentic frameworks -- Claude Code, Codex, and OpenHands -- with four LLMs, and find that all perform poorly: even the strongest agent solves only 21% of easy-level tasks, with less than 10% at higher difficulties. Trajectory analysis reveals a characteristic pattern of failure -- agents successfully invoke simulation machinery but produce physically unstable configurations, fabricate numerical outputs without executing the underlying computation, or abandon tasks prematurely rather than iterating through simulation-specific errors. These failure modes are qualitatively distinct from those observed in general software engineering benchmarks, indicating that fluent code generation does not transfer to grounded physical reasoning.
Vinay Kumar, Satyendra Rajput, Mausam +1
Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi · Department of Computer Science and Engineering, Indian Institute of Technology Delhi · Department of Civil and Environmental Engineering, Indian Institute of Technology Delhi +1
Existing multi-agent Large Language Model (LLM) frameworks for code generation typically use execution feedback and improve iteratively using Input/Output (I/O) test cases. However, this does not work for scientific workflows, where I/O test cases do not exist, and generating them requires solving the very problem at hand. To address this, we introduce MOSAIC, a training-free multi-agent framework for scientific code generation without I/O supervision. Instead of execution feedback, MOSAIC employs a student-teacher knowledge distillation framework that grounds generation through domain-specific examples and structured problem decomposition. To further mitigate hallucinations across chained subproblems, we introduce a Consolidated Context Window (CCW) for maintaining consistent reasoning across agents. Experiments on the SciCode benchmark show that MOSAIC improves accuracy, executability, and numerical precision over existing approaches while relying on lightweight models.
Siddeshwar Raghavan, Tanwi Mallick
Purdue University · Electrical and Computer Engineering · West Lafayette, USA +3