RocketAgent: A Long-Horizon Engineering Agent for Multidisciplinary Design of Liquid-Rocket Thrust Chambers
Authors: Junxiang He, Runze Mao, Kun He, Teng Zhang, Liming Zheng, Ke Xiao, Zhi X. Chen
Organizations: State Key Laboratory of Turbulence and Complex Systems, School of Mechanics and Engineering Science, Peking University, Beijing, 100871, China · College of Engineering, Peking University, Beijing, 100871, China · Legendspace Intelligence Technology Co., Ltd., Beijing, 100080, China · AI for Science Institute (AISI), Beijing, 100084, China
Liquid-rocket thrust-chamber design involves interdependent analyses in which downstream constraints can require earlier design decisions to be revisited. Managing these dependencies across heterogeneous tools requires consistent design information and coordinated updates throughout the workflow. We present RocketAgent, a long-horizon engineering agent for multidisciplinary preliminary design of liquid-rocket thrust chambers. A single plan-owning Coding Agent coordinates engineering skills for performance sizing, subsystem optimization, geometry generation, and multiphysics assessment. A provenance-aware knowledge graph supports method selection, while a typed Design Intermediate Representation maintains shared parameters, artifacts, and decisions. Revision-aware checks invalidate affected results and block superseded inputs, with consequential changes subject to engineering approval. In a representative simulation-based design, RocketAgent continued from an infeasible cooling search through an engineer-authorized operating-point revision, identified feasible subsystem designs, and coordinated subsequent geometry generation and multiphysics assessment to support final configuration selection. Separate module tests assessed surrogate predictions and nozzle adaptation. A two-configuration comparison across three controlled scenarios verified the expected dependency invalidations and superseded-input blocking before solver execution. The representative case demonstrates sustained coordination across a multidisciplinary design workflow, while the controlled tests establish the behavior of the revision mechanisms supporting that execution.
Figures & tables
Figure 1: RocketAgent architecture: (a) natural-language design intent; (b) the plan-owning Coding Agent and its Think–Act–Observe–Improve cycle; (c) an illustrative thrust-chamber representation; (d) the engineering KG; and (e) specialist calculation, geometry, simulation, and optimization capabilities.
Figure 2: Conceptual workflow of RocketAgent across Conceptual Design, Embodiment Design, and Verification, coordinated by a plan-owning Coding Agent with revision-aware feedback.
Figure 3: KG semantic schema: (A) requirement semantics, (B) engineering objects, (C) cases and source-linked evidence, and (D) methods and agents. The four relation groups connect design intent to an accepted concept and numerical inputs.
Figure 4: Embodiment-design tasks linked to the thrust-chamber regions: performance, geometry, and thermal refinement (left); GA-based cooling-channel search and trade-off selection (center); and GPR-assisted NSGA-II injector optimization (right). The Pareto and surrogate plots are schematic.
Figure 5: Selected four-element thrust-chamber geometry: (a) exterior view showing the igniter port and separate oxidizer and fuel inlets; (b) longitudinal section showing the injectors, combustion chamber, nozzle, and regenerative-cooling channels; (c) enlarged injector-head cutaway showing the liquid collection chambers and internal injector passages.
Figure 6: Representative engineering workflow for the 500 N LOX/RP-1 thrust chamber. Conceptual Design defines the requirements and KG-grounded concept. Embodiment Design realizes the chamber contour, cooling-channel layout, and four- and five-element injector alternatives. Geometry realization then feeds injector VOF, chamber-combustion, and coupled thermal–structural assessments. The evidence cards summarize spray angle/pressure drop, thrust/chamber pressure, and stress/margin checks; passing evidence advances to the selected design shown at right.
Figure 7: Cooling-channel trade-off after the governed rich-side redesign. The axes show coolant pressure loss Δp and maximum wall temperature Tw,max . Hollow markers fail at least one hard gate, whereas green markers are fully feasible. The blue square marks the selected balanced design and the red triangle a temperature-priority alternative.
Figure 8: Water-reference injector-module surrogate and search evidence. (a) Held-out VOF–surrogate parity for spray-cone angle ( n=118 ). (b) Held-out parity for Fuel- and Ox-side pressure drops. (c) Standardized residuals z=(y−μ)/σ for 140 independent, search-conditioned post-selection VOF evaluations; dashed lines at z=±2 mark nominal two-standard-deviation limits, whose empirical coverage is reported in Table 1 . (d) NSGA-II nondominated objective set, colored by the Ox-pressure-drop objective, with the selected module-validation knee point highlighted.
Response
C(1) (%)
C(2) (%)
Mean z
SD z
NLPD
Spray-cone angle
50.71
80.71
0.750
1.175
3.565
Δpw,Fuel
92.86
100.00
−0.007
0.540
11.762
Δpw,Ox
64.29
93.57
−0.430
0.945
10.680
Table 1: Uncertainty assessment on 140 independent, search-conditioned post-selection water-reference VOF evaluations. Gaussian reference coverage is 68.27% at ±1σ and 95.45% at ±2σ . NLPD uses degrees for angle and Pa for pressure; its values are unit-dependent and are not compared across outputs.
Figure 9: Injector-scale VOF verification. (a) Water volume fraction αwater and (b) velocity magnitude ∣U∣ in m/s on the centerline x – z plane at t=0.1 s.
Response
GPR μ±σ
VOF mean
VOF–GPR
Relative deviation
z
Spray angle
87.621±5.523∘
84.674∘
−2.947∘
3.48%
−0.53
Fuel Δp
311.266±44.963 kPa
313.295 kPa
+2.029 kPa
0.65%
+0.05
Ox Δp
253.004±11.646 kPa
246.505 kPa
−6.498 kPa
2.64%
−0.56
Table 2: Independent endpoint comparison between the GPR prediction and the DeepFlame VOF window mean. Deviations are VOF minus GPR; the relative deviation uses the GPR mean, and z is the standardized residual.
Figure 10: Chamber-scale Fluent verification for the four-element configuration. (a) Three-dimensional static-temperature field on a 0–3400 K scale. (b) Corresponding velocity-magnitude field on a 0 – 2200ms−1 scale. Both panels show sections at z=20 , 60, 100, and 140 mm and expose the injector-side entry, chamber, convergent section, and nozzle.
Metric
Design reference
Four elements
Relative error
Five elements
Relative error
pc
1.000 MPa
0.9441 MPa
−5.59%
0.9303 MPa
−6.97%
F
500 N
471.63 N
−5.67%
460.53 N
−7.89%
Isp
180.197 s
169.939 s
−5.69%
165.940 s
−7.91%
c∗
1403.5ms−1
1362.1ms−1
−2.95%
1342.2ms−1
−4.37%
Table 3: Integral chamber-performance comparison from the Fluent reacting-flow calculations. Relative errors are measured against the corresponding design or RocketCEA reference; the preliminary-design screening criterion is ∣δ∣<10% for each listed response.
Figure 11: Coupled thermal–structural response of the selected four-element injector assembly: (a) steady nodal-temperature field obtained from the kerosene- and LOX-side film conditions; (b) von Mises stress field after transfer of the temperature field and application of the pressure loads; and (c) averaged nodal von Mises stress along the declared top-plate path. The path maximum is 254.00 MPa at X=18.175mm , corresponding to a margin of 1.38 relative to the approximately 350 MPa aged-state yield reference and satisfying the minimum criterion of 1.2. The connecting line in panel (c) is for visualization only.
Figure 12: Oxidizer-side branch mass-flow distributions from CFD under the same prescribed Ox inlet flow. Each marker gives the percentage deviation dOx,i=100(m˙Ox,i/m˙Ox−1) ; labels report the integrated Ox outlet flow in kg/s. (a) Four elements: total 0.14736 kg/s, CVOx=1.92% , and δOx,max=3.11% . (b) Five elements: total 0.14664 kg/s, CVOx=5.08% , and δOx,max=6.51% . Both layouts satisfy the flow-acceptance checks for the Ox circuit; the lower within-layout dispersion favors four elements.
Figure 13: Contour comparison after strict 0.1-mm simplification. (a) Chamber and convergent profiles from the injector face to the throat, plotted against x . (b) Divergent profiles from the throat to the exit, plotted against x−xt . Insets illustrate the coordinate definitions and are not to scale. Open markers denote retained profile points; maximum interpolation errors against the quantized source profiles are 0.1000 mm for TOC, 0.09828 mm for the 15∘ conical contour, and 0.1000 mm for the 30∘/10∘ parabolic bell.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 14: Calibration diagnostic for the 140 independent post-selection water-reference VOF evaluations. The diagonal denotes equality between empirical and nominal central-interval coverage. Curves below it indicate under-coverage and curves above it indicate conservative intervals. The candidates are selected by optimization; the curves describe this search-conditioned population, not a random held-out test or a guarantee for real-propellant operation.
Figure 15: Shared Design IR record relationships. Boxes separate fields from usage checks; solid arrows indicate information links, and the dashed arrow denotes authorization of a parameter revision subject to the state-transition contract.
Figure 16: Contract-bound execution: capability selection, skill invocation, adapter/backend execution, and checked state commitment. Dashed arrows show accepted-state reads for subsequent planning and execution.
RocketSmith is an agentic system which intelligently automates the DFAM process for the development of high powered rockets suitable for launch. The system utilizes a large language model to orchestrate the execution of software tools to validate design characteristics such as flight stability and generate the parametric design components for the rocket assembly. A collection of subagents and skills enable optimization workflows of flight parameters via iteration in both zero-shot and human-in-the-loop workflows. With this system, four distinct high power rockets with various motor and assembly configurations were developed utilizing the unique design capabilities of additive manufacturing. These assembly components were fabricated using various FDM printers, manually evaluated for flight readiness, and flight tested at a launch event. From these tests, all rockets achieved a stable launch and two of the four rockets were successfully recovered in reflyable condition. The altimeter data validated that the rockets achieved an altitude 80% of the expected apogee predicted by the agentic system, establishing consistency between simulation and experimentation.
Peter Pak, Jesse Barkley, Rumi Loghmani +3
Department of Mechanical Engineering, Carnegie Mellon University, Pittsburgh, PA, USA · Tripoli Rocketry Association, Pittsburgh, PA, USA · Machine Learning Department, Carnegie Mellon University, Pittsburgh, PA, USA
Large Language Model (LLM) agents are increasingly applied to engineering design tasks, yet existing evaluation frameworks do not adequately address multi-agent systems that combine simulation, retrieval, and manufacturing preparation. We introduce a benchmark suite with three evaluation dimensions: (1) a workflow benchmark with seven prompt styles targeting distinct cognitive demands-including direct tool use, semantic disambiguation, conditional branching, and working-memory tasks; (2) a Retrieval-Augmented Generation (RAG) benchmark with gated scoring isolating retrieval contributions to parameter selection; and (3) an High Performance Computing (HPC) benchmark evaluating end-to-end ML training orchestration on a SLURM cluster. Alongside the benchmark we present EngiAI, a Multi-Agent System (MAS) reference implementation built on LangGraph that operationalizes the benchmark by coordinating seven specialized agents through a supervisor architecture, unifying topology optimization, document retrieval, HPC job orchestration, and 3D printer control. Across four LLM backends and two EngiBench problems, proprietary models achieve 96-97% average task completion on Beams2D, while open-source 4B-parameter models reach 55-78%, with clear generational improvement. Conditional branching proves most challenging, with task completion dropping to 20-53% for the conditional style on Photonics2D. RAG gating confirms near-perfect retrieval-augmented scores (about 1.0) versus near-zero without retrieval, validating the evaluation design. On HPC orchestration, one model completes all pipeline steps in 100% of runs while another drops to 50%, revealing that multi-step instruction following degrades over long-running workflows.
This paper introduces a multi-agent framework guided by Large Language Models (LLMs) to assist in the early stages of engineering design, a phase often characterized by vast parameter spaces and inherent uncertainty. Operating under a human-in-the-loop paradigm and demonstrated on the canonical problem of aerodynamic airfoil design, the framework employs a team of specialized agents: a Coding Assistant, a Design Agent, a Systems Engineering Agent, and an Analyst Agent - all coordinated by a human Manager. Integrated within a set-based design philosophy, the process begins with a collaborative phase where the Manager and Coding Assistant develop a suite of validated tools, after which the agents execute a structured workflow to systematically explore and prune a large set of initial design candidates. A key contribution of this work is the explicit integration of formal risk management, employing the Conditional Value-at-Risk (CVaR) as a quantitative metric to filter designs that exhibit a high probability of failing to meet performance requirements, specifically the target coefficient of lift. The framework automates labor-intensive initial exploration through a global sensitivity analysis conducted by the Analyst agent, which generates actionable heuristics to guide the other agents. The process culminates by presenting the human Manager with a curated final set of promising design candidates, augmented with high-fidelity Computational Fluid Dynamics (CFD) simulations. This approach effectively leverages AI to handle high-volume analytical tasks, thereby enhancing the decision-making capability of the human expert in selecting the final, risk-assessed design.
Varun Kumar, George Em Karniadakis
aSchool of Engineering, Brown University · bDivision of Applied Mathematics, Brown University