Automated Assembly Instruction Generation from CAD Models Using Grounded Large Language Models: A Human-in-the-Loop Framework
Authors: Aaron Dsouza, Mohammed Azeez Khan, Ashutosh Mishra, Arshaan Khan, Neha K. Nair, Amar Kumar Behera
Organizations: Department of Electronics and Communication Engineering, National Institute of Technology Warangal, India · Department of Computer Science and Engineering, National Institute of Technology Warangal, India · Department of Computer Science and Engineering, Nawab Shah Alam Khan College of Engineering and Technology, Hyderabad, India · Department of Physics, National Institute of Technology Warangal, India · Department of Design, Indian Institute of Technology Kanpur, Kanpur, India
Assembly documentation is a downstream manufacturing artifact that is still usually authored by interpreting CAD models by hand. Structured product data and large language models are both available, yet studies of CAD interpretation, assembly sequence planning, instruction writing, and human oversight have largely proceeded separately. This paper formulates CAD-grounded assembly instruction generation: the production of natural-language assembly procedures constrained by structured engineering information extracted from CAD models. The proposed framework maps a STEP assembly to a typed ProductGraph intermediate representation, derives a precedence order by deterministic topological sorting, realizes each step as language conditioned only on selected graph context, attaches per-step visual documentation, and applies rule-based and model-assisted checks. PDF export remains disabled until a human reviewer resolves every quality flag. The case study establishes endto-end feasibility on a built-in six-part reference assembly: the pipeline preserves a reported assembly order and carries quantity, material, and torque into an exported manual page. Generalization and geometric validation remain open empirical questions. The contribution is an architecture that separates engineering state, deterministic reasoning, grounded language realization, verification, and human release.
Figures & tables
Study
CAD
Order
Text
Ground.
Visual
Gate
Pan et al. [ 18 ]
STEP
GA
–
–
–
–
Wang et al. [ 21 ]
CAD
Interactive
–
Semantic
–
–
Jiang et al. [ 7 ]
–
Case reuse
LLM
KG/RAG
–
–
Kelm et al. [ 10 ]
–
–
LLM
MTM text
Assist.
–
Jonek et al. [ 8 ]
Design
RAG eval.
LLM
RAG
Sim. eval.
–
Lin et al. [ 13 ]
Demo./3D
Demo.
LLM
Demo.
AR
–
Table 1: Placement of the proposed framework relative to related work. “Gate” means a reported control that blocks release until a person resolves review items. Dashes mark functions that are not the reported contribution of that study.
Figure 1: Overview of the CAD-grounded assembly instruction generation framework. STEP-derived engineering information is written into a typed ProductGraph, ordered by deterministic assembly reasoning, realized as grounded instructions, and released only after automated checks and human verification. Gold marks the ProductGraph. The remaining shades distinguish CAD input, ordering, language realization, and verification.
Field
Type
Role
parts
list of parts
Identity, material, quantity, bounding box
constraints
list of constraints
Precedence and optional torque
bom
list of entries
Aggregated bill of materials
assembly_steps
list of steps
Order, sentence, diagram, flags
metadata
dictionary
Job and pipeline metadata
Table 2: ProductGraph fields used as the contract between pipeline stages.
Figure 2: Grounding path for the fastener step of the reference assembly. Structured fields enter a step-specific ProductGraph slice. The language-realization rules permit those fields and exclude facts the slice does not contain. The sentence is the fastener instruction visible on the exported manual page, which preserves quantity and torque.
Figure 3: Implemented precedence for the six-part reference assembly. The left panel is the mapping from constraint categories to directed edges. The right panel is the unique topological order returned by Kahn’s algorithm; arrows follow that order, and the fastener record stores a torque of 2.5 Nm. The interpreter does not parse STEP AP214 kinematic mates, so the figure does not assign a kinematic type to each arrow.
Figure 4: Human-in-the-loop verification path. Deterministic rules and a model-assisted review attach flags to each generated instruction. A reviewer may edit, resolve, flag, or reorder. PDF export stays closed until the required flags are resolved.
Stage
Component
Instantiation
ΓS
CAD parser
CadQuery/OCCT, or the reference assembly
State
ProductGraph
Pydantic schema; JSON in SQLite
ΓO
Sequence stage
NetworkX; Algorithm 1
ΓV
Diagram stage
Blender/Cycles, else Pillow
ΓL
Writer stage
Llama 3.3 70B via Groq; SHA-256 cache
ΓQ
QA stage
Regex rules, then a JSON model review
Table 3: Mapping from framework stages to the proof-of-concept implementation.
Figure 5: Exported manual bands for the motor bracket, the DC motor, and the fastener of the reference assembly, taken from the stored PDF page and rendered with the schematic fallback. Each band shows the instruction. The motor-related bands also show the electrical warning, and the fastener band shows the 2.5 Nm torque callout.
Stage
Wall-clock time
CAD parsing (reference assembly)
<0.1 s
Sequence inference
<0.1 s
Diagram rendering (Pillow)
∼ 3–5 s
Instruction writing (6 calls)
∼ 8–15 s
QA review (6 calls)
∼ 8–15 s
Total
∼ 20–35 s
Table 4: Approximate pipeline latency for the six-part reference assembly on one laptop, using reference-assembly parsing, the Pillow schematic renderer, and remote model calls. Ranges reflect observed run-to-run variation, not a confidence interval.
Figure 6: Implemented framework and open research directions. Solid boxes are stages exercised by the prototype. Dashed boxes are proposed extensions, and the dashed link indicates where richer CAD interpretation would enter the ProductGraph.
Recent advances in large language models and programmatic CAD have significantly improved Text-to-CAD generation for individual parts. However, production-ready mechanical assembly generation remains largely unsolved. Unlike single-part modeling, assemblies require coordinated reasoning over multiple components, functional interfaces, assembly relations, engineering principles, and physical consistency. Consequently, directly generating executable CAD code is insufficient for constructing mechanically valid and reusable assemblies. We present AssemCAD, an axiom-grounded framework for production-ready CAD assembly generation from natural language. Instead of representing an assembly as monolithic CAD code, AssemCAD first constructs an axiomatic Assembly Specification consisting of typed parts, geometry-backed ports, executable mates, and engineering axioms. Each assembly relation is explicitly grounded in one or more engineering principles, making the resulting specification interpretable, reusable, and verifiable. To realize this specification, AssemCAD introduces a port- and mate-based CAD assembly library that executes symbolic assembly relations through deterministic mate transformations and validates declared interfaces using concrete B-Rep geometric evidence. Built on this representation and library, AssemCAD further supports on-demand synthesis of reusable parametric component factories for both standard and open-world geometries. Experiments on AssemBench show that AssemCAD substantially improves assembly preservation and physical validity over code-centric CAD generation baselines, while generalizing across different foundation-model backbones. By combining axiom-grounded assembly reasoning with deterministic geometric execution, AssemCAD extends Text-to-CAD from isolated part generation toward production-ready mechanical assembly design.
Yurui Dong, Shu Zou, Siqi Li +7
Shanghai Artificial Intelligence Laboratory · Fudan University · The Australian National University +2
Large language models can write plausible CAD scripts, but reliable industrial CAD modeling requires more than syntactically valid code: every feature, placement, and assembly relation must be accepted by an exact geometric kernel while remaining editable as parametric boundary representation geometry. We present Embodied CAD, solver-grounded LLM agents for parametric B-Rep assembly modeling. Instead of generating a complete script in one pass, the agent iteratively selects actions from a stratified L0-L4 CAD skill library, resolves them into typed geometric operations, executes them in a CAD backend, and uses solver feedback to plan, repair, and learn. The framework combines action grammar constraints, deterministic parameter resolution, and solver-derived rewards for supervised warm-up and GRPO-style refinement. We evaluate Embodied CAD on multi-step mechanical, industrial equipment, and mold-oriented assembly tasks using solver-aligned metrics: executable rate, skill accuracy, operation-family accuracy, exact policy accuracy, and task completion success. The results show that solver-grounded planning executes all strong-planner workflows in the current benchmark, while learned controllers reach high executable rates and expose the remaining gap between valid tool calls and exact long-horizon policy prediction.
Turning a CAD design into an assembly plan is still largely done by hand, requiring engineers to reason about geometric feasibility, tool access, stability, and the ergonomics of human assembly. In this work, we encode long-established design for assembly (DfA) principles into a contained, end-to-end approach for generating assembly plans. Our approach takes only a mesh assembly and produces either a step-by-step assembly manual or a structured failure report, requiring no joint metadata, fastener annotations, or additional information. Four major components of a manufacturing plan are addressed autonomously: an assembly tool list, the assembly sequence and subassemblies, an assembly manual, and design feedback for improving assemblability. For determining the sequence plan, we systematically disassemble the object in a physics simulator and apply a cost function that encodes DfA principles. Manual generation, tool labelling, and assembly feedback rely primarily on multimodal large language models. Compared with a baseline that always removes the outermost part first from Tian et al., DfA-aware sequence planning reduces simulated assembly time, measured with a robot-arm assembly-time proxy, by 35% on 136 assemblies of 5 to 30 parts. The correct tool is selected for 88.6% of assembly steps. A vision-language model judge compares the generated manuals against ablated variants, identifying which page elements carry the information a reader needs. The presented approach and open-source code are available for use by engineers or AI agents looking to rapidly accelerate the creation of manufacturing plans for a given product design.
Faustin Arion von Arx, Millicent Schlafly, Mark D. Fuge
ETH Zurich Department of Mechanical and Process Engineering Zurich, Switzerland