Transitioning from a digital design to a robotic assembly process currently requires months of expert manual tuning to reconcile part geometries with robotic constraints. This paper presents an end-to-end, autonomous pipeline for the design and physical construction of bespoke wooden assemblies. A generative AI agent translates user prompts into initial 3D geometries, balancing the visual fidelity of the design with select physical constraints. The assemblability of the design is further improved by a gradient-based repair stage that backpropagates through a graph attention network surrogate to adjust component geometries. In addition to correcting for disjointed and overlapping components, we demonstrate hardware-specific corrections, differentiably optimizing the geometry of components to enable robot screwdriving for 86.7% of 60 novel natural language inputs, significantly outperforming prior work by a factor of ten. For ten of the structures, we physically demonstrate assemblability with two UR5e robots. This work marks a meaningful step toward on-demand robotic manufacturing, enabling the rapid production of customized, low-volume goods.
Figures & tables
Fig. 1: Woodworking On-Demand Pipeline. A. A generative AI agent interprets a natural language input to produce 8 initial geometries, of which many fail due to physics and motion planning checks. B. A differentiable graph attention network, trained from 174,000 block configurations labeled in NVIDIA Isaac Sim, performs gradient-based optimization to satisfy hardware-specific constraints, such as screwdriver clearance. C. Two UR5e robots physically assemble the optimized design to validate the framework’s feasibility for bespoke parts.
Fig. 2: Late-fork GAT architecture: Two shared GATv2Conv layers process the assembly graph, then split into a regression branch predicting standardized overlap and thickness ( y^reg ) and a classification branch driving a binary feasibility head ( ℓ^feas ) and an auxiliary failure-mode head ( ℓ^mode ). Edge features are reused as attention inputs at every GAT layer with hidden width 128, 4 attention heads and dropout 0.1.
Fig. 3: Final Assemblies for Ten Natural Language Inputs . For each prompt, we show the best design candidate produced by the generative AI agent (left) followed by the modified design (right). We physically construct each object to validate assemblability.
Gap/ Concurrence
Structural Tipping
Sufficient Overlap
Thickness
Overall Screwable
End-to-End (Full System)
Ours
93.3%(56)
96.7%(58)
98.3%(59)
86.7%(52)
86.7%(52)
36.7%(22)
Blox-Net
100%(60)
100%(60)
85.0%(51)
8.3%(5)
8.3%(5)
6.7%(4)
p=0.127
p=0.476
* p<0.05
*** p<0.001
*** p<0.001
*** p<0.001
TABLE I: Benchmarking Results From Simulated Assembly Checks. We present the percentage of 60 designs that passed each assembly check followed by the number of designs in parentheses.
Gap/ Concurrence
Structural Tipping
Sufficient Overlap
Thickness
Overall Screwable
Ours
93.3%(56)
96.7%(58)
98.3%(59)
86.7%(52)
86.7%(52)
MLP+KD-Tree Cost
21.6%(13)
71.7%(43)
63.3%(38)
36.7%(22)
11.7%(7)
Ours w/o GenAI
95.0%(57)
86.7%(52)
93.3%(56)
81.7%(49)
66.7%(40)
Ours w/o Perturbations
93.3%(56)
96.7%(58)
96.7%(58)
86.7%(52)
86.7%(52)
Ours w/o GAT (Blox-Net)
100%(60)
100%(60)
85.0%(51)
8.3%(5)
8.3%(5)
TABLE II: Ablation Results From Simulated Assembly Checks. We present the percentage of 60 designs that passed each assembly check followed by the number of designs in parentheses. For the condition without GenAI, only one candidate design is created (as opposed to 8 ).
Fig. 4: Remaining Assembly Failures. A. For the castle , the body of the gripper collides with the remaining assembly. B. The bicycle handlebar post remains longer than the available screws after optimization. C. After assemblability modifications for the wishing well and stairs , the designs are no longer visually recognizable.
Turning a CAD design into an assembly plan is still largely done by hand, requiring engineers to reason about geometric feasibility, tool access, stability, and the ergonomics of human assembly. In this work, we encode long-established design for assembly (DfA) principles into a contained, end-to-end approach for generating assembly plans. Our approach takes only a mesh assembly and produces either a step-by-step assembly manual or a structured failure report, requiring no joint metadata, fastener annotations, or additional information. Four major components of a manufacturing plan are addressed autonomously: an assembly tool list, the assembly sequence and subassemblies, an assembly manual, and design feedback for improving assemblability. For determining the sequence plan, we systematically disassemble the object in a physics simulator and apply a cost function that encodes DfA principles. Manual generation, tool labelling, and assembly feedback rely primarily on multimodal large language models. Compared with a baseline that always removes the outermost part first from Tian et al., DfA-aware sequence planning reduces simulated assembly time, measured with a robot-arm assembly-time proxy, by 35% on 136 assemblies of 5 to 30 parts. The correct tool is selected for 88.6% of assembly steps. A vision-language model judge compares the generated manuals against ablated variants, identifying which page elements carry the information a reader needs. The presented approach and open-source code are available for use by engineers or AI agents looking to rapidly accelerate the creation of manufacturing plans for a given product design.
Faustin Arion von Arx, Millicent Schlafly, Mark D. Fuge
ETH Zurich Department of Mechanical and Process Engineering Zurich, Switzerland
In production processes for consumer products, assembly instructions are essential not only for planning but also for executing the production process. Likewise in robotics, it is crucial for an assembly robot to understand how components fit together and can be assembled. To facilitate these tasks, we contribute a method for constructing scene graphs to represent and characterize assembly relationships between components. Our approach does not rely on semantic data and is capable of handling a very small dataset. To realize this, the output of a Faster R-CNN model is used to create geometric representations, which are then processed by a transformer architecture to generate an adjacency matrix. This matrix serves as input to a Siamese network that uses message passing based on an attentional graph convolutional network (aGCN) architecture to characterize the connections between the components. We validate our method on a study dataset of toy model components which can be assembled into transportation vehicles.
Christoph Jahn, Urs Waldmann, Bastian Goldluecke
1Mercedes-Benz AG, Sindelfingen, Germany · Department of Computer and Information Science, University of Konstanz, Konstanz, Germany
The task of 3D assembly requires translating an understanding of parts and their relationships into precise spatial arrangements. Can pretrained general-purpose agents assemble objects through visual interaction without additional assembly-specific fine-tuning? To investigate this question, we introduce AssemblyWorld, an interactive 3D environment in which agents inspect rendered views and manipulate supplied rigid parts, guided by images or assembly manuals when available. Agents perceive part geometry through 2D views rather than direct access to mesh vertices or faces, while their resulting assemblies are evaluated geometrically. Building on this environment, we construct AssemblyWorldBench, comprising 100 assembly tasks across 80 objects spanning furniture, industrial assembly, and fracture reassembly. Evaluating eight agent systems reveals substantial differences in their capabilities. The strongest system achieves 80.9% part accuracy but 59.4% complete-assembly success. The evaluated open-source systems lag substantially behind their stronger closed-source peers in both execution reliability and assembly accuracy. Analyses of visual references, interaction trajectories, and failures show how agents revise assemblies while leaving residual positioning errors. AssemblyWorld provides a common setting for both assessing the capabilities of interactive assembly agents and characterizing the gap between approximate structure recovery and precise reconstruction.
Jiahao Zhang, Yeying Fan, Moitreya Chatterjee +5
The Australian National University · Tsinghua University · Mitsubishi Electric Research Laboratories (MERL) +1