AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents
Authors: Jiahao Zhang, Yeying Fan, Moitreya Chatterjee, Suhas Lohit, Bernhard Egger, Tim K. Marks, Anoop Cherian, Stephen Gould
Organizations: The Australian National University · Tsinghua University · Mitsubishi Electric Research Laboratories (MERL) · Friedrich-Alexander-Universität Erlangen-Nürnberg
The task of 3D assembly requires translating an understanding of parts and their relationships into precise spatial arrangements. Can pretrained general-purpose agents assemble objects through visual interaction without additional assembly-specific fine-tuning? To investigate this question, we introduce AssemblyWorld, an interactive 3D environment in which agents inspect rendered views and manipulate supplied rigid parts, guided by images or assembly manuals when available. Agents perceive part geometry through 2D views rather than direct access to mesh vertices or faces, while their resulting assemblies are evaluated geometrically. Building on this environment, we construct AssemblyWorldBench, comprising 100 assembly tasks across 80 objects spanning furniture, industrial assembly, and fracture reassembly. Evaluating eight agent systems reveals substantial differences in their capabilities. The strongest system achieves 80.9% part accuracy but 59.4% complete-assembly success. The evaluated open-source systems lag substantially behind their stronger closed-source peers in both execution reliability and assembly accuracy. Analyses of visual references, interaction trajectories, and failures show how agents revise assemblies while leaving residual positioning errors. AssemblyWorld provides a common setting for both assessing the capabilities of interactive assembly agents and characterizing the gap between approximate structure recovery and precise reconstruction.
Figures & tables
Figure 2: Interaction dynamics of Astra and Fable across 100 evaluations each. (a–b) Mean call shares within active evaluation-bins. (c) Median per-evaluation translation and rotation magnitudes; translation is normalized by the largest-part diagonal. (d) Mean PA and SR. All four plots use benchmark source weights and normalized interaction time. In (c–d), colors identify systems; solid and dashed lines use the left and right axes, respectively.
4 Method
N
SCD ↓
PA ↑
SR ↑
Manual-PA ICCV’25
280
4.24
70.04
33.57
AssemblyDyno CVPR’26
280
3.91
71.21
34.64
Astra
279
8.54
78.32
55.91
Table 4: Assembly results on AssemblyBench. N counts evaluated objects; PA/SR are percentages and SCD is scaled by 1000. Astra is evaluated on 279 of 280 objects, excluding one refusal involving a dangerous item.
8 Method
RE ↓
TE ↓
PA ↑
CD ↓
Jigsaw NeurIPS’23
26.30∘
6.43
73.64
10.47
PF++ ICLR’25
20.68∘
4.37
83.33
6.68
GARF ICCV’25
10.62∘
2.10
91.00
2.12
RPF NeurIPS’25
6.32∘
2.18
96.90
2.53
TORA-CKA ECCV’26
3.03∘
0.80
97.28
0.26
SARe-Gen arXiv’26
7.78∘
0.35
98.53
0.23
Table 5: Assembly results on Fantastic Breaks. RE/TE are rotation/translation errors. Units: RE in degrees, PA in percent, TE/CD scaled by 100/1000. RE is Euler-angle root mean square error (RMSE) for Jigsaw, PF++, GARF, and Astra; geodesic otherwise. Astra uses anchor alignment on 150 objects.
Figure 3: GPT-6 Astra interaction trajectories across three assembly domains. Recorded states are re-rendered from the corresponding camera views. Labels show environment-call numbers (#) and elapsed time; selected MCP calls connect successive views. Insets highlight the IKEA end-frame correction. Part IDs and arguments are abbreviated; angles are in degrees.
Figure 4: SR gains from reference images, with paired 95% intervals: benchmark (top) and Astra on larger PartNet sets (bottom). Labels give gains in percentage points.
Figure 5: Recorded failure cases and targets. (a) Logical grouping leaves parts dispersed. (b) Numerical pose placement leaves components detached. (c) Eight of nine parts are correct, with one residual positional error. Red marks incorrect parts; teal marks the matched target in (c). Views in (a–b) are fitted independently.
Method
RE ↓
TE ↓
PA ↑
CD ↓
Agent
11.45
2.54
91.67
6.98
GARF
14.78
5.16
83.67
8.82
Agent → GARF
9.34
4.24
87.00
6.94
GARF → Agent
10.35
1.84
93.00
4.35
Table 6: Bidirectional refinement on 150 Fantastic Breaks objects with shared point samples and anchor-aligned evaluation. Arrows indicate execution order. RE is in degrees, PA in percent, and TE/CD scaled by 100/1000.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Tools
Information or operation
list_objects
Lists object and group identifiers.
get_object , get_scene , get_state
Reports body poses, local bounding extents, scene conventions, groups, and camera state. Mesh vertices and faces are unavailable.
capture_scene
Returns a 1024×768 image with a 38∘ vertical field of view.
move_camera
Orbits, zooms, or pans using yaw, pitch, zoom, and lateral/vertical offsets.
translate_objects
Applies a translation vector in world or camera coordinates.
rotate_objects
Applies X , Y , then Z angles in degrees, in world or camera coordinates, about a specified pivot. The default pivot is the mean of selected body origins.
Appendix
Table A-2: Environment tools shared by all evaluated systems. Coordinates use normalized scene units with world up +Z ; quaternions use (w,x,y,z) . Group identifiers select their constituent parts.
Dataset
N
Median
Interquartile range
90th percentile
Ratio >10 (%)
AssemblyBench
279
4.39
2.29–8.90
15.73
20.1
IKEA-Manual
102
2.20
1.67–3.28
4.52
3.9
PartNet
1472
1.79
1.00–3.22
8.32
8.6
Appendix
Table A-6: Within-object part-size ratios on the evaluated source-level sets. Each ratio divides the largest part PCA bounding-box diagonal by the smallest. PartNet assigns equal weight to objects across all three categories.
Condition
2–5 parts
6–10 parts
11–20 parts
Chair / None
40: 51.50/25.00
358: 61.90/28.50
393: 50.40/17.00
Table / None
136: 85.80/79.40
237: 73.50/49.40
160: 55.90/22.50
Table / Image
136: 93.70/86.80
237: 85.30/66.70
160: 78.30/39.40
Storage / None
1: 60.00/0.000
42: 58.50/28.60
105: 46.60/6.700
Storage / Image
1: 100.0/100.0
42: 69.70/33.30
105: 61.20/14.30
Appendix
Table A-7: Part-count bands within complete PartNet category–reference conditions. Each cell shows sample count and mean PA/SR in percent. Bands keep objects with equal part counts together; category and reference conditions are not pooled.
Setting
N
Mean
Median
k
Contribution (%)
Remaining mean
Chair / NR
791
69.90
2.88
40
94.62
3.96
Table / NR
533
46.98
1.21
27
94.53
2.71
Storage / NR
148
141.60
0.58
8
98.87
1.70
Table / IR
533
20.95
0.49
27
96.02
0.88
Storage / IR
148
184.80
0.53
8
99.32
1.33
Appendix
Table A-8: Concentration of PartNet SCD for Astra. The largest k=⌈0.05N⌉ values define the upper tail; its contribution is the fraction of summed SCD. Mean and median use all objects; the last column excludes the upper tail only as a diagnostic. SCD uses the same scale as Table 2 .
Figure A-1: Correctness-threshold sensitivity for all eight systems. PA/SR use equal source weights; saved registration and Hungarian correspondence remain fixed. The input-shape equivalence threshold remains fixed.
System
Trans. early
Late
Rot. early
Late
Revisit (%)
Astra
0.573
0.051
81.7
44.0
81.7
Fable
1.072
0.053
72.5
19.0
87.5
Opus
1.529
0.332
73.8
45.0
88.1
Sol
0.827
0.273
90.0
85.0
88.9
Sonnet
1.984
0.916
90.0
90.0
90.0
Terra
0.386
0.390
90.0
90.0
75.0
Appendix
Table A-10: Early-to-late interaction changes. Each entry is a source-weighted median of per-evaluation action medians within the first or last third of the recorded interaction interval. Translation is normalized by the largest-part diagonal; rotation is in degrees. Revisit is the median fraction of edits affecting a previously edited part.
Figure A-2: Recorded-state truncation at absolute budgets. Earlier termination retains the final state. Diamonds in the shaded Final column show untruncated outcomes, not an additional time budget; five episodes have pose changes after 60 minutes of end-to-end time. Terminal values agree with the result tables. These are retrospective curves, not new budget-conditioned agent runs.
Figure A-3: Four common objects selected near Astra’s median PA in each source, evaluated by Astra, Fable, Opus, and Sol. Columns retain a shared view direction and per-part colors; each view is fitted to its own geometry, so apparent size is not a shared physical scale. Numbers are PA under the common benchmark protocol. The other four systems appear in Figure A-4 .
Figure A-4: The same four objects evaluated by Sonnet, Terra, Qwen, and DeepSeek, with the same rendering conventions as Figure A-3 .
Figure A-5: Per-object changes in shape CD for the two composition orders. Each point is one of the same 150 objects; points below the diagonal improve. Both axes are logarithmic.
Figure A-6: GPT-6 Astra under task variations on five fixed samples per benchmark block, with one run per condition. PartNet reference conditions share objects and perturbations. Missing-part evaluation measures only retained parts; Fantastic Breaks is excluded from this condition.
Source
Object
Original
Layout A
Layout B
Distractor
Missing
PartNet NR
40074
81.80/0
100.0/1
81.80/0
72.70/0
30.00/0
2738
0.000/0
0.000/0
0.000/0
0.000/0
0.000/0
23814
18.20/0
18.20/0
18.20/0
18.20/0
10.00/0
22320
76.90/0
100.0/1
100.0/1
100.0/1
75.00/0
46475
54.50/0
54.50/0
54.50/0
27.30/0
20.00/0
PartNet IR
40074
81.80/0
81.80/0
81.80/0
100.0/1
60.00/0
Appendix
Table A-12: Paired task-variation results. Each cell reports PA (%) / SR (0 or 1); the missing-part column instead reports retained-part PA / retained-part completeness. A dash indicates an untested condition. Each condition is run once.
Assembling objects from parts requires understanding multimodal instructions, linking them to 3D components, and predicting physically plausible 6-DoF motions for each assembly step. Existing datasets focus on simplified scenarios, overlooking shape complexities and assembly trajectories in industrial assemblies. We introduce AssemblyBench, a synthetic dataset of 2,789 industrial objects with multimodal instruction manuals, corresponding 3D part models, and part assembly trajectories. We also propose a transformer-based model, AssemblyDyno, which uses the instructional manual and the 3D shape of each part to jointly predict assembly order and part assembly trajectories. AssemblyDyno outperforms prior works in both assembly pose estimation and trajectory feasibility, where the latter is evaluated by our physics-based simulations.
Danrui Li, Jiahao Zhang, Bernhard Egger +4
Rutgers, The State University of New Jersey, USA · The Australian National University, Australia · Friedrich-Alexander-Universität Erlangen-Nürnberg, Germany +1
Robotic assembly in high-mixture settings requires adaptable systems that can handle diverse parts, yet current approaches typically rely on policies specialized to each insertion task. Although this can reach high success rates, it makes the process of deploying systems for new problems tedious and time consuming. We present a framework for generalizable insertion using world models that combine robot proprioceptive information with raw visual observations captured by a wrist-mounted camera. Our model-based approach trains a single world model on up to 90 insertion tasks with geometrically diverse parts, achieving 56% zero-shot success on unseen objects with unknown geometry compared to just 7% with a model-free baseline. Importantly, performance improves as more objects are included in the training dataset, demonstrating strong scalability. Lastly, finetuning the generalist model on held-out objects significantly enhances data-efficiency compared to training from scratch and, in some cases, achieves better asymptotic performance. To our knowledge, this is the first system capable of assembling unseen objects in an entirely data-driven manner, and thus represents a significant step toward scalable, generalizable robotic assembly systems.
Nicklas Hansen, Iretiayo Akinola, Yijie Guo +7
NVIDIA · University of California San Diego · Work completed during an internship at NVIDIA +1
We introduceWorkBenchMark, a LEGO Duplo-based robotic assembly benchmark motivated by the RoboCup Smart Manufacturing League. Robotic assembly couples low-level manipulation with task-level symbolic reasoning under physical constraints, a combination that current end-to-end learning methods do not yet solve reliably. The benchmark provides 400 tasks across four complexity tiers. We provide an open-vocabulary perception, Assembly-by-Disassembly baseline solution. Our planning-based pipeline outperforms a modern vision-language-action approach across all tiers. The benchmark, simulation environment, and baseline implementation will be released openly to support the broader robotic assembly community.
Wenbo Ma, Daniel Swoboda, Matteo Tschesche +1
Chair of Machine Learning and Reasoning (i6), RWTH Aachen University, Aachen, Germany · MASCOR Institute, FH Aachen University of Applied Sciences, Aachen, Germany