We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at http://www.brickben.ch
Figures & tables
Figure 1 : We evaluate coding agents on their ability to design LEGO assemblies from a text prompt. We introduce three settings with varying constraints: Model ( ≤ 400 parts), Set (400–4000 parts), and Alt-Build (the part inventory of set 10698). Each assembly above can be built.
Figure 2 : BrickNet validation. We evaluate GPT-5.6 Luna on the BrickNet text-to-assembly validation set ( Kulits & Schmid, 2026 ) . While it was not specialized for this dataset, we find it exceeds task-specific baselines including BrickGPT ( Pun et al., 2025 ) across metrics and nears the alignment of the ground-truth objects. (a) Sample inputs and model outputs. (b) Caption alignment following BrickNet’s evaluation protocol. PE ( Bolya et al., 2025 ) and SigLIP 2 ( Tschannen et al., 2025 ) score the similarity between the embeddings of the renders and of the caption, and VQAScore ( Lin et al., 2024 ) the probability that a VQA model answers yes when asked whether the renders show the caption. See Kulits & Schmid (2026) for details on metric computation.
Table 1 : BrickBench. We report metrics for each of the three evaluation settings and in the aggregate ( Overall ). Valid requires that the assembly satisfies the part requirement of the setting and is both collision-free and stable. VQA is the proportion of VQA questions satisfied ( Sec. 4.4 ). ELO , Align ELO , and Design ELO are ratings estimated from pairwise VLM judgments ( Sec. 4.4 ). Cost is the API cost in US dollars, or its equivalent for BrickGPT ( Pun et al., 2025 ) , of producing one assembly. *The Valid metric of BrickNet-14B ( Kulits & Schmid, 2026 ) doesn’t include stability as the model doesn’t produce assemblies in a canonical frame and the direction of gravity is undefined.
Figure 3 : BrickBench summary. We report summary metrics averaged across the three evaluation settings. See Tab. 1 for disaggregated results. The ELO-plot lines represent 95% intervals.
Figure 4 : BrickBench samples. We show one assembly from each of the eleven agents for three prompts in each setting. We observe the agents vary in their interpretation of the prompts. One assembly is missing for Muse because it did not deliver one within the turn budget.
Figure 5 : ELO. (a) We plot ELO against mean assembly cost. (b) To validate our model-based Design ELO computation, we perform a perceptual study where participants are tasked with judging design quality between a pair of assemblies for a given prompt. The prompt isn’t given to the judge, and they make their decision based only on the images. We find that the rankings agree on 34/36 agent pairs (Kendall τ=0.89 ), suggesting the metric is an effective proxy for human judgment. The study covers the nine reference agents; Opus 5.5 and GPT-6.1 Sol were evaluated after it.
Table 2 : Environment ablation. We ablate the effect of the BrickAgent environment, pooled over the three evaluation settings ( Overall ; 300 prompts). While both agents consistently produce valid assemblies when given tools, they struggle without them to varying degrees: GPT-6 Astra is somewhat robust but its assemblies drop to being valid only 40% of the time while GPT-5.6 Luna’s fall to less than a percent. The environment improves the VQA of Astra but lowers its ELO within noise. Without it, Luna improves on both, suggesting the buildability constraint limits expressivity.
Figure 6 : Human-design evaluation. We perform a perceptual study comparing agent designs with human designs. For each assembly in the Model and Set settings, we select a random assembly from the BrickNet dataset with a comparable number of parts and ask a human rater to identify the one made by a person. We find that raters are able to distinguish between the agent-designed and human-designed assemblies (a); the study covers the nine reference agents and the two baselines. While agent-designed assemblies often satisfy the prompt criteria, there is a clear gap in fine-grained design to those produced by humans (b).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1 : VQA question types. We plot the share of each setting’s questions by DSG type.
Figure A2 : Question-graph sample. We plot a sample question graph for a set prompt. An arrow from one question to another indicates the second is satisfied only if the judge answers yes to the first.
Figure A3 : Judge prompts. The three prompts of Sec. 4.4 as the judge receives them. For the two pairwise questions, the four views of each assembly are tiled into one image and the two images precede the text, and the design question never contains the prompt. For VQA , the eight views precede the text as separate images, and the bracketed fields are filled in for each assembly.
Table A1 : Further assembly statistics. We report further statistics of the assemblies per setting. Parts is the average number of parts in an assembly. Unique is the (average) number of unique parts. Colors is the number of colors. Collisions is the number of parts that nonphysically collide (overlap) with one another. Stable is the proportion of assemblies that remain static under simulation, with every part moving under 3 LDU. Conn. Components is the number of graph components after BrickNet parsing; a valid assembly may have multiple components in order to represent multiple objects.
Figure A4 : BrickBench. We visualize disaggregated BrickBench quantitative results. The baselines BrickNet-14B ( Kulits & Schmid, 2026 ) and BrickGPT ( Pun et al., 2025 ) are only evaluated on the Model setting and the Valid metric reported of BrickNet-14B does not include stability.
Figure A5 : Further assembly statistics. We plot the statistics of Tab. A1 by setting.
Figure A6 : Additional BrickBench samples. We show further assemblies from the eleven agents in the three settings.
Figure A7 : VQA breakdowns. We report score by prompt category (a) and attribute type (b).
Figure A8 : Part-count distributions. We report the density of part counts in the Model ( ≤ 400) and Set (400–4000) settings for each agent. Most agents peak near 400 in Set . GPT-6 Astra uses considerably more parts on average in both settings.
We train a language model to generate LEGO-brick build sequences. While prior work has been restricted to discrete, voxel-like towers, we consider a much broader set of pieces, encompassing thousands of part types with diverse connection semantics. To enable this, we first collect a large-scale dataset of over 100,000 human-designed LDraw brick objects and scenes. The complexity of our setting makes it challenging to autoregressively assemble structures that satisfy physical constraints. When predicting block pose directly, build sequences quickly become invalid after a small number of steps. Although pieces are placed in 3D space, it is the spatial relationships of the parts which define the whole. With this in mind, we design a graph-based program representation that parametrizes structure through connectivity, improving the physical grounding of generated sequences. To enable future applications, we make our dataset and models available for research purposes. https://kulits.github.io/BrickNet
Peter Kulits, Cordelia Schmid
Inria, `Ecole Normale Sup`erieure, CNRS, PSL Research University · Max Planck Institute for Intelligent Systems, Tübingen
We introduceWorkBenchMark, a LEGO Duplo-based robotic assembly benchmark motivated by the RoboCup Smart Manufacturing League. Robotic assembly couples low-level manipulation with task-level symbolic reasoning under physical constraints, a combination that current end-to-end learning methods do not yet solve reliably. The benchmark provides 400 tasks across four complexity tiers. We provide an open-vocabulary perception, Assembly-by-Disassembly baseline solution. Our planning-based pipeline outperforms a modern vision-language-action approach across all tiers. The benchmark, simulation environment, and baseline implementation will be released openly to support the broader robotic assembly community.
Wenbo Ma, Daniel Swoboda, Matteo Tschesche +1
Chair of Machine Learning and Reasoning (i6), RWTH Aachen University, Aachen, Germany · MASCOR Institute, FH Aachen University of Applied Sciences, Aachen, Germany
We dream of AI agents that can read arbitrary designs and construct real-world objects from reusable building blocks. As a first step toward this vision, we study whether multimodal large language models (MLLMs) possess the visual grounding and spatial reasoning capabilities required for brick assembly. We formulate brick assembly as a sequential decision-making problem, where each step involves two subtasks: brick selection, identifying the target brick from candidate components, and brick pose estimation, predicting where and how the selected brick should be placed. To support this study, we introduce BC-Bench (Brick Construction Benchmark), the first benchmark for evaluating MLLMs on assembly with diverse bricks. Experiments show that current state-of-the-art MLLMs remain far from reliable builders, struggling with fine-grained brick selection and failing at precise pose estimation. To bridge this gap, we propose Brick-Composer, a learning framework that equips MLLMs with assembly skills through three complementary signals: Human Design Sparks, which provide affordance-rich construction demonstrations; World Feedback, which grounds predicted actions in visual and physical consequences; and Synthetic Experience, which scales learning beyond existing object designs. Brick-Composer improves brick selection accuracy by over three times, substantially reduces pose estimation errors, and raises strict step-level assembly success from less than 1% to around 15%. After training, a Qwen-3-8B can correctly compose up to 42% of the steps for a complete object, suggesting that MLLMs can acquire assembly capabilities through targeted, physically grounded learning.
Jiateng Liu, Bingxuan Li, Zhenhailong Wang +8
UIUC1 · Stevens Institute of Technology2 · Northwestern University3