We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at http://www.brickben.ch
Figures & tables
Figure 1 : We evaluate coding agents on their ability to design LEGO assemblies from a text prompt. We introduce three settings with varying constraints: Model ( ≤ 400 parts), Set (400–4000 parts), and Alt-Build (the part inventory of set 10698). Each assembly above can be built.
Figure 2 : BrickNet validation. We evaluate GPT-5.6 Luna on the BrickNet text-to-assembly validation set ( Kulits & Schmid, 2026 ) . While it was not specialized for this dataset, we find it exceeds task-specific baselines including BrickGPT ( Pun et al., 2025 ) across metrics and nears the alignment of the ground-truth objects. (a) Sample inputs and model outputs. (b) Caption alignment following BrickNet’s evaluation protocol. PE ( Bolya et al., 2025 ) and SigLIP 2 ( Tschannen et al., 2025 ) score the similarity between the embeddings of the renders and of the caption, and VQAScore ( Lin et al., 2024 ) the probability that a VQA model answers yes when asked whether the renders show the caption. See Kulits & Schmid (2026) for details on metric computation.
Table 1 : BrickBench. We report metrics for each of the three evaluation settings and in the aggregate ( Overall ). Valid requires that the assembly satisfies the part requirement of the setting and is both collision-free and stable. VQA is the proportion of VQA questions satisfied ( Sec. 4.4 ). ELO , Align ELO , and Design ELO are ratings estimated from pairwise VLM judgments ( Sec. 4.4 ). Cost is the API cost in US dollars, or its equivalent for BrickGPT ( Pun et al., 2025 ) , of producing one assembly. *The Valid metric of BrickNet-14B ( Kulits & Schmid, 2026 ) doesn’t include stability as the model doesn’t produce assemblies in a canonical frame and the direction of gravity is undefined.
Figure 3 : BrickBench summary. We report summary metrics averaged across the three evaluation settings. See Tab. 1 for disaggregated results. The ELO-plot lines represent 95% intervals.
Figure 4 : BrickBench samples. We show one assembly from each of the eleven agents for three prompts in each setting. We observe the agents vary in their interpretation of the prompts. One assembly is missing for Muse because it did not deliver one within the turn budget.
Figure 5 : ELO. (a) We plot ELO against mean assembly cost. (b) To validate our model-based Design ELO computation, we perform a perceptual study where participants are tasked with judging design quality between a pair of assemblies for a given prompt. The prompt isn’t given to the judge, and they make their decision based only on the images. We find that the rankings agree on 34/36 agent pairs (Kendall τ=0.89 ), suggesting the metric is an effective proxy for human judgment. The study covers the nine reference agents; Opus 5.5 and GPT-6.1 Sol were evaluated after it.
Table 2 : Environment ablation. We ablate the effect of the BrickAgent environment, pooled over the three evaluation settings ( Overall ; 300 prompts). While both agents consistently produce valid assemblies when given tools, they struggle without them to varying degrees: GPT-6 Astra is somewhat robust but its assemblies drop to being valid only 40% of the time while GPT-5.6 Luna’s fall to less than a percent. The environment improves the VQA of Astra but lowers its ELO within noise. Without it, Luna improves on both, suggesting the buildability constraint limits expressivity.
Figure 6 : Human-design evaluation. We perform a perceptual study comparing agent designs with human designs. For each assembly in the Model and Set settings, we select a random assembly from the BrickNet dataset with a comparable number of parts and ask a human rater to identify the one made by a person. We find that raters are able to distinguish between the agent-designed and human-designed assemblies (a); the study covers the nine reference agents and the two baselines. While agent-designed assemblies often satisfy the prompt criteria, there is a clear gap in fine-grained design to those produced by humans (b).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1 : VQA question types. We plot the share of each setting’s questions by DSG type.
Figure A2 : Question-graph sample. We plot a sample question graph for a set prompt. An arrow from one question to another indicates the second is satisfied only if the judge answers yes to the first.
Figure A3 : Judge prompts. The three prompts of Sec. 4.4 as the judge receives them. For the two pairwise questions, the four views of each assembly are tiled into one image and the two images precede the text, and the design question never contains the prompt. For VQA , the eight views precede the text as separate images, and the bracketed fields are filled in for each assembly.
Table A1 : Further assembly statistics. We report further statistics of the assemblies per setting. Parts is the average number of parts in an assembly. Unique is the (average) number of unique parts. Colors is the number of colors. Collisions is the number of parts that nonphysically collide (overlap) with one another. Stable is the proportion of assemblies that remain static under simulation, with every part moving under 3 LDU. Conn. Components is the number of graph components after BrickNet parsing; a valid assembly may have multiple components in order to represent multiple objects.
Figure A4 : BrickBench. We visualize disaggregated BrickBench quantitative results. The baselines BrickNet-14B ( Kulits & Schmid, 2026 ) and BrickGPT ( Pun et al., 2025 ) are only evaluated on the Model setting and the Valid metric reported of BrickNet-14B does not include stability.
Figure A5 : Further assembly statistics. We plot the statistics of Tab. A1 by setting.
Figure A6 : Additional BrickBench samples. We show further assemblies from the eleven agents in the three settings.
Figure A7 : VQA breakdowns. We report score by prompt category (a) and attribute type (b).
Figure A8 : Part-count distributions. We report the density of part counts in the Model ( ≤ 400) and Set (400–4000) settings for each agent. Most agents peak near 400 in Set . GPT-6 Astra uses considerably more parts on average in both settings.
Chair of Machine Learning and Reasoning (i6), RWTH Aachen University, Aachen, Germany · MASCOR Institute, FH Aachen University of Applied Sciences, Aachen, Germany