A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively writes and executes Blender code, inspects scenes and renderings, and revises the program. To evaluate end-to-end scene recovery, we introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse indoor and outdoor scenes. LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance. Its simulator-grounded design enables extensibility and precise automatic evaluation. Among evaluated agents, GPT-6-astra achieves the strongest overall results, with 53.4% indoor and 39.6% outdoor scores, yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance. Analysis of agent construction trajectories reveals three recurring issues: weak scene initialization, regressive edits during iteration, and unreliable self-evaluation. These findings motivate LEGO-Plugin, a training-free harness plugin for more controlled iterative scene construction, which improves all six evaluated models, with relative gains of up to 62.7% in overall score. Finally, we test whether reconstructed scenes can represent natural images and support vision tasks. In LEGO-World, we derive object detections, instance masks, and relative depth as deterministic queries on scenes reconstructed by GPT-6-astra. These readouts show non-trivial performance across all three tasks but fall well short of specialized vision models, suggesting that program-constructed scenes from current coding agents are a promising but not yet sufficiently precise representation of natural images.
Figures & tables
Figure 1: From image to scene program. Given a single RGB image, a coding agent iteratively plans, edits Blender code, renders, and inspects the result against the input, producing an editable and executable scene program. The resulting scene program can also serve as a representation that can be queried directly for detection, segmentation, and depth estimation.
Benchmark
Image Input
Scene-level Eval.
Executable / Interactive
Indoor/Outdoor Coverage
Natural-Image Style
Controlled Difficulty
3DCodeBench ( Gao et al., 2026 )
✓
✗
✗
✗
✗
✗
WorldCoder-Bench ( Lu et al., 2026b )
✗
✓
✓
✗
✗
✗
P3D-Bench ( Yang et al., 2026 )
✓
✗
✗
✗
✗
✗
SEIG ( He et al., 2026 )
✓
✗
✓
✗
✗
✗
SceneActBench ( Zhao et al., 2026 )
✓
✓
✓
✗
✗
✗
LEGO-Bench (Ours)
✓
✓
✓
✓
✓
✓
Table 1: Comparison of related 3D benchmarks. ✓ : supported; ✗ : not supported. LEGO-Bench uniquely targets end-to-end evaluation of executable scene reconstruction across diverse, realistic scenes with controlled difficulty.
Figure 2: Agentic scene construction, abstracted from a GPT-6-astra trajectory. Four blocks define the process: the Agent Workspace manages artifacts, the Action Space exposes agent operations, the Blender Code Structure organizes editable scene state, and the Scene Construction Workflow iterates from reference interpretation to validated final artifacts. Together, they enable executable, feedback-driven scene reconstruction.
Figure 4
Indoor
Outdoor
Model (+ Harness)
V ↑
R ↑
A ↑
S ↑
V ↑
R ↑
A ↑
S ↑
General-Purpose Coding Agents
GPT-6-astra + Codex
100.0 (0.0)
52.4 (1.1)
54.4 (0.7)
53.4 (0.8)
98.0 (1.0)
34.0 (1.5)
45.5 (0.5)
39.6 (0.8)
GPT-6-sol + Codex
99.4 (1.1)
22.2 (2.4)
42.5 (0.3)
32.3 (1.0)
99.7 (0.6)
19.8 (1.0)
28.7 (0.5)
24.2 (0.4)
GPT-6-luna + Codex
99.7 (0.5)
15.8 (1.2)
30.5 (0.3)
23.2 (0.7)
99.7 (0.6)
13.2 (0.6)
21.4 (0.5)
17.3 (0.5)
GPT-5.6-sol + Codex
97.8 (2.3)
9.8 (2.1)
19.8 (0.6)
14.8 (1.0)
98.7 (0.6)
12.8 (1.8)
17.7 (0.4)
15.3 (0.8)
Table 3: Performance ( % ) on LEGO-Bench. Higher is better for every metric. As a group, general-purpose coding agents reliably produce valid artifacts across both Indoor and Outdoor splits, with GPT-6-astra achieving the strongest overall performance.
Indoor
Outdoor
Tier
R
A
R
A
Easy
23.3 (0.1)
30.4 (0.6)
18.4 (0.9)
25.3 (0.4)
Medium
19.7 (1.2)
27.9 (0.6)
14.2 (1.0)
22.8 (0.9)
Hard
18.8 (0.4)
27.4 (0.5)
12.7 (1.0)
22.4 (0.3)
Table 4: Performance across scene-complexity tiers.
Figure 5: Scene quality over construction time. GPT-6-astra reaches an evaluable scene earlier and refines it steadily, whereas GPT-5.6-sol starts later and regresses more often. The gap shows that reliable construction depends on fast initialization and protection against regressive edits.
Figure 6: Judge agreement with metric directions. Reconstruction judgments remain near or below chance, and self-judging provides no consistent advantage over cross-model judging.
Figure 7: LEGO-Plugin. A harness-compatible plugin adds Enhanced Initialization, Grounded Refinement, and Version Control to the vanilla coding-agent harness. These modules directly address weak initialization, unreliable self-evaluation, and regressive edits without changing the underlying agent.
Figure 9: Qualitative LEGO-Plugin examples. Relative to the best of three GPT-5.6-sol base runs, LEGO-Plugin better matches both references and more than doubles overall scores ( 0.147→0.325 ; 0.109→0.278 ).
Task
Metric
Methods
Score
Detection
AP ↑
DINO
59.88
COCO BBox
LEGO-Anything
30.14
Segmentation
AP ↑
Segment Anything 3
53.96
LVIS Mask
LEGO-Anything
14.75
Table 5: Task readout from frozen reconstructed scenes on three 100-image tracks. Without task-specific training, the same scenes support detection, segmentation, and depth, but remain substantially behind specialist models.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: Cost–performance frontier for coding agents on LEGO-Bench. Each point is one coding-agent configuration, using full-benchmark mean Final score and a standardized token-cost proxy per input. The dashed line connects the non-dominated frontier.
Figure 11: Trajectory diagnostics on the common-task intersection. GPT-6-astra reaches the first evaluable scene later in wall-clock time but earlier relative to its own runtime, and it has the lowest regression rate and best-to-final regret. Points and bars denote means and scene-family bootstrap 95% intervals.
Figure 12: A severe best-to-last regression example. Each model is shown on the same House bedroom task at its best and last observed checkpoints. Even the strongest model can regress late in the trajectory. Displayed Q is the ungated average of Reconstruction and Appearance.
Model + Harness
V ↑
R ↑
A ↑
S ↑
GPT-5.6-sol + Codex
93.3 (5.8)
11.2 (7.9)
24.3 (3.3)
16.4 (3.6)
GPT-5.6-terra + Codex
86.7 (5.8)
0.9 (1.3)
23.0 (4.2)
10.7 (2.8)
GPT-5.6-luna + Codex
80.0 (10.0)
7.6 (6.0)
21.1 (2.9)
11.4 (2.8)
GPT-6-astra + Codex
100.0 (0.0)
24.0 (5.0)
46.6 (0.3)
35.3 (2.5)
GPT-6-sol + Codex
100.0 (0.0)
17.8 (2.6)
40.0 (0.8)
28.9 (1.0)
GPT-6-luna + Codex
100.0 (0.0)
6.7 (4.7)
34.4 (0.7)
20.0 (2.7)
Appendix
Table 6: Performance ( % ) on the Bird’s-Eye reconstruction stress split. Higher is better for every metric.
Figure 14: Representative LEGO-Bench inputs. LEGO-Bench includes indoor, outdoor, and Bird’s-Eye reconstruction settings with controlled scene complexity.
Figure 15: Scored-object count per logical scene. Counts increase within paired families, but global count is not used as the formal difficulty definition.
Figure 16: Distribution of scored-object size. Object scale spans small manipulable items through architectural and city-scale structures.
Manifest statistic
Count
Canonical assets
443
Confirmed static
368
Validated articulated
75
Joint requirements
327
Continuous
75
Revolute
55
Appendix
Table 7: Coverage of the canonical articulation GT manifest. Joint requirements are parsed from the exact human-accepted candidate URDFs.
Figure 17: Human-in-the-loop articulation annotation process. The figure traces how the 443 canonical assets move across review states over 15 human annotation rounds, from initial review to static acceptance, articulated acceptance, or regeneration. The second round deliberately re-examines both first-pass articulation positives and unresolved assets, so the articulated count need not increase monotonically. After 15 rounds, all assets are resolved as 368 static and 75 articulated.
Figure 18: Overview of the LEGO-Bench metrics. Validity checks usable scene artifacts; the main table also treats unresolved headline-evaluation failures as invalid. Reconstruction is illustrated with step-wise point-cloud states from the evaluator: visible GT surfaces from private depth, visible prediction surfaces from the submitted scene, and per-object sets induced by private masks. Appearance compares a fresh evaluator rerender against the reference image without geometric warping or camera refitting.
Figure 19: Reconstruction threshold.
Figure 23: Human study interface used for metric validation. Left: annotator login. Middle: a worked example shown before annotation. Right: the main annotation page, where the annotator compares two blinded full-scene reconstructions against a shared reference and records one A/B/Tie judgment for each marked object or region.
Figure 24: The vanilla user prompt for LEGO-Bench.
Builder / Judge
GPT-6-astra
GPT-6-sol
GPT-6-luna
GPT-5.6-sol
GPT-5.6-terra
GPT-5.6-luna
GPT-6-astra
61.7 / 76.7
51.7 / 78.3
50.0 / 83.3
58.3 / 86.7
51.7 / 76.7
53.3 / 78.3
GPT-6-sol
45.0 / 75.0
45.0 / 78.3
46.7 / 76.7
43.3 / 75.0
48.3 / 81.7
45.0 / 63.3
GPT-6-luna
51.7 / 71.7
56.7 / 76.7
43.3 / 66.7
46.7 / 71.7
41.7 / 70.0
41.7 / 65.0
GPT-5.6-sol
41.7 / 48.3
45.0 / 60.0
38.3 / 53.3
46.7 / 48.3
46.7 / 55.0
46.7 / 46.7
GPT-5.6-terra
45.0 / 58.3
40.0 / 66.7
36.7 / 56.7
40.0 / 70.0
43.3 / 65.0
43.3 / 55.0
GPT-5.6-luna
41.7 / 51.7
43.3 / 46.7
45.0 / 35.0
35.0 / 43.3
41.7 / 41.7
35.0 / 38.3
Appendix
Table 8: Builder–judge directional agreement (%). Each cell reports Reconstruction / Appearance agreement on 60 archived checkpoint pairs. Rows are builders; columns are judges. both_bad and failed judgments count as incorrect.
Track
Attempted images
Scene outputs scored
Matched diagnostic pairs
COCO boxes
100
94
93
LVIS masks
100
94
93
ETH3D depth
100
99
99
Appendix
Table 9: Coverage of the natural-image comparisons. Official box/mask AP retains all 100 subset images, including missing scene predictions. Matched diagnostics use only jointly valid pairs.