A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively writes and executes Blender code, inspects scenes and renderings, and revises the program. To evaluate end-to-end scene recovery, we introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse indoor and outdoor scenes. LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance. Its simulator-grounded design enables extensibility and precise automatic evaluation. Among evaluated agents, GPT-6-astra achieves the strongest overall results, with 53.4% indoor and 39.6% outdoor scores, yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance. Analysis of agent construction trajectories reveals three recurring issues: weak scene initialization, regressive edits during iteration, and unreliable self-evaluation. These findings motivate LEGO-Plugin, a training-free harness plugin for more controlled iterative scene construction, which improves all six evaluated models, with relative gains of up to 62.7% in overall score. Finally, we test whether reconstructed scenes can represent natural images and support vision tasks. In LEGO-World, we derive object detections, instance masks, and relative depth as deterministic queries on scenes reconstructed by GPT-6-astra. These readouts show non-trivial performance across all three tasks but fall well short of specialized vision models, suggesting that program-constructed scenes from current coding agents are a promising but not yet sufficiently precise representation of natural images.
Figures & tables
Figure 1: From image to scene program. Given a single RGB image, a coding agent iteratively plans, edits Blender code, renders, and inspects the result against the input, producing an editable and executable scene program. The resulting scene program can also serve as a representation that can be queried directly for detection, segmentation, and depth estimation.
Benchmark
Image Input
Scene-level Eval.
Executable / Interactive
Indoor/Outdoor Coverage
Natural-Image Style
Controlled Difficulty
3DCodeBench ( Gao et al., 2026 )
✓
✗
✗
✗
✗
✗
WorldCoder-Bench ( Lu et al., 2026b )
✗
✓
✓
✗
✗
✗
P3D-Bench ( Yang et al., 2026 )
✓
✗
✗
✗
✗
✗
SEIG ( He et al., 2026 )
✓
✗
✓
✗
✗
✗
SceneActBench ( Zhao et al., 2026 )
✓
✓
✓
✗
✗
✗
LEGO-Bench (Ours)
✓
✓
✓
✓
✓
✓
Table 1: Comparison of related 3D benchmarks. ✓ : supported; ✗ : not supported. LEGO-Bench uniquely targets end-to-end evaluation of executable scene reconstruction across diverse, realistic scenes with controlled difficulty.
Figure 2: Agentic scene construction, abstracted from a GPT-6-astra trajectory. Four blocks define the process: the Agent Workspace manages artifacts, the Action Space exposes agent operations, the Blender Code Structure organizes editable scene state, and the Scene Construction Workflow iterates from reference interpretation to validated final artifacts. Together, they enable executable, feedback-driven scene reconstruction.
Figure 4
Indoor
Outdoor
Model (+ Harness)
V ↑
R ↑
A ↑
S ↑
V ↑
R ↑
A ↑
S ↑
General-Purpose Coding Agents
GPT-6-astra + Codex
100.0 (0.0)
52.4 (1.1)
54.4 (0.7)
53.4 (0.8)
98.0 (1.0)
34.0 (1.5)
45.5 (0.5)
39.6 (0.8)
GPT-6-sol + Codex
99.4 (1.1)
22.2 (2.4)
42.5 (0.3)
32.3 (1.0)
99.7 (0.6)
19.8 (1.0)
28.7 (0.5)
24.2 (0.4)
GPT-6-luna + Codex
99.7 (0.5)
15.8 (1.2)
30.5 (0.3)
23.2 (0.7)
99.7 (0.6)
13.2 (0.6)
21.4 (0.5)
17.3 (0.5)
GPT-5.6-sol + Codex
97.8 (2.3)
9.8 (2.1)
19.8 (0.6)
14.8 (1.0)
98.7 (0.6)
12.8 (1.8)
17.7 (0.4)
15.3 (0.8)
Table 3: Performance ( % ) on LEGO-Bench. Higher is better for every metric. As a group, general-purpose coding agents reliably produce valid artifacts across both Indoor and Outdoor splits, with GPT-6-astra achieving the strongest overall performance.
Indoor
Outdoor
Tier
R
A
R
A
Easy
23.3 (0.1)
30.4 (0.6)
18.4 (0.9)
25.3 (0.4)
Medium
19.7 (1.2)
27.9 (0.6)
14.2 (1.0)
22.8 (0.9)
Hard
18.8 (0.4)
27.4 (0.5)
12.7 (1.0)
22.4 (0.3)
Table 4: Performance across scene-complexity tiers.
Figure 5: Scene quality over construction time. GPT-6-astra reaches an evaluable scene earlier and refines it steadily, whereas GPT-5.6-sol starts later and regresses more often. The gap shows that reliable construction depends on fast initialization and protection against regressive edits.
Figure 6: Judge agreement with metric directions. Reconstruction judgments remain near or below chance, and self-judging provides no consistent advantage over cross-model judging.
Figure 7: LEGO-Plugin. A harness-compatible plugin adds Enhanced Initialization, Grounded Refinement, and Version Control to the vanilla coding-agent harness. These modules directly address weak initialization, unreliable self-evaluation, and regressive edits without changing the underlying agent.
Figure 9: Qualitative LEGO-Plugin examples. Relative to the best of three GPT-5.6-sol base runs, LEGO-Plugin better matches both references and more than doubles overall scores ( 0.147→0.325 ; 0.109→0.278 ).
Task
Metric
Methods
Score
Detection
AP ↑
DINO
59.88
COCO BBox
LEGO-Anything
30.14
Segmentation
AP ↑
Segment Anything 3
53.96
LVIS Mask
LEGO-Anything
14.75
Table 5: Task readout from frozen reconstructed scenes on three 100-image tracks. Without task-specific training, the same scenes support detection, segmentation, and depth, but remain substantially behind specialist models.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: Cost–performance frontier for coding agents on LEGO-Bench. Each point is one coding-agent configuration, using full-benchmark mean Final score and a standardized token-cost proxy per input. The dashed line connects the non-dominated frontier.
Figure 11: Trajectory diagnostics on the common-task intersection. GPT-6-astra reaches the first evaluable scene later in wall-clock time but earlier relative to its own runtime, and it has the lowest regression rate and best-to-final regret. Points and bars denote means and scene-family bootstrap 95% intervals.
Figure 12: A severe best-to-last regression example. Each model is shown on the same House bedroom task at its best and last observed checkpoints. Even the strongest model can regress late in the trajectory. Displayed Q is the ungated average of Reconstruction and Appearance.
Model + Harness
V ↑
R ↑
A ↑
S ↑
GPT-5.6-sol + Codex
93.3 (5.8)
11.2 (7.9)
24.3 (3.3)
16.4 (3.6)
GPT-5.6-terra + Codex
86.7 (5.8)
0.9 (1.3)
23.0 (4.2)
10.7 (2.8)
GPT-5.6-luna + Codex
80.0 (10.0)
7.6 (6.0)
21.1 (2.9)
11.4 (2.8)
GPT-6-astra + Codex
100.0 (0.0)
24.0 (5.0)
46.6 (0.3)
35.3 (2.5)
GPT-6-sol + Codex
100.0 (0.0)
17.8 (2.6)
40.0 (0.8)
28.9 (1.0)
GPT-6-luna + Codex
100.0 (0.0)
6.7 (4.7)
34.4 (0.7)
20.0 (2.7)
Appendix
Table 6: Performance ( % ) on the Bird’s-Eye reconstruction stress split. Higher is better for every metric.
Figure 14: Representative LEGO-Bench inputs. LEGO-Bench includes indoor, outdoor, and Bird’s-Eye reconstruction settings with controlled scene complexity.
Figure 15: Scored-object count per logical scene. Counts increase within paired families, but global count is not used as the formal difficulty definition.
Figure 16: Distribution of scored-object size. Object scale spans small manipulable items through architectural and city-scale structures.
Manifest statistic
Count
Canonical assets
443
Confirmed static
368
Validated articulated
75
Joint requirements
327
Continuous
75
Revolute
55
Appendix
Table 7: Coverage of the canonical articulation GT manifest. Joint requirements are parsed from the exact human-accepted candidate URDFs.
Figure 17: Human-in-the-loop articulation annotation process. The figure traces how the 443 canonical assets move across review states over 15 human annotation rounds, from initial review to static acceptance, articulated acceptance, or regeneration. The second round deliberately re-examines both first-pass articulation positives and unresolved assets, so the articulated count need not increase monotonically. After 15 rounds, all assets are resolved as 368 static and 75 articulated.
Figure 18: Overview of the LEGO-Bench metrics. Validity checks usable scene artifacts; the main table also treats unresolved headline-evaluation failures as invalid. Reconstruction is illustrated with step-wise point-cloud states from the evaluator: visible GT surfaces from private depth, visible prediction surfaces from the submitted scene, and per-object sets induced by private masks. Appearance compares a fresh evaluator rerender against the reference image without geometric warping or camera refitting.
Figure 19: Reconstruction threshold.
Figure 23: Human study interface used for metric validation. Left: annotator login. Middle: a worked example shown before annotation. Right: the main annotation page, where the annotator compares two blinded full-scene reconstructions against a shared reference and records one A/B/Tie judgment for each marked object or region.
Figure 24: The vanilla user prompt for LEGO-Bench.
Builder / Judge
GPT-6-astra
GPT-6-sol
GPT-6-luna
GPT-5.6-sol
GPT-5.6-terra
GPT-5.6-luna
GPT-6-astra
61.7 / 76.7
51.7 / 78.3
50.0 / 83.3
58.3 / 86.7
51.7 / 76.7
53.3 / 78.3
GPT-6-sol
45.0 / 75.0
45.0 / 78.3
46.7 / 76.7
43.3 / 75.0
48.3 / 81.7
45.0 / 63.3
GPT-6-luna
51.7 / 71.7
56.7 / 76.7
43.3 / 66.7
46.7 / 71.7
41.7 / 70.0
41.7 / 65.0
GPT-5.6-sol
41.7 / 48.3
45.0 / 60.0
38.3 / 53.3
46.7 / 48.3
46.7 / 55.0
46.7 / 46.7
GPT-5.6-terra
45.0 / 58.3
40.0 / 66.7
36.7 / 56.7
40.0 / 70.0
43.3 / 65.0
43.3 / 55.0
GPT-5.6-luna
41.7 / 51.7
43.3 / 46.7
45.0 / 35.0
35.0 / 43.3
41.7 / 41.7
35.0 / 38.3
Appendix
Table 8: Builder–judge directional agreement (%). Each cell reports Reconstruction / Appearance agreement on 60 archived checkpoint pairs. Rows are builders; columns are judges. both_bad and failed judgments count as incorrect.
Track
Attempted images
Scene outputs scored
Matched diagnostic pairs
COCO boxes
100
94
93
LVIS masks
100
94
93
ETH3D depth
100
99
99
Appendix
Table 9: Coverage of the natural-image comparisons. Official box/mask AP retains all 100 subset images, including missing scene predictions. Matched diagnostics use only jointly valid pairs.
Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing render can hide incorrect spatial relations, intersecting objects, or unintended modifications. We introduce Code4Scene, a benchmark of 190 Unreal Engine cases built from human-assembled scenes that evaluates coding agents on two complementary settings under a shared execution interface. Construction tests scene-level spatial reasoning from open-ended language specifications, where many realizations are valid; editing tests precise control of scene state, where the agent must recover the target scene from reference images while preserving everything else. Rather than scoring code or rendered views, Code4Scene evaluates the generated engine-native scene for task fulfillment, artifact integrity, and static physical validity, with edits additionally compared against withheld ground truth. Across 14 coding-agent configurations on the 95-case public set, construction and editing performance are strongly correlated but not interchangeable (Spearman ρ=0.78): Claude Fable 5.1 leads construction, Gemini 3.8 Flash leads editing, and GPT-6 Astra narrowly leads overall. Spatial Composition is the weakest construction category for every agent, while editing remains imprecise: the best Repair F1 is only 0.527, and 35.8% of edits that fully recover the target still introduce unintended changes elsewhere in the scene. These results expose a gap between plausible 3D generation and reliable spatial reasoning and state control.
Xiaokang Ye, Siddhant Hitesh Mantri, Zimeng Chen +6
We present LiteReality-Agent, an agentic system for reconstructing real indoor environments as realistic, articulated, and simulation-ready 3D scenes from RGB-D scans. At its core, LiteReality-Agent formulates 3D reconstruction as a coding problem, in which a coding agent gathers evidence using specialised tools and iteratively edits a Python script, Room.py, which can be executed to produce a 3D digital twin of the room. With this formulation, we develop a robust observe-edit-verify harness that supports evidence gathering, measurement, verification, layout optimisation, simulation readiness, and quality control throughout the reconstruction process. LiteReality-Agent produces high-quality reconstructions suitable for simulation and downstream embodied AI tasks. Furthermore, as agent capabilities continue to improve rapidly, the system introduced by LiteReality-Agent remains a strong orchestration framework for future agents: it equips them with specialised tools, structured workflows, and robust verification mechanisms that substantially improve reconstruction quality and reliability. We demonstrate that LiteReality-Agent produces reconstructions that are more geometrically accurate, visually realistic, and simulation-compatible than those generated by recent frontier models, such as Astra and Fable. We therefore view LiteReality-Agent as a practical and important building block for robust real-to-sim systems. Both the source code and the data-capture application are publicly available. Code:https://github.com/LiteReality/LiteReality-Agent/
Zhening Huang, Yueyan Li, Johnathan Chiu +5
University of Cambridge · Imperial College London · Independent Researcher
Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and embodied AI a simulation-ready replica of the real environment whose objects can be manipulated individually. Existing pipelines decompose the task into three steps---parse the observations into instances, generate an asset for each, and place each asset back---but every step presumes an input that a cluttered capture rarely provides: accurate instance geometry, unoccluded views, and assets that accurately match the observations. We propose Lucida, which keeps this order but redistributes the requirements, so each step consumes only what a real capture reliably provides and precision is reached at the end of the pipeline rather than demanded at its start. Lucida parses the video into a scene graph whose nodes carry per-instance multi-view evidence, generates a complete asset for each instance from its evidence, and places assets with GizmoAct, a VLM policy that casts placement as multi-turn GUI interaction, manipulating the object's gizmo in a closed loop and deciding itself when alignment is reached. Across scene-level 3D object detection, object pose estimation, and scene reconstruction, Lucida improves mAP over Boxer by 69% on R2S-Scene, raises ADD-SB@0.05 from 57.8% to 83.4% on CA-1M, and increases scene F-Score from 0.794 for SAM3D to 0.924.
Minghan Qin, Yuang Wang, Xiuyu Yang +6
1ByteDance Seed · 2Peking University · 3Zhejiang University