Single-image scene generation aims to produce a complete 3D scene mesh from a single image, including surfaces the camera did not observe. While pretrained 3D object generators encode a strong shape prior, they are mainly designed for isolated objects in a fixed canonical volume and focus mostly on indoor scenes, since diverse 3D data for outdoor scenes are quite limited. In this work, we present a method that redesigns such an object-centric generator, e.g., Trellis 2, to work on both indoor and outdoor scenes while retaining its prior. We accomplish this by (a) partitioning the scene into adaptive chunks that scale relative to the distance to the camera; nearby chunks have a smaller size to keep the finer detail, while distant structures, e.g., buildings, are covered by large chunks; (b) making the generator capture explicit 2D-3D correspondence by lifting image features and making the model aware of the free space, observed surface, and unobserved region; (c) synthesizing around 4,000 outdoor scenes to broaden the training data, as existing scene datasets are largely indoor. Experiments on Tanks and Temples, ScanNet++, and in-the-wild images show that our method outperforms all baselines in geometric accuracy and perceptual quality across both indoor and outdoor scenes.
Figures & tables
Figure 1. Generating a complete 3D mesh from a single image. Our method reconstructs detailed scene meshes from one photograph. The method completes geometry beyond the observed view and supports indoor, outdoor, and large-scale scenes. We show the Colosseum reconstruction from multiple viewpoints (left) and diverse reconstruction results (right).
Figure 2. Overview of our method. (a) Scene representation. We lift the input image using an estimated point map and partition the scene into adaptive 3D chunks. (b) Chunk generation. For each chunk, the generator first predicts the sparse structure and then its structured geometry latents. Both stages are conditioned on surface-aligned lifted image features, while the first stage further receives a depth-based visibility indicator to distinguish free space vs. occluded regions. (c) Autoregressive scene generation. We generate neighboring chunks sequentially while reusing latents in overlapping regions, allowing each new chunk to complete unseen geometry conditioned on the previously generated scene chunk. The assembled latents are finally decoded into the full scene mesh.
Figure 3. Scene chunking strategies. A single chunk (a) compresses the entire scene into one fixed volume, sacrificing geometric detail, while fixed-size chunks (b) require many chunks to cover large scenes. Our adaptive strategy (c) increases the physical chunk size with depth, preserving finer resolution nearby while efficiently covering distant regions. Best viewed zoomed in.
Figure 4. Learning from imperfect synthetic data. Our model reconstructs more faithful geometry than the imperfect training data from the agentic framework.
Figure 5. Visual comparisons. Our method preserves global scene layout while recovering finer details across diverse inputs. World Tracing ( Zhang et al., 2026 ) produces noisy occluded regions and lacks detail. GenRecon ( Schmid et al., 2026 ) often imposes a room-like prior for outdoor scenes. For Lyra 2 ( Shen et al., 2026 ) , we overlay RGB on the normal map. Although the rendering appears plausible, the underlying geometry is inaccurate. VolFill ( Ngo et al., 2026 ) has line-like surface artifacts and fails on outdoor scenes.
Tanks and Temples
ScanNet++
In-the-wild
Method
CD ↓
F1 ↑
CD ↓
F1 ↑
DreamSim ↓
CLIP-N ↑
Win rate (%) ↑
Iterative 3D completion
EvoScene ( Zheng et al., 2026 )
4.61
0.241
5.70
0.690
0.379
0.797
90.7
Extend3D ( Yoon et al., 2026 )
8.37
0.165
5.54
0.674
0.375
0.805
94.4
Object composition
3D-RE-GEN ( Sautter et al., 2026 )
7.17
0.146
15.84
0.291
0.429
0.810
88.9
Table 1. Quantitative comparison on Tanks and Temples, ScanNet++ and In-the-wild images. Our approach outperforms all baselines across different datasets and metrics.
Layout
Indoor
Outdoor
Time
CD ↓
F1 ↑
CD ↓
F1 ↑
(min) ↓
Single chunk
2.87
0.759
49.9
0.138
1
Fixed 3m
3.13
0.782
38.5
0.100
15
Adaptive
2.86
0.792
37.4
0.240
5
Table 2. Comparison of chunking strategies. Adaptive chunking improves reconstruction and is faster than fixed-size chunks.
Table 4. Ablation of dataset contribution. Our outdoor data improves generalization without sacrificing indoor quality.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1. Chunk allocation. We partition this Colosseum into overlapping depth regions and allocate chunks to cover the estimated point cloud within the camera frustum.
Figure A2. Comparison of chunking strategies. A single chunk compresses the entire scene into a fixed canonical volume, distorting depth and producing wall-like geometry. Fixed-size chunks require too many chunks for large scenes, and it ran out of memory on an 80G GPU. Our adaptive strategy efficiently covers the scene while preserving geometric detail.
Figure A3. Additional qualitative comparisons on diverse scenes. We compare surface-normal renderings across indoor and outdoor environments. Our method better preserves scene layout and local geometry, and occluded regions. For Lyra 2, we overlay its RGB rendering on the normal map.
Figure A4. Additional qualitative comparisons on diverse scenes. We evaluate scenes containing people, animals, large architectural structures, and unusual object configurations. Our method more consistently recovers coherent geometry and scene structure, while the baselines often exhibit incomplete geometry, simplified layouts, or coarse surfaces.
Figure A5. Agentic framework for outdoor scene construction. Given an input image, (a) a VLM first produces a scene plan describing objects, their dimensions and spatial relations, and ground regions. (b) Each visible object is cropped from the image, completed to recover occluded geometry, and reconstructed into a 3D mesh. (c) Ground geometry is recovered from the estimated point map, semantic layout, and height map. (d) A layout solver places the reconstructed assets while satisfying spatial constraints and avoiding collisions. (e) A VLM critic iteratively inspects the assembled scene and proposes corrections to object placement and missing content, producing the final 3D scene.
Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context from inherently ambiguous visual evidence. Despite recent progress in joint layout-and-mesh generation, existing methods often rely on holistic or weakly decomposed pipelines that entangle many factors at once and demand extensive scene-level supervision, limiting their generalization to complex real-world environments. We propose a multi-agent orchestration framework that decomposes single-image 3D scene generation into three structured stages: scene initialization, environment construction, and multi-agent refinement. The initialization stage extracts image-derived object masks, builds object-level 3D representations, and predicts an initial spatial layout to form a coarse 3D scene. The environment-construction stage then leverages this initialization together with point-map geometry to build an environmental scaffold of supporting surfaces, room boundaries, materials, and illumination. Finally, in the refinement stage, a planner agent identifies structural and visual inconsistencies, applies simple corrections directly, and dispatches specialist agents for complex localized revisions that are reintegrated into the global scene. To provide reliable structural initialization while reducing reliance on scene-level annotations, we further introduce a geometry-aware layout predictor supervised by sparse geometric priors derived from point maps. Unlike fully supervised layout generators, the predictor can be trained from segmentation-level data and generalizes robustly to diverse real-world scenes. Extensive experiments on benchmark datasets show that our method consistently outperforms prior approaches in geometric accuracy, spatial consistency, and perceptual realism.
Jeonghwan Kim, Yushi Lan, Yongwei Chen +3
Nanyang Technological University · University of Oxford · Meshy AI
Recent advances in single image-to-3D generation have enabled high-quality asset synthesis, yet extending these capabilities to indoor scene generation remains challenging. Existing methods focus on asset-level generation while neglecting the structural layout, which is essential for downstream applications and serves as the spatial anchor for grounding assets. However, a single image with a limited field of view lacks the spatial coverage to recover a coherent global layout. To this end, we use a 360° image represented in equirectangular projection (ERP) and propose InSpace, a structure-aware framework for 3D indoor scene generation. InSpace comprises three stages: (1) estimating partial scene geometry as spatial priors, (2) generating coarse scene structure with view-selective cross-attention, and (3) producing detailed layout and asset geometry with textures through a global-local hybrid attention, using flow matching. We also propose ERP-FRONT, a paired ERP-Image-to-3D indoor scene dataset based on 3D-FRONT. Experiments show that InSpace generates complete 3D indoor scenes with structural layout, along with separate textured assets from a single ERP image, achieving strong performance across 3D and 2D metrics. Project Page: https://kookie12.github.io/InSpace-Project-Page/
Gwanhyeong Koo, Hyunsu Kim, Youngji Kim +5
1KAIST · *Work done during an internship at NAVER LABS. · 2NAVER LABS +1
Geometry-conditioned 3D scene generation enables the creation of 3D environments from user-provided geometry, offering direct control over scene structure and object layout. To generate such 3D scenes, current methods commonly adopt a three-stage design that first defines a view schedule, then synthesizes multi-view observations along the scheduled views, and finally reconstructs a 3D representation from the generated images. However, defining the view schedule becomes a major bottleneck for outdoor scenes, where large, unstructured, and unbounded geometry makes it difficult to obtain views that provide sufficient coverage while supporting stable generation. To address this bottleneck, we present SceneFrom3D, a framework that automatically schedules views from outdoor input geometries. SceneFrom3D constructs a directed generation graph whose nodes represent anchor views and whose edges represent interpolation trajectories, defining which views to synthesize, which view pairs to interpolate, and in which order generation should proceed. Beyond automatic view scheduling, SceneFrom3D further improves controllability through object-level conditioning, assigning each object an identity image for appearance guidance and a geometry-adherence parameter for region-wise control over the input geometry. Experiments demonstrate that SceneFrom3D achieves state-of-the-art geometry-conditioned outdoor 3D scene generation, producing high-quality scenes with controllable object appearance and geometry adherence.
Geonung Kim, Jeongeun Park, Nuri Ryu +2
POSTECH, Republic of Korea · Meta Reality Labs, United States of America