SPOON: Towards Coherent Compositional 3D Scene Generation from Uncalibrated Multi-view Images
Authors: Guibiao Liao, Mochu Xiang, Heng Li, Ken Deng, Zijie Wang, Guanbin Li, Ping Tan, Shenghua Gao, +1 more
Organizations: The University of Hong Kong · Shenzhen Loop Area Institute · TranscEngram · The Hong Kong University of Science and Technology · Sun Yat-sen University
Compositional 3D scene generation aims to recover complete 3D object shapes and their spatial arrangement from visual observations. Recent image-conditioned 3D generators provide strong priors for producing high-quality object geometry, making the generation of complex scenes increasingly practical. A central challenge is therefore to spatially organize these generated assets into a globally coherent scene while remaining consistent with multi-view observations. Existing approaches either entangle scene layout with object generation or separately estimate spatial placement from view-specific observations, where pose hypotheses may remain ambiguous and inconsistent across views, often resulting in an incoherent object-camera soup. We introduce SPOON, a framework that reformulates multi-view compositional 3D generation as scene-level, geometry-grounded pose reasoning. Rather than treating view-specific object pose hypotheses independently, SPOON coordinates them using reconstruction-derived multi-view geometry through a Guide-Route-Reconcile paradigm. This progressively organizes object poses and camera configurations into a coherent scene-level spatial arrangement. Extensive experiments on ARSG-110K and MIDI-3D-Front demonstrate consistent improvements in object placement and scene composition across varying numbers of input views. On ARSG-110K, SPOON reduces scene-level and object-level Chamfer distances by 12.7% and 17.7%, respectively, compared with a strong baseline.
Figures & tables
Figure 1: Given multi-view images, current 3D generation pipelines often yield a soup of objects. SPOON can recover a coherent compositional 3D scene with accurate object geometry and layout.
Figure 2: Overview of SPOON. Given multiple uncalibrated images, 3D reconstruction methods first recover the scene geometry as camera poses and depth maps. Then, for each object, we prompt the mesh generation process with multi-view visual cues and guide the pose recovery with reconstruction guidance, which is then placed using an optimal layout hypothesis through routing . Finally, the object layouts and camera poses are jointly optimized and reconciled to provide a coherent compositional 3D scene.
Method
Views
ARSG-110K
MIDI-3D-Front
CDS↓
FSS↑
CDO↓
FSO↑
IoU↑
CDS↓
FSS↑
CDO↓
FSO↑
IoU↑
Gen3DSR
1
0.265
46.72
0.546
31.95
0.304
0.123
40.07
0.157
38.11
0.363
MIDI
1
0.801
15.35
0.179
35.99
0.033
0.080
50.19
0.103
53.58
0.518
I-Scene
1
0.277
34.69
0.563
21.25
0.228
0.311
10.76
0.187
69.90
0.025
3D-Fixer
1
0.159
68.82
0.197
57.85
0.519
0.069
78.67
0.032
94.39
0.492
SAM3D
1
0.146
65.32
0.257
48.48
0.445
0.084
77.06
0.053
92.62
0.430
Table 1: Main results on the ARSG-110K and MIDI-3D-Front test sets under varying numbers of input views. We report scene- and object-level metrics together with 3D bounding-box IoU for spatial layout evaluation. Best results within each input-view setting are highlighted in bold.
Figure 3: Qualitative comparison on the MIDI-3D-Front. Compared with existing single-view and multi-view methods, our method produces more accurate object placements and scene layouts.
Figure 4: Qualitative comparison on the ARSG-110K test set.
Figure 6
Variant
CDS↓
FSS↑
CDO↓
FSO↑
IoU ↑
Baseline
0.083
87.84
0.020
96.76
0.597
w/o Guide
0.073
91.47
0.018
97.58
0.677
w/o Route
0.071
90.63
0.019
97.10
0.669
w/o Reconcile
0.072
90.20
0.014
97.50
0.617
Ours
0.069
91.73
0.016
97.66
0.684
Table 4: Component ablation of our method.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Visualization of pose generation. The highlighted object exhibits an inaccurate orientation without guidance, while pose guidance steers the generated pose toward an orientation more consistent with the object layout.
α
CDS↓
FSS↑
IoU ↑
1
0.072
90.08
0.619
2
0.072
90.20
0.617
3
0.072
90.07
0.618
Appendix
Table 5: Sensitivity analysis of the routing exponent α on MIDI with four input views, with multi-view reconciliation disabled.
Figure 7: Qualitative comparison on the MIDI-3D-Front test set.
Figure 8: Qualitative comparison on the ARSG-110K test set.
Generating complete 3D scenes from sparse, unconstrained views is a fundamental challenge in 3D vision which requires reasoning beyond observed content while remaining computationally tractable. Existing feed-forward reconstruction methods are inherently limited to content visible in the input images, while 3D generative modeling is hindered by the high computational cost of dense volumetric representations and the scarcity of large-scale 3D supervision. We introduce SPAR3S, a sparse voxel-aligned 3D latent generative model for conditional scene completion without requiring ground-truth 3D data for supervision. Our key insight is to formulate 3D scene generation in a structured, compact, voxel-aligned 3D latent space where only occupied voxels are represented. We learn this sparse latent space directly from multi-view images using photometric supervision via differentiable 3D Gaussian Splatting. Given a partial set of observed voxels encoded from sparse input views, scene completion reduces to predicting the missing latent tokens and their spatial support within the voxel grid. To this end, we train a masked autoregressive transformer that jointly models voxel occupancy and latent token values, enabling efficient and spatially consistent generation of unseen regions. We demonstrate the effectiveness of our method on synthetic indoor scenes, achieving higher novel-view quality than prior work. We further validate its generalization on RealEstate10k, highlighting its applicability to real-world data.
Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel +4
1NAVER LABS Europe, France · 2NAVER LABS, South Korea · 3Carnegie Mellon University, USA
High-quality 3D scene assets are critical for embodied applications such as robotic manipulation, navigation, and simulation. Despite their strong object priors, recent single-image 3D generation models such as SAM3D remain insufficient for real-world scenes, where severe occlusions, redundant observations, and cross-view inconsistencies make reliable scene generation challenging. We introduce Scene-SAM3D, a training-free framework that extends SAM3D from single-view object generation to calibrated multi-view scene asset generation. Scene-SAM3D selects a compact set of complementary views, reducing observation redundancy while providing additional evidence for regions occluded in individual views. Based on the selected views, it performs step-efficient latent velocity fusion to integrate multi-view evidence and suppress cross-view conflicts in canonical space. Finally, a lightweight rigid-object Gaussian optimization refines the scene layout within 200 iterations while preserving the generated object geometry. Experiments on Replica and ScanNet++ demonstrate consistent improvements at both instance and scene levels, with our method reducing scene-level CD by 43.8% on Replica and 30.9% on ScanNet++, while cutting flow-model sampling FLOPs and wall-time latency by nearly 20% under the same multi-view setting. Code will be released at https://github.com/xibi777/Scene-SAM3D.
Yuqi Zhang, Yadan Luo, Xiangyu Sun +3
The University of Queensland · Shanghai AI Laboratory · East China Normal University
Single-image-to-3D generative models can now produce high-quality geometry, yet conditioning on a single view inevitably introduces ambiguity about unseen regions. Multi-view conditioning can reduce this ambiguity, but existing methods either require fixed canonical viewpoints or rely on external reconstruction modules that impose heavy training costs and limit generation quality. We observe that pretrained single-view models already possess strong 2D-to-3D grounding that can be reused for multi-view conditioning. However, a closer analysis reveals that their conditioning mechanism entangles orientation control with geometry transfer, two functions that conflict when images from different viewpoints are naively combined. Based on this analysis, we propose ROAR-3D, a lightweight method that upgrades a pretrained single-view model to accept an arbitrary number of unposed images. A token-wise view router assigns each 3D latent token to its most relevant view, implicitly establishing 2D-to-3D correspondences without explicit pose input. A dual-stream attention design preserves the pretrained primary-view behavior while routing auxiliary views through a separate path dedicated to geometric enrichment. An orientation perturbation strategy ensures the auxiliary path learns orientation-independent geometry transfer. These components introduce minimal trainable parameters and add negligible inference overhead relative to the single-view baseline. ROAR-3D achieves state-of-the-art multi-view 3D generation quality and supports test-time view scaling from 1 to 12+ views with consistent improvements.
Hanxiao Sun, Mingxin Yang, Shuhui Yang +5
The Hong Kong University of Science and Technology · Tencent Hunyuan