SPOON: Towards Coherent Compositional 3D Scene Generation from Uncalibrated Multi-view Images
Organizations: The University of Hong Kong · Shenzhen Loop Area Institute · TranscEngram · The Hong Kong University of Science and Technology · Sun Yat-sen University
Abstract
Compositional 3D scene generation aims to recover complete 3D object shapes and their spatial arrangement from visual observations. Recent image-conditioned 3D generators provide strong priors for producing high-quality object geometry, making the generation of complex scenes increasingly practical. A central challenge is therefore to spatially organize these generated assets into a globally coherent scene while remaining consistent with multi-view observations. Existing approaches either entangle scene layout with object generation or separately estimate spatial placement from view-specific observations, where pose hypotheses may remain ambiguous and inconsistent across views, often resulting in an incoherent object-camera soup. We introduce SPOON, a framework that reformulates multi-view compositional 3D generation as scene-level, geometry-grounded pose reasoning. Rather than treating view-specific object pose hypotheses independently, SPOON coordinates them using reconstruction-derived multi-view geometry through a Guide-Route-Reconcile paradigm. This progressively organizes object poses and camera configurations into a coherent scene-level spatial arrangement. Extensive experiments on ARSG-110K and MIDI-3D-Front demonstrate consistent improvements in object placement and scene composition across varying numbers of input views. On ARSG-110K, SPOON reduces scene-level and object-level Chamfer distances by 12.7% and 17.7%, respectively, compared with a strong baseline.
Figures & tables
| Method | Views | ARSG-110K | MIDI-3D-Front | ||||||||
| Gen3DSR | 1 | 0.265 | 46.72 | 0.546 | 31.95 | 0.304 | 0.123 | 40.07 | 0.157 | 38.11 | 0.363 |
| MIDI | 1 | 0.801 | 15.35 | 0.179 | 35.99 | 0.033 | 0.080 | 50.19 | 0.103 | 53.58 | 0.518 |
| I-Scene | 1 | 0.277 | 34.69 | 0.563 | 21.25 | 0.228 | 0.311 | 10.76 | 0.187 | 69.90 | 0.025 |
| 3D-Fixer | 1 | 0.159 | 68.82 | 0.197 | 57.85 | 0.519 | 0.069 | 78.67 | 0.032 | 94.39 | 0.492 |
| SAM3D | 1 | 0.146 | 65.32 | 0.257 | 48.48 | 0.445 | 0.084 | 77.06 | 0.053 | 92.62 | 0.430 |
| Variant | IoU | ||||
| Baseline | 0.083 | 87.84 | 0.020 | 96.76 | 0.597 |
| w/o Guide | 0.073 | 91.47 | 0.018 | 97.58 | 0.677 |
| w/o Route | 0.071 | 90.63 | 0.019 | 97.10 | 0.669 |
| w/o Reconcile | 0.072 | 90.20 | 0.014 | 97.50 | 0.617 |
| Ours | 0.069 | 91.73 | 0.016 | 97.66 | 0.684 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| IoU | |||
| 1 | 0.072 | 90.08 | 0.619 |
| 2 | 0.072 | 90.20 | 0.617 |
| 3 | 0.072 | 90.07 | 0.618 |