cs.CVSep 30, 2026

SPOON: Towards Coherent Compositional 3D Scene Generation from Uncalibrated Multi-view Images

Authors: Guibiao Liao, Mochu Xiang, Heng Li, Ken Deng, Zijie Wang, Guanbin Li, Ping Tan, Shenghua Gao, +1 more

Organizations: The University of Hong Kong · Shenzhen Loop Area Institute · TranscEngram · The Hong Kong University of Science and Technology · Sun Yat-sen University

Abstract

Compositional 3D scene generation aims to recover complete 3D object shapes and their spatial arrangement from visual observations. Recent image-conditioned 3D generators provide strong priors for producing high-quality object geometry, making the generation of complex scenes increasingly practical. A central challenge is therefore to spatially organize these generated assets into a globally coherent scene while remaining consistent with multi-view observations. Existing approaches either entangle scene layout with object generation or separately estimate spatial placement from view-specific observations, where pose hypotheses may remain ambiguous and inconsistent across views, often resulting in an incoherent object-camera soup. We introduce SPOON, a framework that reformulates multi-view compositional 3D generation as scene-level, geometry-grounded pose reasoning. Rather than treating view-specific object pose hypotheses independently, SPOON coordinates them using reconstruction-derived multi-view geometry through a Guide-Route-Reconcile paradigm. This progressively organizes object poses and camera configurations into a coherent scene-level spatial arrangement. Extensive experiments on ARSG-110K and MIDI-3D-Front demonstrate consistent improvements in object placement and scene composition across varying numbers of input views. On ARSG-110K, SPOON reduces scene-level and object-level Chamfer distances by 12.7% and 17.7%, respectively, compared with a strong baseline.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Sparse auto-regressive modeling for scene generation from multi-view images

    Sep 3, 2026Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel +43D Scene Generation3D Generative Models

  2. Scene-SAM3D: Multi-View Scene Asset Generation Without Fine-Tuning

    Jul 18, 2026Yuqi Zhang, Yadan Luo, Xiangyu Sun +33D Scene GenerationIndoor Scene Generation

  3. ROAR-3D: Routing Arbitrary Views for High-Fidelity 3D Generation

    May 20, 2026Hanxiao Sun, Mingxin Yang, Shuhui Yang +5High-Fidelity 3D Generation3D Generation