Object-centric scene reconstruction requires completing partial object observations while preserving metric alignment and avoiding collisions with the surrounding. Existing generation-based methods are often image-conditioned and suffer from scale ambiguity and insufficient geometric constraints. We propose COOL, a framework for COllision-aware and Observation-aLigned reconstruction. Based on an object generation model, COOL conditions the generation on instance and background point clouds. Instance geometry anchors generation in scene coordinates, while background geometry provides local context for scene-consistent completion. We further introduce an explicit collision loss and use joint optimization and resampling to reduce collisions during inference. Experiments on 3D-Front and Scan2CAD demonstrate strong scene-level fidelity, observation alignment, and collision reduction. Moreover, additional studies validate its robustness to mask errors and its applicability to real-world scene replicas.
Figures & tables
Fig. 1: Image-conditioned generation-based methods for object-centric scene reconstruction produce inaccurate object locations and relative sizes. Our method fully exploits scene-coordinate metric geometry to gain an advantage in generating observation-aligned and collision-aware objects.
Fig. 2: Overview of COOL. Using segmented and normalized scene point cloud as input, COOL extracts the instance condition and context condition from instance and background point clouds through the condition extraction network. In conditional generation module, these conditions are introduced through MHA layers in the flow transformer. The generated objects are placed back to the reconstructed scene to compute the geometric overlap as a collision loss, optimizing collisions via a joint optimization and resampling strategy during inference.
Fig. 3: A case of joint optimization and resampling. (i) Initial generation with collision. (ii) Gradient-guided optimization converged to a local optima. (iii) Resampling escaping the local optima. (iv) Gradient-guided optimization achieving a collision-free reconstruction.
Method
3D-Front
Scan2CAD
CDS↓
FSS↑
CDO↓
FSO↑
IoUB↑
ColO↓
ColS↓
CDS↓
FSS↑
CDO↓
FSO↑
IoUB↑
ColO↓
ColS↓
AdaPoinTr
0.002
0.965
0.026
0.895
0.853
0.010
0.078
0.018
0.841
0.062
0.532
0.597
0.329
0.023
InstPIFu
0.065
0.633
0.088
0.507
0.269
0.460
0.503
0.087
0.518
0.072
0.494
0.219
0.510
0.159
MIDI
0.008
0.929
0.026
0.834
0.579
0.072
0.076
–
–
–
–
–
–
–
Ours
0.001
0.989
0.011
0.904
0.895
0.002
0.003
0.015
0.894
0.036
0.745
0.641
0.027
0.000
TABLE I: Quantitative comparisons on the 3D-Front dataset and the Scan2CAD dataset
Fig. 4: Qualitative Comparisons on the 3D Front Dataset.
Fig. 5: Qualitative Comparisons on the Scan2CAD Dataset.
Method
CDS↓
FSS↑
CDO↓
FSO↑
IoUB↑
ColO↓
ColS↓
Gen3DSR
0.050
0.655
0.085
0.443
0.398
0.291
0.093
SAM3D
0.032
0.747
0.040
0.691
0.382
0.322
0.050
3D-Fixer
0.042
0.791
0.059
0.595
0.464
0.286
0.067
Ours(zs)
0.029
0.841
0.068
0.587
0.549
0.041
0.000
TABLE II: Quantitative comparisons on the Scan2CAD dataset with zero-shot methods
C.C.
C.L.
R.S.
CDS↓
FSS↑
CDO↓
FSO↑
IoUB↑
ColO↓
ColS↓
tinfer↓
✗
✗
✗
0.019
0.876
0.040
0.710
0.618
0.245
0.016
3.2s
✓
✗
✗
0.014
0.896
0.037
0.744
0.624
0.203
0.015
3.4s
✓
✓
✗
0.016
0.891
0.038
0.742
0.633
0.095
0.003
3.9s
✓
✗
✓
0.018
0.885
0.038
0.740
0.622
0.089
0.009
5.9s
✓
✓
✓
0.015
0.894
0.036
0.745
0.641
0.027
0.000
7.8s
TABLE III: Ablation Studies on the Scan2CAD dataset
Fig. 6: Qualitative Results on Collision Loss Guidance.
Mask
CDS↓
FSS↑
CDO↓
FSO↑
IoUB↑
SAM [ 27 ]
0.010
0.985
0.024
0.810
0.727
PanoRecon [ 28 ]
0.002
0.994
0.016
0.810
0.766
GT
0.001
0.995
0.015
0.816
0.767
TABLE IV: Study on Robustness under Perception Noise
Fig. 7: Qualitative Results under Perception Noise.
Fig. 8: Study on Multi-view Condition on the Scan2CAD Dataset.
Fig. 9: Cases on Geometric Replicas of Real-World Scenes.
Accurately reconstructing complex full multi-object scenes from sparse observations remains a core challenge in computer vision and a key step toward scalable and reliable simulation for robotics. In this work, we introduce RecGen, a generative framework for probabilistic joint estimation of object and part shapes, as well as their pose under occlusion and partial visibility from one or multiple RGB-D images. By leveraging compositional synthetic scene generation and strong 3D shape priors, RecGen generalizes across diverse object types and real-world environments. RecGen achieves state-of-the-art performance on complex, heavily occluded datasets, robustly handling severe occlusions, symmetric objects, object parts, and intricate geometry and texture. Despite using nearly 80% fewer training meshes than the previous state of the art SAM3D, RecGen outperforms it by 30.1% in geometric shape quality, 9.1% in texture reconstruction, and 33.9% in pose estimation.
Robots operating safely in cluttered everyday environments often need to infer scene geometry from partial observations. Methods that detect objects in 2D and reconstruct them independently struggle in such scenes: a missed object is never reconstructed, a merged detection can fuse two objects, and separately reconstructed meshes may overlap or fail to touch their supporting surfaces. We introduce CODA (Complete Once, Decompose Afterward), a generative model that instead reconstructs the complete scene geometry from a single unsegmented RGB-D image, then separates the surface into the surrounding environment and movable objects. Still, generated scene geometry can drift from the observed partial point cloud. To reduce this drift, CODA uses two explicit 3D grounding mechanisms to keep reconstructed geometry consistent with observed surfaces while completing unseen regions. Experiments on HomebrewedDB and our custom cluttered-scene dataset show more accurate reconstructions and a higher fraction of objects remaining in place under simulated gravity than both object-first and scene-first baselines.
Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open challenge. A central difficulty lies in how object layout is represented. Holistic methods absorb placement into a scene-level generation process, sacrificing object-level detail. Compositional methods preserve object fidelity by decoupling geometry from layout, but typically parameterize layout as sparse, unbounded pose variables that are difficult to learn and generalize poorly under scarce scene-level supervision. We present Mira-Scene, a compositional 3D scene reconstruction framework that replaces sparse pose regression with dense, bounded correspondence recovery. At its core is the Canonical Coordinate Map (CCM), a pixel-aligned field that maps each visible object pixel to a surface coordinate in the object's bounded canonical space. When paired with a scene-space Point Cloud Map (PCM) from monocular geometry estimation, CCM induces dense canonical-to-scene correspondences from which object transformations are recovered through robust geometric alignment. Because CCM operates in bounded canonical space, it provides a stable prediction target that can be trained from scalable object-level 3D data without requiring scene-level layout annotations. Mira-Scene further introduces a multimodal diffusion transformer that jointly generates object geometry and CCMs, using modality-specific expert streams with shared attention and positional encoding to promote geometry-layout consistency. Experiments on indoor, outdoor, synthetic, and in-the-wild scenes show that Mira-Scene substantially outperforms strong baselines in layout accuracy, achieving relative gains of 39.8% in 3D-IoU and 16.5% in 2D-IoU over SAM3D, using limited open-source training data.