We propose Tetris3D, a generative framework for single-image 3D scene reconstruction that recovers objects which are physically and geometrically coherent as a scene. Existing methods often generate objects independently or couple them implicitly, providing limited guidance for ensuring fine-grained spatial compatibility between neighboring objects that interact with one another. To address this, we explicitly condition the generation of each object on the geometry of surrounding objects and their physical relationships, guiding its shape and pose to remain geometrically and physically plausible within the scene. Moreover, we introduce ComOb, a physics simulation-based dataset of 1.2M scenes featuring physical interactions across diverse object categories, with per-object meshes and pairwise physical relation annotations. Comprehensive experiments on synthetic and realworld scenes show that Tetris3D recovers coherent object shapes and poses even when interacting regions are occluded, and achieves state-of-the-art performance in both generation quality and physical stability.
Figures & tables
Figure 1: Teaser. Tetris3D reconstructs 3D scenes by generating objects that fit together as a scene. We generate objects autoregressively, explicitly conditioning each on neighboring geometry and physical relations to achieve geometric and physical coherence among objects, even under occlusion.
Figure 2: Overview of Tetris3D. Tetris3D reconstructs scenes autoregressively in a topological order determined by VLM-inferred physical relations and dependencies (Sec. 3.5 ). Each object is generated through a two-stage pipeline consisting of sparse structure and structured latent generation, with two key components in each DiT: (1) Per-Token Injection (Sec. 3.2 ): We derive cdepth from the depth map, cimg from back-projected image features, and cint from object interactions on the same 3D grid used for generation and these conditions are injected into the corresponding tokens. (2) V2I Attention (Sec. 3.3 ): Invisible tokens identified by object mask attend to visible tokens through cross-attention, using the aggregated information to complete occluded regions.
Figure 3: Schematic of object grid G and pose-aligned generation.
Method
Reconstruction Quality
Generation Quality
Physical Stability
CD-S ↓
CD-O ↓
F1-S ↑
F1-O ↑
IoU-B ↑
ICP-Rot ↓
MMD ↓
COV ↑
P-FID ↓
Uni3D ↑
ULIP ↑
PD ↓
Dmean↓
Epeak↓
Scene Generation
SAM-3D
25.19
1.70
0.2462
0.6296
0.3817
20.37
2.840
71.24
5.222
0.6004
0.6314
1.354
192.4
0.8633
ShapeR
5.72
1.61
0.6480
0.6302
0.6408
10.65
2.590
71.61
12.45
0.5165
0.5594
0.8862
115.1
0.4675
WorldSculpt
8.68
1.98
0.4302
0.5976
0.5369
17.31
2.973
70.56
4.816
0.5454
0.5776
0.1177
103.5
0.4812
Amodal Generation + Pose Estimation
Table 1: Quantitative comparison on Toys4K ( Stojanov et al., 2021 ) . † denotes methods whose generated objects are aligned to the scene using FoundationPose ( Wen et al., 2024 ) .
Method
MessyKitchens
Picasso
Reconstruction Quality
Physical Stability
Reconstruction Quality
Physical Stability
CD-S ↓
CD-O ↓
F1-S ↑
F1-O ↑
IoU-B ↑
ICP-Rot ↓
PD ↓
Dmean↓
Epeak↓
CD-S ↓
CD-O ↓
F1-S ↑
F1-O ↑
IoU-B ↑
ICP-Rot ↓
PD ↓
Dmean↓
Epeak↓
Scene Generation
SAM-3D
0.25
0.64
0.8749
0.8187
0.5417
9.43
0.0270
215.9
0.9949
0.66
1.12
0.7432
0.7254
0.5631
12.87
0.1221
1033.4
6.389
ShapeR
0.23
1.57
0.9089
0.6976
0.6152
9.38
0.0253
388.8
1.825
0.65
1.90
0.7880
0.6293
0.6887
12.04
0.0666
932.2
4.335
WorldSculpt
0.32
1.19
0.8601
0.7445
0.5552
9.88
0.0012
266.3
1.291
0.54
1.69
0.7834
0.6975
0.5087
18.66
0.0111
900.3
4.125
Table 2: Quantitative comparison on MessyKitchens ( Ansari et al., 2026 ) and Picasso ( Yu et al., 2026 ) . † denotes methods whose generated objects are aligned to the scene using FoundationPose ( Wen et al., 2024 ) .
Figure 4: Qualitative comparison. Input scene images (top) and complete target images (bottom) are shown for reference. For each scene, the top row shows reconstructed shapes and poses in scene space, and the bottom row shows individual objects from a common viewpoint.
Figure 5: Qualitative comparison. The first column shows the scene image (top) and per-object masks (bottom), with colors matching the corresponding objects in the generated scenes. For each scene, we show the reconstruction and its final state after physics simulation ( Todorov et al., 2012 ) .
Configuration
Reconstruction Quality
Generation Quality
Physical Stability
CD-S ↓
CD-O ↓
F1-S ↑
F1-O ↑
IoU-B ↑
ICP-Rot ↓
MMD ↓
COV ↑
P-FID ↓
Uni3D ↑
ULIP ↑
PD ↓
Dmean↓
Epeak↓
Full model (Ours)
3.90
1.28
0.7963
0.7691
0.7603
5.50
1.563
84.3
1.875
0.6704
0.7069
0.0192
39.22
0.1193
w/o Interaction Cond.
8.56
1.37
0.7599
0.7134
0.6918
8.66
1.841
80.66
1.938
0.6385
0.6766
0.4980
69.24
0.3240
w/o V2I Attn.
4.46
1.37
0.7905
0.7591
0.7506
6.00
1.629
84.01
1.949
0.6474
0.6897
0.0270
42.04
0.1396
w/ Amodal3R Attn.
4.29
1.34
0.7909
0.7659
0.7496
5.82
1.573
84.25
1.951
0.6615
0.7018
0.0281
40.67
0.1322
Table 3: Ablation study on interaction conditioning and amodal generation. We evaluate the effect of interaction conditions and architectural designs for amodal generation on Toys4K ( Stojanov et al., 2021 ) . For w/ Amodal3R Attn , we replace our V2I attention with the occlusion-aware attention module of Amodal3R ( Wu et al., 2025 ) .
Figure 9
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Examples of ComOb . We showcase scenes from the ComOb dataset. Each row corresponds to a different interaction type. From left to right, we show the complete target-object image, the normal map of the 3D scene, and images rendered from viewpoints used during training.
Method
Reconstruction Quality
Generation Quality
Physical Stability
CD-S ↓
CD-O ↓
F1-S ↑
F1-O ↑
IoU-B ↑
ICP-Rot ↓
MMD ↓
COV ↑
P-FID ↓
Uni3D ↑
ULIP ↑
PD ↓
Dmean↓
Epeak↓
Scene Generation
SAM-3D
24.80
1.73
0.2495
0.6250
0.3861
21.27
2.872
70.46
5.245
0.5981
0.6292
1.203
194.09
0.8269
ShapeR
10.54
2.28
0.4955
0.5241
0.4841
19.42
3.676
64.76
16.64
0.4453
0.4831
0.8408
129.81
0.5989
WorldSculpt
12.53
2.11
0.3736
0.5818
0.4593
19.48
3.113
69.95
4.941
0.5382
0.5683
0.2280
135.1
0.6655
Amodal Generation + Pose Estimation
Appendix
Table 4: Quantitative comparison on Toys4K ( Stojanov et al., 2021 ) with estimated depth. We compare Tetris3D with baselines using depth maps estimated by MoGe3 ( Kong et al., 2026 ) . All methods requiring depth use the same depth maps. † denotes methods whose generated objects are aligned to the scene using FoundationPose ( Wen et al., 2024 ) .
Method
Reconstruction Quality
Generation Quality
Physical Stability
CD-S ↓
CD-O ↓
F1-S ↑
F1-O ↑
IoU-B ↑
ICP-Rot ↓
MMD ↓
COV ↑
P-FID ↓
Uni3D ↑
ULIP ↑
PD ↓
Dmean↓
Epeak↓
Ours (GT)
3.68
0.97
0.8407
0.8084
0.7918
5.29
1.501
83.11
1.599
0.6884
0.7182
0.0036
38.91
0.1721
Ours (VLM)
3.76
0.98
0.8399
0.8063
0.7905
5.37
1.538
83.38
1.651
0.6891
0.7195
0.0038
41.25
0.1780
Appendix
Table 5: Quantitative comparison on Toys4K ( Stojanov et al., 2021 ) with VLM-based reasoning. We compare results using ground-truth physical relation annotations with those using relations inferred by a VLM ( Bai et al., 2025 ) . Tetris3D maintains comparable performance with VLM-inferred relations.
Figure 9: Qualitative comparison with agentic 3D generation. We compare Tetris3D with an LLM agent-based approach to 3D mesh generation ( OpenAI, 2026 ) . Although this approach can produce broadly similar shapes, it often fails to faithfully recover the object geometry depicted in the input image and struggles to generate physically consistent scene layouts.
Figure 10: Qualitative comparison with baselines on Toys4K ( Stojanov et al., 2021 ) .
Figure 11: Qualitative comparison with baselines on Toys4K ( Stojanov et al., 2021 ) .
Figure 12: Qualitative comparison with baselines on Toys4K ( Stojanov et al., 2021 ) .
Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense object hierarchies and irregular affordances. Here, an interactive scene denotes a physically valid, collision-free environment directly loadable into physics simulators. Existing methods, ranging from decoupled symbolic solvers to end-to-end regression models, often suffer from error propagation or overfitting to noisy supervision containing widespread physical violations. To address these limitations, we introduce PhyScene3D, a framework that reformulates generation as a Human-Mimetic Constructive Process. The proposed Cognitive Topological Reasoning Chain (CTRC) factorizes scene synthesis into a sequential, anchor-conditioned process. It employs a 3D AABB-based placement scheme that imposes a strong structural inductive bias. To address imperfect supervision and physical infeasibility, we introduce Physics-Aware Denoising Alignment (PADA). It integrates a differentiable Signed Distance Field (SDF) with Test-Time Optimization (TTO) to project generated scenes onto a physics-feasible manifold while preserving semantic intent. Experiments demonstrate that PhyScene3D outperforms state-of-the-art approaches in both semantic accuracy and physical validity, achieving a 40% reduction in scene-wise collision rate relative to the human-annotated training data.
Weixing Chen, Zhuoqian Feng, Yang Liu +6
Sun Yat-sen University, China · Guangdong Key Laboratory of Big Data Analysis and Processing · Huawei +1
Reconstructing interactive, simulation-ready 3D scenes from a single image is a critical bottleneck for robotic manipulation. While recent single-image lifters recover plausible per-object shapes, composing them yields scenes that collapse under physical simulation due to interpenetrating, hovering, or sinking objects. Existing physics-aware methods address this strictly as a post-hoc layout correction, leaving the underlying geometric errors unresolved. To address this, we introduce SimuScene, a compositional 3D reconstruction pipeline that puts physics in the loop of shape and layout estimation. Rather than using physics merely for layout cleanup, we utilize the physics engine as a diagnostic measurement tool during the generative process itself. By diagnostically simulating reconstructed objects under gravity, we convert penetration and support failures into quantitative correction signals that drive gravity-axis stretching and amodal shape resampling. This physics-informed feedback loop mitigates accumulated reconstruction errors and produces a stable, simulation-ready compositional 3D scene. Extensive experiments demonstrate state-of-the-art performance on physical stability and geometric alignment benchmarks. We further highlight SimuScene's utility by deploying reconstructed environments in humanoid control and robot-arm manipulation tasks.
Reconstructing physically stable 3D scenes from a single RGB image enables casual images to be converted into simulation-ready digital assets for applications such as immersive interaction and content creation. However, existing single-image reconstruction methods fall short in capturing the physical structure of a scene. As a result, they often produce geometrically plausible but physically inconsistent results, including object floating and penetration, which lead to unstable behavior in physics simulations. Image-conditioned scene generation methods improve physical plausibility but often rely on strong scene priors, yielding plausible yet inaccurate object arrangements that fail to match the input image. We propose REST3D, a single-image reconstruction framework that can reconstruct physically stable 3D scenes by integrating physical scene understanding with physics-constrained refinement. We first introduce an agentic physical scene understanding technique that constructs a scene-tree representation capturing object physical states and inter-object relationships from a gravity-support perspective, providing a structural prior for reconstruction. Leveraging this structure, we initialize the scene using image-to-3D models, followed by scene-tree-guided alignment and physics-constrained optimization to resolve physical violations while preserving visual consistency with the input image. Experiments show that our method significantly reduces physical errors and improves simulation stability on both synthetic and real-world datasets while maintaining strong reconstruction quality. We further demonstrate the reconstructed scenes in VR-based human-object interaction, showing their potential for immersive applications.