We propose Tetris3D, a generative framework for single-image 3D scene reconstruction that recovers objects which are physically and geometrically coherent as a scene. Existing methods often generate objects independently or couple them implicitly, providing limited guidance for ensuring fine-grained spatial compatibility between neighboring objects that interact with one another. To address this, we explicitly condition the generation of each object on the geometry of surrounding objects and their physical relationships, guiding its shape and pose to remain geometrically and physically plausible within the scene. Moreover, we introduce ComOb, a physics simulation-based dataset of 1.2M scenes featuring physical interactions across diverse object categories, with per-object meshes and pairwise physical relation annotations. Comprehensive experiments on synthetic and realworld scenes show that Tetris3D recovers coherent object shapes and poses even when interacting regions are occluded, and achieves state-of-the-art performance in both generation quality and physical stability.
Figures & tables
Figure 1: Teaser. Tetris3D reconstructs 3D scenes by generating objects that fit together as a scene. We generate objects autoregressively, explicitly conditioning each on neighboring geometry and physical relations to achieve geometric and physical coherence among objects, even under occlusion.
Figure 2: Overview of Tetris3D. Tetris3D reconstructs scenes autoregressively in a topological order determined by VLM-inferred physical relations and dependencies (Sec. 3.5 ). Each object is generated through a two-stage pipeline consisting of sparse structure and structured latent generation, with two key components in each DiT: (1) Per-Token Injection (Sec. 3.2 ): We derive cdepth from the depth map, cimg from back-projected image features, and cint from object interactions on the same 3D grid used for generation and these conditions are injected into the corresponding tokens. (2) V2I Attention (Sec. 3.3 ): Invisible tokens identified by object mask attend to visible tokens through cross-attention, using the aggregated information to complete occluded regions.
Figure 3: Schematic of object grid G and pose-aligned generation.
Method
Reconstruction Quality
Generation Quality
Physical Stability
CD-S ↓
CD-O ↓
F1-S ↑
F1-O ↑
IoU-B ↑
ICP-Rot ↓
MMD ↓
COV ↑
P-FID ↓
Uni3D ↑
ULIP ↑
PD ↓
Dmean↓
Epeak↓
Scene Generation
SAM-3D
25.19
1.70
0.2462
0.6296
0.3817
20.37
2.840
71.24
5.222
0.6004
0.6314
1.354
192.4
0.8633
ShapeR
5.72
1.61
0.6480
0.6302
0.6408
10.65
2.590
71.61
12.45
0.5165
0.5594
0.8862
115.1
0.4675
WorldSculpt
8.68
1.98
0.4302
0.5976
0.5369
17.31
2.973
70.56
4.816
0.5454
0.5776
0.1177
103.5
0.4812
Amodal Generation + Pose Estimation
Table 1: Quantitative comparison on Toys4K ( Stojanov et al., 2021 ) . † denotes methods whose generated objects are aligned to the scene using FoundationPose ( Wen et al., 2024 ) .
Method
MessyKitchens
Picasso
Reconstruction Quality
Physical Stability
Reconstruction Quality
Physical Stability
CD-S ↓
CD-O ↓
F1-S ↑
F1-O ↑
IoU-B ↑
ICP-Rot ↓
PD ↓
Dmean↓
Epeak↓
CD-S ↓
CD-O ↓
F1-S ↑
F1-O ↑
IoU-B ↑
ICP-Rot ↓
PD ↓
Dmean↓
Epeak↓
Scene Generation
SAM-3D
0.25
0.64
0.8749
0.8187
0.5417
9.43
0.0270
215.9
0.9949
0.66
1.12
0.7432
0.7254
0.5631
12.87
0.1221
1033.4
6.389
ShapeR
0.23
1.57
0.9089
0.6976
0.6152
9.38
0.0253
388.8
1.825
0.65
1.90
0.7880
0.6293
0.6887
12.04
0.0666
932.2
4.335
WorldSculpt
0.32
1.19
0.8601
0.7445
0.5552
9.88
0.0012
266.3
1.291
0.54
1.69
0.7834
0.6975
0.5087
18.66
0.0111
900.3
4.125
Table 2: Quantitative comparison on MessyKitchens ( Ansari et al., 2026 ) and Picasso ( Yu et al., 2026 ) . † denotes methods whose generated objects are aligned to the scene using FoundationPose ( Wen et al., 2024 ) .
Figure 4: Qualitative comparison. Input scene images (top) and complete target images (bottom) are shown for reference. For each scene, the top row shows reconstructed shapes and poses in scene space, and the bottom row shows individual objects from a common viewpoint.
Figure 5: Qualitative comparison. The first column shows the scene image (top) and per-object masks (bottom), with colors matching the corresponding objects in the generated scenes. For each scene, we show the reconstruction and its final state after physics simulation ( Todorov et al., 2012 ) .
Configuration
Reconstruction Quality
Generation Quality
Physical Stability
CD-S ↓
CD-O ↓
F1-S ↑
F1-O ↑
IoU-B ↑
ICP-Rot ↓
MMD ↓
COV ↑
P-FID ↓
Uni3D ↑
ULIP ↑
PD ↓
Dmean↓
Epeak↓
Full model (Ours)
3.90
1.28
0.7963
0.7691
0.7603
5.50
1.563
84.3
1.875
0.6704
0.7069
0.0192
39.22
0.1193
w/o Interaction Cond.
8.56
1.37
0.7599
0.7134
0.6918
8.66
1.841
80.66
1.938
0.6385
0.6766
0.4980
69.24
0.3240
w/o V2I Attn.
4.46
1.37
0.7905
0.7591
0.7506
6.00
1.629
84.01
1.949
0.6474
0.6897
0.0270
42.04
0.1396
w/ Amodal3R Attn.
4.29
1.34
0.7909
0.7659
0.7496
5.82
1.573
84.25
1.951
0.6615
0.7018
0.0281
40.67
0.1322
Table 3: Ablation study on interaction conditioning and amodal generation. We evaluate the effect of interaction conditions and architectural designs for amodal generation on Toys4K ( Stojanov et al., 2021 ) . For w/ Amodal3R Attn , we replace our V2I attention with the occlusion-aware attention module of Amodal3R ( Wu et al., 2025 ) .
Figure 9
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Examples of ComOb . We showcase scenes from the ComOb dataset. Each row corresponds to a different interaction type. From left to right, we show the complete target-object image, the normal map of the 3D scene, and images rendered from viewpoints used during training.
Method
Reconstruction Quality
Generation Quality
Physical Stability
CD-S ↓
CD-O ↓
F1-S ↑
F1-O ↑
IoU-B ↑
ICP-Rot ↓
MMD ↓
COV ↑
P-FID ↓
Uni3D ↑
ULIP ↑
PD ↓
Dmean↓
Epeak↓
Scene Generation
SAM-3D
24.80
1.73
0.2495
0.6250
0.3861
21.27
2.872
70.46
5.245
0.5981
0.6292
1.203
194.09
0.8269
ShapeR
10.54
2.28
0.4955
0.5241
0.4841
19.42
3.676
64.76
16.64
0.4453
0.4831
0.8408
129.81
0.5989
WorldSculpt
12.53
2.11
0.3736
0.5818
0.4593
19.48
3.113
69.95
4.941
0.5382
0.5683
0.2280
135.1
0.6655
Amodal Generation + Pose Estimation
Appendix
Table 4: Quantitative comparison on Toys4K ( Stojanov et al., 2021 ) with estimated depth. We compare Tetris3D with baselines using depth maps estimated by MoGe3 ( Kong et al., 2026 ) . All methods requiring depth use the same depth maps. † denotes methods whose generated objects are aligned to the scene using FoundationPose ( Wen et al., 2024 ) .
Method
Reconstruction Quality
Generation Quality
Physical Stability
CD-S ↓
CD-O ↓
F1-S ↑
F1-O ↑
IoU-B ↑
ICP-Rot ↓
MMD ↓
COV ↑
P-FID ↓
Uni3D ↑
ULIP ↑
PD ↓
Dmean↓
Epeak↓
Ours (GT)
3.68
0.97
0.8407
0.8084
0.7918
5.29
1.501
83.11
1.599
0.6884
0.7182
0.0036
38.91
0.1721
Ours (VLM)
3.76
0.98
0.8399
0.8063
0.7905
5.37
1.538
83.38
1.651
0.6891
0.7195
0.0038
41.25
0.1780
Appendix
Table 5: Quantitative comparison on Toys4K ( Stojanov et al., 2021 ) with VLM-based reasoning. We compare results using ground-truth physical relation annotations with those using relations inferred by a VLM ( Bai et al., 2025 ) . Tetris3D maintains comparable performance with VLM-inferred relations.
Figure 9: Qualitative comparison with agentic 3D generation. We compare Tetris3D with an LLM agent-based approach to 3D mesh generation ( OpenAI, 2026 ) . Although this approach can produce broadly similar shapes, it often fails to faithfully recover the object geometry depicted in the input image and struggles to generate physically consistent scene layouts.
Figure 10: Qualitative comparison with baselines on Toys4K ( Stojanov et al., 2021 ) .
Figure 11: Qualitative comparison with baselines on Toys4K ( Stojanov et al., 2021 ) .
Figure 12: Qualitative comparison with baselines on Toys4K ( Stojanov et al., 2021 ) .