Reconstructing simulation-ready 3D scenes from real-world observations enables robotics, gaming, and immersive applications, yet existing methods largely assume rigid objects. This leaves an important gap for deformables, whose simulation-ready geometry depends on dimensionality (curves, surfaces, or volumes) and whose behavior may require models beyond elasticity. We present CoDimRecon, an agentic framework that reconstructs editable scenes containing rigid, articulated, and deformable objects from multi-view RGB observations. Scene-level geometric priors ground scale and layout, while object-level generated meshes guide the agent toward detailed, compact geometry; articulated rigid objects are decomposed into movable parts with explicit joints. For deformables, category-wise agent sessions reconstruct curves as centerlines with radii, surfaces as manifold shells with thickness, and volumes as watertight solids for volumetric meshing. Reusable simulator skills initialize compatible physical models and parameters, while agent-guided behavioral tests expose mismatches and trigger targeted revisions of motion, geometry, numerics, or material modeling. On evaluated Replica and ScanNet++ scenes, CoDimRecon achieves competitive compositional reconstruction accuracy while additionally producing deformable assets for rod, shell, and solid simulation. We further demonstrate robot interactions across all three representations, including a controlled paper-folding case in which behavioral testing motivates plastic bending.
Figures & tables
Figure 1: CoDimRecon reconstructs a simulation-ready scene from multi-view RGB images. Left: input views of an office. Center: top-down reconstruction with an articulated drawer and deformable examples—a coiled telephone cord (curve), plastic bag (surface), and chair cushion (volume). The bin and the bag are separate objects, rigid and deformable, respectively. Right: robot interactions with the reconstructed cord, plastic bag, and cushion.
Method
Input
Geometric Guidance
Generative Prior
Auto Instance Discovery
Training- Free
Agentic
Deformable Simulation
HoloScene
RGB, Mask, Cam
✓
✓
✗
✗
✗
✗
SimRecon
RGB
✓
✓
✓
✓
✗
✗
ReplicateAnyScene
RGB
✓
✓
✓
✓
✗
✗
VIGA
Single-view RGB
✗
✓
✓
✓
✓
✗
Lucida
RGB
✓
✓
✓
✗
✓
✗
Lumera
Single-view RGB
✓
✓
✓
✗
✓
✗
Table 1: Comparison of compositional scene-reconstruction methods. Cam: camera parameters; RGBD: RGB images with depth.
Figure 2: Overview. (a) Geometric context and generated meshes guide editable primitive-based scene reconstruction. (b) Render–evaluate–refine improves appearance and pose; articulated rigid objects receive joints, and rigid bodies are settled under gravity. (c) Separate sessions reconstruct solver-compatible curves, surfaces, and volumes; reusable skills initialize physical models, and agent-guided robot tests diagnose motion, geometry, or numerical issues before material changes.
Object
Model
ρ
E
ν
r/h
κY
δ
μ
(kg/m3)
(MPa)
(mm)
(m−1)
(mm)
Telephone cord
Discrete elastic rod
1600
40
0.35
1.3
–
0.15
0.60
Paper
StVK membrane, plastic hinges
761.9
2250
0.15
0.105
25
0.50
0.40
Chair cushion
StVK–Hencky solid
100
0.03
0.20
–
–
0.50
0.40
Beanbag
Stable Neo-Hookean solid
120
0.02
0.30
–
–
0.50
0.40
Table 2: Deformable simulation parameters. ρ : density; E : Young’s modulus (paper: membrane and bending; cord: stretching, bending, and twisting); ν : Poisson’s ratio; r/h : cord radius or shell thickness; κY : plastic-hinge yield curvature; δ : contact activation distance; μ : friction coefficient. Dashes denote inapplicable entries. The cord’s stress-free natural shape is the reconstructed coil, so the rod keeps its coiled shape without load. Values are effective simulation parameters, not measured material properties; paper plasticity and the cushion’s StVK–Hencky model were selected through behavioral testing. Appendix B.4 gives discretization, boundary conditions, and tasks.
Datasets
Method
Geometry
Rendering
Latent Similarity
Rigid Stability
CD ↓
F1 ↑
NC ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Layout ↑
Motion ↑
Stable
Stable
(Ground) %↑
(All) %↑
Replica
HoloScene
6.79
53.48
78.76
21.22
0.6881
0.4171
94.46
95.71
77.77
49.41
ReplicateAnyScene
41.88
18.74
61.82
11.17
0.5529
0.6461
76.84
80.88
96.97
69.47
GPT-6 Astra
11.79
52.91
73.12
14.07
0.5671
0.5440
97.24
95.57
92.27
87.97
Our Method
6.63
52.24
79.76
15.06
0.5695
0.4817
95.14
96.90
100.00
97.54
Table 3: Compositional scene-reconstruction results. CoDimRecon produces the most physically stable reconstructions, with the lowest CD and highest NC on both datasets, while keeping input-view rendering second only to HoloScene in PSNR and SSIM. Dark and light blue mark the best and second-best results; ties share the same color.
Figure 3: Qualitative compositional scene-reconstruction comparison. We compare rendered appearance (top) and geometry (bottom) on ScanNet++ scene 67d702f2e8 . Our method reconstructs complete, compact object geometry at faithful scale, from the rear door and wall shelf to the bookcase contents and swivel chair. By contrast, HoloScene’s surfaces are fragmented, ReplicateAnyScene omits the doors and poster, and GPT-6 Astra enlarges the rear door. Our rendering is clean, without the blur and artifacts around the chair and bookcase in HoloScene’s rendering.
Figure 7Figure 8
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Initial reconstruction with geometric and generative references. VGGT-Omega supplies cameras, depth, and point maps, while CropFormer masks are clustered into 3D tracks. SAM3D generates and aligns a reference mesh from a selected view of each track; the agent then rebuilds the object with editable Blender primitives.
Figure 9: Rest-shape settling. Reconstructed beanbag at t=0 and 6 s. (a) Vertex distance to equilibrium, normalized by bounding-box diagonal D=1.75 m; equilibrium is averaged over t∈[5.5,6] s. (b) Finite-difference kinetic energy. Maximum distance remains below 10−4D after 2.3 s.
Asset
Observed failure
Attributed to
Revision
Material changed
Telephone cord
No collision-free path in the first two plans
Motion
Re-planned trajectory
No
Cushion, indenter study
Element inversion and excess strain under large compression
Material model
Hencky logarithmic-strain elasticity
Yes
Cushion, press study
Element inversion and excess strain during the press
Motion
Revised press motion
No
Chair cushion
Back fragments fell under gravity; the solve did not converge
Geometry, boundary
Volume re-classified; full seat underside bonded
No
Chair cushion
Excess stretch and shallow indentation at the third point
Tool
70 mm rounded tool
No
Chair cushion
Excess stretch at the third point
Contact location
Third point moved to the cushion front
No
Appendix
Table 6: Revisions during agent-guided behavioral testing. Rows are chronological within each asset. Earlier cushion studies supply the Hencky model and E=30 kPa used in the final cushion; setup failures, repeats, interrupted runs, and rendering are omitted.
Figure 10: ReplicateAnyScene reimplementation on the official hallway example. (a) Six sampled input frames. (b) Reconstruction with default VLM + SAM3 segmentation; the room shell is hidden to expose individual objects.
Figure 11: Qualitative comparison on ScanNet++ scene 7831862f02 .
Figure 12: Qualitative comparison on ScanNet++ scene acd69a1746 .
Figure 13: Qualitative comparison on Replica room_0 .
Figure 14: Qualitative comparison on Replica room_1 .
Figure 15: Qualitative comparison on Replica room_2 .
Reconstructing interactive, simulation-ready 3D scenes from a single image is a critical bottleneck for robotic manipulation. While recent single-image lifters recover plausible per-object shapes, composing them yields scenes that collapse under physical simulation due to interpenetrating, hovering, or sinking objects. Existing physics-aware methods address this strictly as a post-hoc layout correction, leaving the underlying geometric errors unresolved. To address this, we introduce SimuScene, a compositional 3D reconstruction pipeline that puts physics in the loop of shape and layout estimation. Rather than using physics merely for layout cleanup, we utilize the physics engine as a diagnostic measurement tool during the generative process itself. By diagnostically simulating reconstructed objects under gravity, we convert penetration and support failures into quantitative correction signals that drive gravity-axis stretching and amodal shape resampling. This physics-informed feedback loop mitigates accumulated reconstruction errors and produces a stable, simulation-ready compositional 3D scene. Extensive experiments demonstrate state-of-the-art performance on physical stability and geometric alignment benchmarks. We further highlight SimuScene's utility by deploying reconstructed environments in humanoid control and robot-arm manipulation tasks.
This work addresses the problem of recovering complete, simulatable object geometry from reconstructed real-world scenes, enabling physics-based interaction with objects embedded in the scene. While modern multi-view reconstruction methods can produce visually accurate environments, objects are often incomplete due to occlusions and limited observations, making them unsuitable for physics simulation. To address this limitation, we propose SAM3D-Phys, a framework that integrates scene reconstruction with generative 3D priors of SAM3D to recover physically simulatable objects. Our approach first reconstructs the scene from multi-view images to obtain scene geometry and partial observations of objects. We then leverage SAM3D to infer complete object geometry from these partial observations. To ensure that the recovered objects remain consistent with the reconstructed scene, we restore scene-consistent object states through two complementary strategies: a physics-constrained spatial optimization algorithm that iteratively aligns the recovered object to its original location, and a mask-guided appearance distillation module that refines texture fidelity based on the observed images. By recovering complete object geometry and restoring its pose and appearance within the scene, SAM3D-Phys produces clean object representations suitable for physics-based simulation, enabling simultaneous and physically consistent interactive simulation of multiple objects within a reconstructed scene. Project page: https://chnxindong.github.io/sam3d-phys/
Xin Dong, Weijian Deng, Lihan Zhang +3
Shenzhen International Graduate School, Tsinghua University · Pengcheng Laboratory
Converting multi-view RGB observations into simulation-ready 3D environments remains challenging because current reconstruction pipelines produce monolithic scene representations without explicit physical structure. They are typically defined up to an arbitrary global rotation and entangle rigid foreground objects with background geometry, which hinders stable physical interaction. Existing solutions often recover interactivity by replacing reconstructed objects with retrieved CAD assets, but this introduces a slow retrieval-and-replacement stage and weakens scene-specific geometric fidelity. We propose GARDEN, an RGB-only framework that reformulates reconstruction as physically-grounded scene factorization and outputs a structured hybrid scene representation. The key idea is to use gravity as a universal physical prior: we first align the reconstruction to a unified Gravity-View frame to resolve gauge ambiguity, then recover object-centric rigid meshes with accurate 6-DoF placement, and finally remove duplicate object geometry from the background through conditional 3D point classification. The resulting representation combines explicit rigid bodies with a decoupled background, enabling direct physics simulation while preserving visual realism. Experiments on both simulated and real multi-view scenes show that GARDEN improves object placement reliability, disentanglement quality, and rendering-simulation efficiency compared with retrieval-based baselines. Project page: https://sunjiahaovo.github.io/garden/