Reconstructing simulation-ready 3D scenes from real-world observations enables robotics, gaming, and immersive applications, yet existing methods largely assume rigid objects. This leaves an important gap for deformables, whose simulation-ready geometry depends on dimensionality (curves, surfaces, or volumes) and whose behavior may require models beyond elasticity. We present CoDimRecon, an agentic framework that reconstructs editable scenes containing rigid, articulated, and deformable objects from multi-view RGB observations. Scene-level geometric priors ground scale and layout, while object-level generated meshes guide the agent toward detailed, compact geometry; articulated rigid objects are decomposed into movable parts with explicit joints. For deformables, category-wise agent sessions reconstruct curves as centerlines with radii, surfaces as manifold shells with thickness, and volumes as watertight solids for volumetric meshing. Reusable simulator skills initialize compatible physical models and parameters, while agent-guided behavioral tests expose mismatches and trigger targeted revisions of motion, geometry, numerics, or material modeling. On evaluated Replica and ScanNet++ scenes, CoDimRecon achieves competitive compositional reconstruction accuracy while additionally producing deformable assets for rod, shell, and solid simulation. We further demonstrate robot interactions across all three representations, including a controlled paper-folding case in which behavioral testing motivates plastic bending.
Figures & tables
Figure 1: CoDimRecon reconstructs a simulation-ready scene from multi-view RGB images. Left: input views of an office. Center: top-down reconstruction with an articulated drawer and deformable examples—a coiled telephone cord (curve), plastic bag (surface), and chair cushion (volume). The bin and the bag are separate objects, rigid and deformable, respectively. Right: robot interactions with the reconstructed cord, plastic bag, and cushion.
Method
Input
Geometric Guidance
Generative Prior
Auto Instance Discovery
Training- Free
Agentic
Deformable Simulation
HoloScene
RGB, Mask, Cam
✓
✓
✗
✗
✗
✗
SimRecon
RGB
✓
✓
✓
✓
✗
✗
ReplicateAnyScene
RGB
✓
✓
✓
✓
✗
✗
VIGA
Single-view RGB
✗
✓
✓
✓
✓
✗
Lucida
RGB
✓
✓
✓
✗
✓
✗
Lumera
Single-view RGB
✓
✓
✓
✗
✓
✗
Table 1: Comparison of compositional scene-reconstruction methods. Cam: camera parameters; RGBD: RGB images with depth.
Figure 2: Overview. (a) Geometric context and generated meshes guide editable primitive-based scene reconstruction. (b) Render–evaluate–refine improves appearance and pose; articulated rigid objects receive joints, and rigid bodies are settled under gravity. (c) Separate sessions reconstruct solver-compatible curves, surfaces, and volumes; reusable skills initialize physical models, and agent-guided robot tests diagnose motion, geometry, or numerical issues before material changes.
Object
Model
ρ
E
ν
r/h
κY
δ
μ
(kg/m3)
(MPa)
(mm)
(m−1)
(mm)
Telephone cord
Discrete elastic rod
1600
40
0.35
1.3
–
0.15
0.60
Paper
StVK membrane, plastic hinges
761.9
2250
0.15
0.105
25
0.50
0.40
Chair cushion
StVK–Hencky solid
100
0.03
0.20
–
–
0.50
0.40
Beanbag
Stable Neo-Hookean solid
120
0.02
0.30
–
–
0.50
0.40
Table 2: Deformable simulation parameters. ρ : density; E : Young’s modulus (paper: membrane and bending; cord: stretching, bending, and twisting); ν : Poisson’s ratio; r/h : cord radius or shell thickness; κY : plastic-hinge yield curvature; δ : contact activation distance; μ : friction coefficient. Dashes denote inapplicable entries. The cord’s stress-free natural shape is the reconstructed coil, so the rod keeps its coiled shape without load. Values are effective simulation parameters, not measured material properties; paper plasticity and the cushion’s StVK–Hencky model were selected through behavioral testing. Appendix B.4 gives discretization, boundary conditions, and tasks.
Datasets
Method
Geometry
Rendering
Latent Similarity
Rigid Stability
CD ↓
F1 ↑
NC ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Layout ↑
Motion ↑
Stable
Stable
(Ground) %↑
(All) %↑
Replica
HoloScene
6.79
53.48
78.76
21.22
0.6881
0.4171
94.46
95.71
77.77
49.41
ReplicateAnyScene
41.88
18.74
61.82
11.17
0.5529
0.6461
76.84
80.88
96.97
69.47
GPT-6 Astra
11.79
52.91
73.12
14.07
0.5671
0.5440
97.24
95.57
92.27
87.97
Our Method
6.63
52.24
79.76
15.06
0.5695
0.4817
95.14
96.90
100.00
97.54
Table 3: Compositional scene-reconstruction results. CoDimRecon produces the most physically stable reconstructions, with the lowest CD and highest NC on both datasets, while keeping input-view rendering second only to HoloScene in PSNR and SSIM. Dark and light blue mark the best and second-best results; ties share the same color.
Figure 3: Qualitative compositional scene-reconstruction comparison. We compare rendered appearance (top) and geometry (bottom) on ScanNet++ scene 67d702f2e8 . Our method reconstructs complete, compact object geometry at faithful scale, from the rear door and wall shelf to the bookcase contents and swivel chair. By contrast, HoloScene’s surfaces are fragmented, ReplicateAnyScene omits the doors and poster, and GPT-6 Astra enlarges the rear door. Our rendering is clean, without the blur and artifacts around the chair and bookcase in HoloScene’s rendering.
Figure 7Figure 8
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Initial reconstruction with geometric and generative references. VGGT-Omega supplies cameras, depth, and point maps, while CropFormer masks are clustered into 3D tracks. SAM3D generates and aligns a reference mesh from a selected view of each track; the agent then rebuilds the object with editable Blender primitives.
Figure 9: Rest-shape settling. Reconstructed beanbag at t=0 and 6 s. (a) Vertex distance to equilibrium, normalized by bounding-box diagonal D=1.75 m; equilibrium is averaged over t∈[5.5,6] s. (b) Finite-difference kinetic energy. Maximum distance remains below 10−4D after 2.3 s.
Asset
Observed failure
Attributed to
Revision
Material changed
Telephone cord
No collision-free path in the first two plans
Motion
Re-planned trajectory
No
Cushion, indenter study
Element inversion and excess strain under large compression
Material model
Hencky logarithmic-strain elasticity
Yes
Cushion, press study
Element inversion and excess strain during the press
Motion
Revised press motion
No
Chair cushion
Back fragments fell under gravity; the solve did not converge
Geometry, boundary
Volume re-classified; full seat underside bonded
No
Chair cushion
Excess stretch and shallow indentation at the third point
Tool
70 mm rounded tool
No
Chair cushion
Excess stretch at the third point
Contact location
Third point moved to the cushion front
No
Appendix
Table 6: Revisions during agent-guided behavioral testing. Rows are chronological within each asset. Earlier cushion studies supply the Hencky model and E=30 kPa used in the final cushion; setup failures, repeats, interrupted runs, and rendering are omitted.
Figure 10: ReplicateAnyScene reimplementation on the official hallway example. (a) Six sampled input frames. (b) Reconstruction with default VLM + SAM3 segmentation; the room shell is hidden to expose individual objects.
Figure 11: Qualitative comparison on ScanNet++ scene 7831862f02 .
Figure 12: Qualitative comparison on ScanNet++ scene acd69a1746 .
Figure 13: Qualitative comparison on Replica room_0 .
Figure 14: Qualitative comparison on Replica room_1 .
Figure 15: Qualitative comparison on Replica room_2 .