LIVIN: Benchmarking Spatial and Embodied Intelligence in Digital Twins of Lived-In Homes
Organizations: ShanghaiTech University · Deemos Technology
Abstract
Realistic household simulation must capture not only diverse environments but also the lived-in object arrangements and spatial constraints that shape robot motion and interaction. Existing resources often trade off scale, real-world correspondence, and interaction readiness, leaving a gap in faithful, interactive replicas of how real homes are actually arranged. To this end, we introduce LIVIN, a benchmark for spatial and embodied intelligence built on digital twins of 30 diverse lived-in homes. These replicas preserve observed room layouts, furniture configurations, and everyday belongings. To construct them, we design a human-in-the-loop workflow comprising instance recognition, architectural reconstruction, and object generation and placement, with intermediate results reviewed and corrected by humans against the source observations at each stage. We evaluate four tasks in LIVIN: 3D detection, 3D reconstruction, navigation, and loco-manipulation. Our evaluations show that current methods remain challenged by the dense object arrangements, occlusions, limited free space, and constrained interaction regions found in realistic lived-in homes. We hope LIVIN will help advance embodied AI in real-world homes, from spatial understanding to robotic interaction, and ultimately bring embodied intelligence into everyday home environments.
Figures & tables
| Properties | Statistics | User Study | |||||||
| Dataset | Real-World | Inst.-Spec. | Artic. | Avg. Obj. | Obj. Dens. | Artic. Obj. | Plaus. | Realism | Mesh Qual. |
| Grounding | Mesh | Struct. | / Room | (obj./m 2 ) | (%) | (%) | (%) | (%) | |
| 3D-FRONT ( Fu et al., 2021 ) | ✗ | ✗ | ✗ | 6.93 | 0.39 | – | 15.33 | 10.10 | 12.60 |
| ProcTHOR ( Deitke et al., 2022 ) | ✗ | ✗ | ✓ | 18.33 | 0.82 | 11.11 | 5.23 | 8.36 | 2.03 |
| HSSD ( Khanna et al., 2024 ) | ✗ | ✗ | ✓ | 20.17 | 1.58 | 10.08 | 13.24 | 21.26 | 18.70 |
| ReplicaCAD ( Szot et al., 2021 ) | ✓ | ✗ | ✓ | 30.48 | 0.38 | 5.38 | 10.80 | 4.18 | 3.66 |
| 3D Detection | |||||
|---|---|---|---|---|---|
| Method | mAP | [email protected]/0.5 | [email protected]/0.5 | [email protected]/0.5 | 3D [email protected]/0.5 |
| SceneScript ( Avetisyan et al., 2024 ) | – | 0.608 / 0.428 | 0.083 / 0.058 | 0.146 / 0.103 | 0.591 / 0.678 |
| EFM3D ( Straub et al., 2024 ) | 0.071 | 0.617 / 0.362 | 0.096 / 0.056 | 0.166 / 0.097 | 0.541 / 0.657 |
| SpatialLM ( Mao et al., 2025 ) | – | 0.743 / 0.522 | 0.111 / 0.078 | 0.192 / 0.135 | 0.601 / 0.698 |
| Boxer ( DeTone et al., 2026 ) | 0.306 | 0.721 / 0.477 | 0.417 / 0.276 | 0.528 / 0.349 | 0.568 / 0.659 |
| Method | SR (%) | SPL (%) | Collision Count | NE (m) |
|---|---|---|---|---|
| OmniNav ( Xue et al., 2026b ) | 6.67 | 3.66 | 80.59 | 6.69 |
| DualVLN ( Wei et al., 2026a ) | 4.44 | 2.56 | 42.11 | 5.97 |
| GPT-6 Astra ( OpenAI, 2026a ) | 41.48 | 23.59 | 55.41 | 2.16 |
| Method | Grasping and placement | Pushing and pulling | Articulated-object manipulation | Long-horizon tasks | |
|---|---|---|---|---|---|
| SR (%) | SR (%) | SR (%) | SR (%) | PS | |
| DeepSeek-V4.1-Flash ( DeepSeek-AI, 2026 ) | 6.7 | 20.0 | 19.0 | 0.0 | 2.8 |
| GPT-6 Astra ( OpenAI, 2026a ) | 24.4 | 33.3 | 61.9 | 12.5 | 15.9 |
| ( Wei et al., 2026b ) | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Setting |
|---|---|
| Generator Tier | Gen-2.5-High |
| Mesh Mode | Raw |
| Quality | Low |
| Output | GLB |
| Material | PBR |
| Texture Mode | High |
| Skill | Guidance for the Generated Blender Scene |
|---|---|
| Coordination | Assign stage-specific room tasks, preserve source identities and accepted bases, and collect complete room outputs before publication. |
| Structure | Import the initial shell; check cameras and scale; refine walls, openings, recesses, and visible architectural details; reconstruct windows at the reference pose. |
| Structure-Aware Fitting and Articulation | Jointly model object geometry, materials, placement, internal structures, and movable parts in the surrounding room context; record articulation mechanisms, support relations, and observed or inferred parameters. |
| Object Placement | Import assigned meshes, preserve their parts and materials, place supports before supported objects, ensure collision-free placement, and record root transforms and support relations. |
| Phase | Inspection Aspects | Correction and Completion |
|---|---|---|
| Structure | Room boundaries, wall connections, openings, scale, and alignment with source observations. | Select the relevant structural element and provide localized correction feedback. Each room is explicitly confirmed before proceeding. |
| Articulated Objects | Geometry, dimensions, materials, placement, structural fit, and articulation. | Select the relevant object or part for correction, adjust its pose, and resolve inter-object penetration. |
| Non-Articulated Objects | Appearance, geometry, placement, support, and relations to surrounding objects. | Adjust object poses, resolve inter-object penetration, or regenerate defective assets. |
| Method | Output Coverage | 3D Box Recall | |
|---|---|---|---|
| Gen3DSR ( Ardelean et al., 2025 ) | 132/306 (43.14%) | 34.97% | 0.2586 |
| RecGen ( Zadaianchuk et al., 2026 ) | 306/306 (100.00%) | 42.48% | 0.3558 |
| SAM3D ( Chen et al., 2026b ) | 306/306 (100.00%) | 53.59% | 0.3686 |
| Fire3D ( Xia et al., 2026a ) | 292/306 (95.42%) | 72.55% | 0.5774 |
| ShapeR ( Siddiqui et al., 2026 ) | 306/306 (100.00%) | 85.95% | 0.7510 |
| Method | Unrecovered Blockage (%) | Fall (%) | Goal-Completion Mismatch (%) | Other Failures (%) | Failed NE (m) |
|---|---|---|---|---|---|
| GPT-6 Astra ( OpenAI, 2026a ) | 1.27 | 17.72 | 77.22 | 3.80 | 3.15 |
| OmniNav ( Xue et al., 2026b ) | 69.05 | 2.38 | 15.08 | 13.49 | 7.12 |
| DualVLN ( Wei et al., 2026a ) | 14.73 | 0.00 | 71.32 | 13.95 | 6.22 |
| Object grasping and placement (15 tasks) | |
|---|---|
| LM1 | Move a saucer from the placemat to an empty spot on the table. |
| LM2 | Grasp and lift a small green cream jar and keep holding it. |
| LM3 | Grasp and lift a handheld massager and hold it stably. |
| LM4 | Grasp and lift a gray insulated cup and keep holding it. |
| LM5 | Place the wine glass on the table to the right of the placemat. |
| LM6 | Lift a toy elephant with both hands and keep holding it. |