Realistic household simulation must capture not only diverse environments but also the lived-in object arrangements and spatial constraints that shape robot motion and interaction. Existing resources often trade off scale, real-world correspondence, and interaction readiness, leaving a gap in faithful, interactive replicas of how real homes are actually arranged. To this end, we introduce LIVIN, a benchmark for spatial and embodied intelligence built on digital twins of 30 diverse lived-in homes. These replicas preserve observed room layouts, furniture configurations, and everyday belongings. To construct them, we design a human-in-the-loop workflow comprising instance recognition, architectural reconstruction, and object generation and placement, with intermediate results reviewed and corrected by humans against the source observations at each stage. We evaluate four tasks in LIVIN: 3D detection, 3D reconstruction, navigation, and loco-manipulation. Our evaluations show that current methods remain challenged by the dense object arrangements, occlusions, limited free space, and constrained interaction regions found in realistic lived-in homes. We hope LIVIN will help advance embodied AI in real-world homes, from spatial understanding to robotic interaction, and ultimately bring embodied intelligence into everyday home environments.
Figures & tables
Figure 1: Overview of LIVIN. Built from 30 lived-in homes, LIVIN captures realistic household layouts, furnishings, and daily-use objects, enabling evaluation of 3D detection, reconstruction, navigation, and loco-manipulation in realistic domestic environments.
Properties
Statistics
User Study
Dataset
Real-World
Inst.-Spec.
Artic.
Avg. Obj.
Obj. Dens.
Artic. Obj.
Plaus. ↑
Realism ↑
Mesh Qual. ↑
Grounding
Mesh
Struct.
/ Room
(obj./m 2 )
(%)
(%)
(%)
(%)
3D-FRONT ( Fu et al., 2021 )
✗
✗
✗
6.93
0.39
–
15.33
10.10
12.60
ProcTHOR ( Deitke et al., 2022 )
✗
✗
✓
18.33
0.82
11.11
5.23
8.36
2.03
HSSD ( Khanna et al., 2024 )
✗
✗
✓
20.17
1.58
10.08
13.24
21.26
18.70
ReplicaCAD ( Szot et al., 2021 )
✓
✗
✓
30.48
0.38
5.38
10.80
4.18
3.66
Table 1: Comparison with existing indoor scene datasets. Inst.-Spec. Mesh denotes instance-specific object meshes, and Artic. Struct. denotes modeled articulation. Obj. Dens. denotes object density (obj./m 2 ), and Artic. Obj. denotes the percentage of articulated objects. Plaus., Realism, and Mesh Qual. denote human preference rates for physical plausibility, scene realism, and mesh quality.
Figure 2: Qualitative results for 3D detection and reconstruction. Top: 3D detection predictions from SceneScript, EFM3D, SpatialLM, and Boxer compared with ground truth, where green boxes denote correct detections, yellow boxes indicate missed objects, and red boxes denote incorrect predictions. Bottom: 3D reconstructions from ShapeR, Gen3DSR, SAM3D, RecGen, and Fire3D compared with ground-truth.
Table 2: Quantitative evaluation of 3D detection and reconstruction on LIVIN . F-Score-S uses a surface-distance threshold of 0.05m , and F1 uses a 3D box-matching threshold of 0.25 .
Figure 3: Humanoid navigation in LIVIN . LIVIN evaluates humanoid navigation in faithful replicas of real lived-in homes under naturally occurring spatial constraints. From left to right, examples show the robot turning into a room, moving sideways through a narrow passage, and stepping over a doorway threshold.
Method
SR (%) ↑
SPL (%) ↑
Collision Count ↓
NE (m) ↓
OmniNav ( Xue et al., 2026b )
6.67
3.66
80.59
6.69
DualVLN ( Wei et al., 2026a )
4.44
2.56
42.11
5.97
GPT-6 Astra ( OpenAI, 2026a )
41.48
23.59
55.41
2.16
Table 3: Quantitative comparison of navigation methods on LIVIN . SR denotes success rate and SPL denotes success weighted by path length; both are reported in percent. Collision Count is the mean number of robot–environment collision events, and NE denotes the mean geodesic distance from the robot’s final position to the goal in meters.
Figure 4: Loco-manipulation tasks. Examples of object grasping and placement, object pushing and pulling, articulated-object manipulation, and long-horizon tasks in household environments.
Method
Grasping and placement
Pushing and pulling
Articulated-object manipulation
Long-horizon tasks
SR (%) ↑
SR (%) ↑
SR (%) ↑
SR (%) ↑
PS ↑
DeepSeek-V4.1-Flash ( DeepSeek-AI, 2026 )
6.7
20.0
19.0
0.0
2.8
GPT-6 Astra ( OpenAI, 2026a )
24.4
33.3
61.9
12.5
15.9
Ψ0 ( Wei et al., 2026b )
0.0
0.0
0.0
0.0
0.0
Table 4: Loco-manipulation performance across task families on LIVIN . SR denotes end-to-end success rate in percent and is reported for all four task families. PS denotes the mean process score across trials (0–100) and is reported only for long-horizon tasks. Higher is better.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Gallery. Top-down view and close-up view.
Figure 6: Articulated objects in LIVIN . Each quartet shows original appearance and motion frames. Colors identify moving assemblies consistently within each object; gray denotes stationary structure.
Figure 7: Statistics of the LIVIN scene collection. We report (a) the distribution of room types and (b) the distribution of articulated objects across room types.
Figure 8: Input overview. Each home is represented by a 2D floorplan, a metric scan, and posed multi-view images.
Parameter
Setting
Generator Tier
Gen-2.5-High
Mesh Mode
Raw
Quality
Low
Output
GLB
Material
PBR
Texture Mode
High
Appendix
Table 5: Rodin parameters used for object generation.
Skill
Guidance for the Generated Blender Scene
Coordination
Assign stage-specific room tasks, preserve source identities and accepted bases, and collect complete room outputs before publication.
Structure
Import the initial shell; check cameras and scale; refine walls, openings, recesses, and visible architectural details; reconstruct windows at the reference pose.
Structure-Aware Fitting and Articulation
Jointly model object geometry, materials, placement, internal structures, and movable parts in the surrounding room context; record articulation mechanisms, support relations, and observed or inferred parameters.
Object Placement
Import assigned meshes, preserve their parts and materials, place supports before supported objects, ensure collision-free placement, and record root transforms and support relations.
Appendix
Table 6: Responsibilities of the construction skills. Each skill specifies task instructions and required deliverables for scene construction.
Figure 9: Construction ablation. Qualitative comparison of Code-Only, Single-Round Refinement, Iterative Refinement, No Refinement, and our human-guided construction strategy.
Phase
Inspection Aspects
Correction and Completion
Structure
Room boundaries, wall connections, openings, scale, and alignment with source observations.
Select the relevant structural element and provide localized correction feedback. Each room is explicitly confirmed before proceeding.
Articulated Objects
Geometry, dimensions, materials, placement, structural fit, and articulation.
Select the relevant object or part for correction, adjust its pose, and resolve inter-object penetration.
Non-Articulated Objects
Appearance, geometry, placement, support, and relations to surrounding objects.
Adjust object poses, resolve inter-object penetration, or regenerate defective assets.
Appendix
Table 7: 3D reconstruction review process. Each phase is reviewed for its corresponding reconstruction aspects, followed by targeted correction and completion.
Figure 10: Snapshot of the human review interface. The reconstructed scene is shown on the left, the corresponding posed reference image on the right, and an additional panel on the far right displays object material information for inspection.
Figure 11: Selected small clutter objects from the LIVIN test set. Top: full-scene views. Bottom: corresponding close-ups. Highlighted objects belong to the common evaluation subset used by all methods.
Method
Output Coverage ↑
3D Box Recall ↑
Fobj↑
Gen3DSR ( Ardelean et al., 2025 )
132/306 (43.14%)
34.97%
0.2586
RecGen ( Zadaianchuk et al., 2026 )
306/306 (100.00%)
42.48%
0.3558
SAM3D ( Chen et al., 2026b )
306/306 (100.00%)
53.59%
0.3686
Fire3D ( Xia et al., 2026a )
292/306 (95.42%)
72.55%
0.5774
ShapeR ( Siddiqui et al., 2026 )
306/306 (100.00%)
85.95%
0.7510
Appendix
Table 8: Output coverage and reconstruction accuracy on the clutter-like proxy subset of LIVIN . Output coverage reports the number and percentage of generated objects, while 3D box recall and Fobj evaluate object-level reconstruction accuracy.
Method
Unrecovered Blockage (%)
Fall (%)
Goal-Completion Mismatch (%)
Other Failures (%)
Failed NE (m) ↓
GPT-6 Astra ( OpenAI, 2026a )
1.27
17.72
77.22
3.80
3.15
OmniNav ( Xue et al., 2026b )
69.05
2.38
15.08
13.49
7.12
DualVLN ( Wei et al., 2026a )
14.73
0.00
71.32
13.95
6.22
Appendix
Table 9: Failure case analysis of navigation methods on LIVIN . Failure categories are reported as percentages of failed episodes for each method. Failed NE denotes the mean navigation error over failed episodes.
Figure 12: Navigation trajectories with GR00T and the mixed execution setting. Left: GPT-6 Astra+GR00T fails to pass the narrow passage on the right side of the dining room. Right: GPT-6 Astra+Mixed uses CAT for local traversal, allowing the robot to pass through the constrained right-side passage and continue toward the goal.
Object grasping and placement (15 tasks)
LM1
Move a saucer from the placemat to an empty spot on the table.
LM2
Grasp and lift a small green cream jar and keep holding it.
LM3
Grasp and lift a handheld massager and hold it stably.
LM4
Grasp and lift a gray insulated cup and keep holding it.
LM5
Place the wine glass on the table to the right of the placemat.
LM6
Lift a toy elephant with both hands and keep holding it.
Appendix
Table 10: Loco-manipulation tasks. Instructions are condensed from the task definitions.
Indoor scene generation is crucial for robot simulation and modern interior design. However, complex layouts together with scarce 3D scene data make learning-based generation challenging. Existing methods often rely on hand-crafted rules or focus on isolated sub-tasks (e.g., floorplan synthesis or single-room furnishing), producing whole-home scenes that lack global coherence, realism, and simulation readiness. To mitigate these limitations, we propose a unified hierarchical framework that decomposes indoor scene synthesis into controllable stages. First, we curate a large-scale dataset of 300K real residential floorplans to train a large language model for whole-home floorplan generation. With detailed descriptions and a K-D tree-based representation, our method enables fine-grained, controllable whole-home floorplan generation. Building upon the generated whole-home floorplan, we leverage image generation models to draft furniture layouts from multi-level roaming viewpoints, and then generate the layouts of small manipulable objects on different supporting surfaces (e.g., cabinets, desks, and dining tables) for embodied AI simulation. During furniture and object layout generation, a VLM-based refiner iteratively corrects furniture and object placement, and a 3D generative model enables flexible replacement of individual assets. We further attach basic physical attributes and simple surface texture and lighting setups to complete the pipeline for embodied AI use. Experiments and user studies demonstrate that our pipeline produces indoor spaces with greater layout diversity and stronger 3D design appeal, outperforming prior methods on both quantitative and qualitative metrics. Finally, alongside our generation pipeline, we will release the floorplan dataset and 5K fully furnished scenes to the community. Project Page: https://kairos-homeworld.github.io/
Wenbo Li, Xiaoliang Ju, Zipeng Qin +2
Ace Robotics · CUHK MMLab · Shenzhen Loop Area Institute
Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3DBench, a robot-centric benchmark targeting low-level spatial intelligence in embodied 3D environments. To systematically evaluate these foundational perceptual capabilities, the benchmark includes 6 task categories divided into two core groups: Spatial Structural Understanding (Grounding, Spatial Relation Prediction, and Multi-view Correspondence) and Interaction-Oriented Perception (Affordance Prediction, Grasp Point Prediction, and Trajectory Prediction). The benchmark spans 12 subcategories and contains over 21k high-quality question-answer pairs. We evaluate 13 state-of-the-art models, and the results show that while current models exhibit relatively strong high-level spatial reasoning, such as understanding object-to-object positional relations, they remain fragile in interaction-oriented perception, highlighting a significant lack of robust 3D-aware interaction priors. To actively bridge this capability gap revealed by our benchmark, we further synthesize a large-scale training dataset comprising 1.3M QA pairs. Notably, fine-tuning on this dataset yields significant improvements in low-level spatial intelligence. Ultimately, Embodied3DBench fills a critical gap by providing both a systematic evaluation framework and a scalable data solution, setting a clear target for the development of interaction-aware multimodal systems.
We present LiteReality-Agent, an agentic system for reconstructing real indoor environments as realistic, articulated, and simulation-ready 3D scenes from RGB-D scans. At its core, LiteReality-Agent formulates 3D reconstruction as a coding problem, in which a coding agent gathers evidence using specialised tools and iteratively edits a Python script, Room.py, which can be executed to produce a 3D digital twin of the room. With this formulation, we develop a robust observe-edit-verify harness that supports evidence gathering, measurement, verification, layout optimisation, simulation readiness, and quality control throughout the reconstruction process. LiteReality-Agent produces high-quality reconstructions suitable for simulation and downstream embodied AI tasks. Furthermore, as agent capabilities continue to improve rapidly, the system introduced by LiteReality-Agent remains a strong orchestration framework for future agents: it equips them with specialised tools, structured workflows, and robust verification mechanisms that substantially improve reconstruction quality and reliability. We demonstrate that LiteReality-Agent produces reconstructions that are more geometrically accurate, visually realistic, and simulation-compatible than those generated by recent frontier models, such as Astra and Fable. We therefore view LiteReality-Agent as a practical and important building block for robust real-to-sim systems. Both the source code and the data-capture application are publicly available. Code:https://github.com/LiteReality/LiteReality-Agent/
Zhening Huang, Yueyan Li, Johnathan Chiu +5
University of Cambridge · Imperial College London · Independent Researcher