LiteReality-Agent: An Agentic System for Interactable 3D Indoor Scene Reconstruction
Organizations: University of Cambridge · Imperial College London · Independent Researcher
Abstract
We present LiteReality-Agent, an agentic system for reconstructing real indoor environments as realistic, articulated, and simulation-ready 3D scenes from RGB-D scans. At its core, LiteReality-Agent formulates 3D reconstruction as a coding problem, in which a coding agent gathers evidence using specialised tools and iteratively edits a Python script, Room.py, which can be executed to produce a 3D digital twin of the room. With this formulation, we develop a robust observe-edit-verify harness that supports evidence gathering, measurement, verification, layout optimisation, simulation readiness, and quality control throughout the reconstruction process. LiteReality-Agent produces high-quality reconstructions suitable for simulation and downstream embodied AI tasks. Furthermore, as agent capabilities continue to improve rapidly, the system introduced by LiteReality-Agent remains a strong orchestration framework for future agents: it equips them with specialised tools, structured workflows, and robust verification mechanisms that substantially improve reconstruction quality and reliability. We demonstrate that LiteReality-Agent produces reconstructions that are more geometrically accurate, visually realistic, and simulation-compatible than those generated by recent frontier models, such as Astra and Fable. We therefore view LiteReality-Agent as a practical and important building block for robust real-to-sim systems. Both the source code and the data-capture application are publicly available. Code:https://github.com/LiteReality/LiteReality-Agent/
Figures & tables
| Layout | Object | Overall | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | Hum. | A1 | A2 | Hum. | A1 | A2 | Hum. | A1 | A2 |
| LiteReality | 2.03 | 3.00 | 3.00 | 1.95 | 1.70 | 1.70 | 2.11 | 1.80 | 1.80 |
| Fable 5.1 (+Scan) | 2.80 | 3.10 | 3.00 | 2.66 | 2.50 | 2.30 | 2.82 | 3.00 | 2.30 |
| GPT-6 (+Scan) | 3.11 | 4.10 | 3.60 | 2.98 | 3.20 | 2.90 | 3.00 | 3.60 | 2.90 |
| LiteReality-Agent | 4.60 | 5.00 | 4.90 | 4.58 | 4.80 | 4.20 | 4.54 | 4.90 | 4.20 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Capture component | Role in reconstruction |
|---|---|
| RGB frames | Appearance references for objects, materials, and fixtures. |
| Depth and confidence | Geometric evidence and visibility checks. |
| Camera calibration | Intrinsics and poses for projection and matched-view rendering; timestamps associate observations. |
| Derived point cloud | A spatial view of the captured surfaces and missing regions. |
| RoomPlan detections | Room structure, openings, and semantic object boxes with metric dimensions and poses. |