OpenSplatGraph: From Dense Semantic Maps to Structured Scene Graphs for Open-Vocabulary Robot Perception
Authors: Binh Long Nguyen, Kien Nguyen, Clinton Fookes, Peyman Moghadam
Organizations: School of Electrical Engineering and Robotics, Queensland University of Technology (QUT), Brisbane, QLD 4000, Australia · CSIRO Robotics, CSIRO, Brisbane, QLD 4069, Australia
Dense 3D mapping with semantic understanding is essential for robotic perception in complex environments. Recent 3D Gaussian Splatting-based mapping approaches enable high-fidelity geometry and efficient open-vocabulary perception, but typically represent semantics as unstructured feature fields that limit object-centric reasoning. In contrast, 3D scene graphs explicitly model objects and their relationships for structured reasoning, but are commonly constructed from sparse geometric representations that do not fully exploit dense semantic maps. In this work, we present OpenSplatGraph, a unified framework that constructs persistent 3D scene graphs directly from an online Gaussian-based open-vocabulary semantic map. The proposed framework augments the dense semantic map with a reliability-aware semantic field that maintains lightweight observation statistics for confidence-aware, query-conditioned object extraction. Extracted object instances are associated with persistent graph nodes, allowing object attributes and relationships to be incrementally updated across observations and queries. By tightly coupling dense semantic mapping with persistent object-centric representations, our framework supports both language-guided object grounding and structured relational reasoning while preserving the geometric fidelity of Gaussian-based mapping. Comprehensive evaluations on standard 3D scene understanding benchmarks and real-world robotic experiments demonstrate that OpenSplatGraph achieves competitive performance for online open-vocabulary perception and downstream robotic tasks. Project page: https://csiro-robotics.github.io/OpenSplatGraph.
Figures & tables
Figure 1 : Comparison of complementary representations for open-vocabulary 3D scene understanding. Dense semantic maps provide high-fidelity geometry and language-aligned features but represent semantics as unstructured feature fields. Scene graphs explicitly model objects and their relationships for structured reasoning, but typically rely on sparse geometric representations. Our framework bridges these paradigms by constructing a persistent 3D scene graph directly from an online dense semantic map.
Figure 2 : Overview of OpenSplatGraph . Starting from a multi-view RGB-D sequence, the framework incrementally constructs an online open-vocabulary dense semantic map together with a reliability-aware semantic field for robust query-conditioned object extraction. Extracted object instances are then associated across queries and incrementally accumulated into a persistent 3D scene graph. By jointly maintaining the dense semantic field and the scene graph, OpenSplatGraph enables structured relational reasoning while preserving high-fidelity geometric and semantic information for downstream robotic tasks.
Mode
Method
ScanNet
Replica
ScanNet20
NYUv2-40
ScanNet20
NYUv2-40
mIoU ↑
mAcc ↑
mIoU ↑
mAcc ↑
FPS ↑
mIoU ↑
mAcc ↑
mIoU ↑
mAcc ↑
FPS ↑
Offline
LangSplat [ 23 ]
3.78
9.11
3.46
8.43
–
3.84
8.75
3.56
8.26
–
OpenGaussian [ 33 ]
30.08
45.59
27.53
42.17
–
25.63
36.24
23.45
31.59
–
InstanceGaussian [ 17 ]
40.32
55.78
36.88
52.91
–
32.24
45.99
29.88
41.01
–
Online
ConceptFusion [ 11 ]
12.80
22.62
11.72
20.93
0.52
13.05
29.89
11.51
27.66
0.49
Table 1 : Quantitative comparison of 3D semantic segmentation on ScanNet and Replica datasets. (Best and second-best results are highlighted in bold and underlined .)
Figure 3 : Qualitative results of 3D semantic segmentation on ScanNet and Replica datasets.
Query Type
CLIP Retrieval
LLM Retrieval
ConceptGraphs
OpenSplatGraph
ConceptGraphs
OpenSplatGraph
R@1
R@2
R@3
R@1
R@2
R@3
R@1
R@2
R@3
R@1
R@2
R@3
Descriptive
0.59
0.82
0.86
0.45
0.82
0.92
0.61
0.64
0.64
0.50
0.59
0.69
Affordance
0.43
0.57
0.63
0.54
0.61
0.79
0.57
0.63
0.66
0.61
0.71
0.75
Negation
0.26
0.60
0.71
0.51
0.64
0.82
0.80
0.89
0.97
0.80
0.89
0.94
Table 2 : Quantitative comparison of 3D object grounding on Replica across different query types and retrieval strategies. (ConceptGraphs results are reported from [ 7 ] .)
Figure 4 : Qualitative visualization of open-vocabulary 3D scene graph construction on the Replica dataset. For visualization, asymmetric relations are displayed from nodes with lower instance IDs to those with higher IDs.
Figure 5 : Real-world robotic navigation experiments with OpenSplatGraph . The framework constructs an open-vocabulary semantic map online and supports language-guided navigation in indoor environments. For each query, we visualize the rendered target object from the reference view and from an additional arbitrary viewpoint.
Configuration
mIoU ↑
mAcc ↑
R@1 ↑
Baseline
30.97
42.39
–
+ Reliability-aware semantic modeling
36.06
50.74
–
+ Persistent object association
36.06
50.74
0.36
+ Scene graph reasoning (Full)
36.06
50.74
0.50
Table 3 : Progressive ablation of OpenSplatGraph components.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Method
ScanNet
Replica
mIoU ↑
mAcc ↑
mIoU ↑
mAcc ↑
GaussianGraph [ 30 ]
31.09
48.91
31.18
49.14
DynamicGSG [ 6 ]
–
–
31.06 †
54.04 †
OpenSplatGraph (Ours)
36.75
53.82
35.37
47.67
Appendix
Table 4 : Additional open-vocabulary 3D semantic segmentation results on ScanNet and Replica. † DynamicGSG uses a different evaluation mapping.
Method
Node Prec. ↑
Edge Prec. ↑
Dup. Rate (%) ↓
ConceptGraphs [ 7 ]
0.71
0.88
6.77
OpenSplatGraph (Ours)
0.73
0.91
5.94
Appendix
Table 5 : Direct scene-graph evaluation against ConceptGraphs.
Scene graphs are becoming a standard representation for robot navigation, providing hierarchical geometric and semantic scene understanding. However, most scene graph mapping methods rely on depth cameras or LiDAR sensors. In this work, we present LEXI-SG, the first dense monocular visual mapping system for open-vocabulary 3D scene graphs using only RGB camera input. Our approach exploits the semantic priors of open-vocabulary foundation models to partition the scene into rooms, deferring feed-forward reconstruction to when each room is fully observed -- enabling scalable dense mapping without sliding-window scale inconsistencies. We propose a room-based factor graph formulation to globally align room reconstructions while preserving local map consistency and naturally imposing the semantic scene graph hierarchy. Within each room, we further support open-vocabulary object segmentation and tracking. We validate LEXI-SG on indoor scenes from the Habitat-Matterport 3D and self-collected egocentric office sequences. We evaluate its performance against existing feed-forward SLAM methods, as well as established scene graphs baselines. We demonstrate improved trajectory estimation and dense reconstruction, as well as, competitive performance in open-vocabulary segmentation. LEXI-SG shows that accurate, scalable, open-vocabulary 3D scene graphs can be achieved from monocular RGB alone. Our project page and office sequences are available here: https://ori-drs.github.io/lexisg-web/.
Christina Kassab, Hyeonjae Gil, Matías Mattamala +2
Department of Engineering Science, University of Oxford, UK · Department of Mechanical Engineering, Seoul National University, South Korea · School of Informatics, University of Edinburgh, UK
Open-vocabulary 3D scene graph methods typically operate in two stages: first reconstruct, then enrich with vision-language models, leaving the graph unqueryable during exploration. We argue that this sequential coupling is unnecessary and propose an asynchronous architecture in which lightweight online mapping runs concurrently with heavyweight semantic refinement. A probabilistic voxel-based backbone maintains stable object identities incrementally, while background VLM agents progressively enrich the graph. This framework resolves duplicate object tracks through semantic loop closure, attaches fine-grained visual attributes and derives spatial relations between objects. A multi-target frame scheduler amortizes VLM cost by selecting a small set of informative frames that jointly cover multiple targets. The resulting scene graph is queryable during exploration and grows in semantic richness over time. Our method matches or outperforms existing open-vocabulary 3D scene graph methods on semantic segmentation (ScanNet, Replica) and surpasses the prior state-of-the-art across three visual grounding benchmarks (Sr3D+, Nr3D, ScanRefer) by 15.3 to 18.8 A@0.25. Project page: https://denizbickici.github.io/thinkgraphs/
Deniz Bickici, Michael Pabst, Shohei Mori +1
University of Stuttgart, Germany · IMPRS-IS, Germany · Graz University of Technology, Austria
Integrating open-vocabulary semantic information into dynamic 3D scene representations is essential for long-term embodied scene understanding. However, existing methods often suffer from fragile instance association due to incomplete cross-view cues, while their limited ability to handle object-level topological changes restricts long-term robotic task execution. Moreover, current 3D scene understanding methods either rely on simple feature matching without explicit spatial reasoning or assume offline ground-truth 3D geometry. To address these challenges, we present DGSG-Mind, a hybrid instance-aware 3D Gaussian dynamic scene graph system with an embodied reasoning agent. Our system couples a probabilistic voxel grid with explicit 3D Gaussians to enable robust cross-modal instance fusion and incremental semantic mapping. It handles dynamic changes through Gaussian-based visual relocalization and localized masked refinement guided by geometric-semantic consistency. Built on the instance Gaussian map, DGSG-Mind further constructs a hierarchical scene graph and develops the 3D Gaussian Mind, which integrates structural relations, spatial-semantic information, and visually annotated RoI Gaussian renderings for multimodal reasoning. Extensive experiments show that DGSG-Mind achieves the best zero-shot 3DVG performance among methods operating on self-reconstructed maps, while also delivering strong performance in 3D open-vocabulary semantic segmentation and scene reconstruction. We further deploy DGSG-Mind on real-world robots to demonstrate its target-oriented reasoning and dynamic update capabilities. The project page of DGSG-Mind is available at https://icr-lab.github.io/DGSG-Mind
Luzhou Ge, Xiangyu Zhu, Jinyan Liu +1
School of Computer Science, Beijing Institute of Technology, Beijing 100081, China.