cs.CVSep 25, 2026

TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking

Authors: Peder Borge Hellesylt, Albert Gassol Puigjaner, Kostas Alexis, Annette Stahl

Organizations: Norwegian University of Science and Technology (NTNU), Trondheim, Norway

Abstract

Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TRACKGRAPH, an online open-vocabulary system that maintains short-term 2D mask identity directly in the image stream before fusing segments into 3D. FastSAM masks and CLIP features are computed at sparse keyframes, while dense DINOv3 features are used to propagate masks at a high rate in between. The resulting tracked masks are fused into a class-agnostic 3D segment layer within a hierarchical scene graph, with 3D association handling tracking interruptions and long-term revisits. Compact multi-view CLIP embeddings enable open-vocabulary retrieval. Across Replica, ScanNet++, and HM3D, TRACKGRAPH achieves competitive open-vocabulary segmentation and retrieval against state-of-the-art mapping methods, including the highest synonym frequency on Replica (0.50). On the same NVIDIA A100, it is 1.7x faster and uses 3.3x less GPU memory than ViT-H OVI-MAP. Real-world quadruped deployments demonstrate onboard scene graph construction and object search at 7.5Hz, while recorded drone data is used to test the method under aerial viewpoints.

Explore similar work

CardsList
  1. Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs

    Jun 30, 2026Deniz Bickici, Michael Pabst, Shohei Mori +1Open-Vocabulary 3D Segmentation3D SGG

  2. A Scene Language Model for Open-Vocabulary Scene Mapping

    Sep 18, 2026Adam Lilja, Fabio Hübel, Siming He +63D Scene Representation3D VLMs

  3. Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis

    Jul 6, 2026Xianhao Chen, Jiarui Hu, Yuanbo Yang +5Graph Attention NetworksOpen-Vocabulary 3D Segmentation