cs.CVSep 29, 2026

HIGS: Hierarchical Implicit Grids for Joint Geometric and Semantic Scene Understanding

Authors: Hanwen Cao, Wenqiang Wu, Kuang-Ting Tu, Mathias Otnes, Jeffrey Delmerico, Rui Wang, Yulun Tian, Nikolay Atanasov

Organizations: University of California San Diego, La Jolla, CA, USA · Southern University of Science and Technology, Shenzhen, Guangzhou, China · Norwegian University of Science and Technology, Trondheim, Norway · Microsoft Spatial AI Lab, Zürich, Switzerland · University of Michigan, Ann Arbor, MI, USA

Abstract

Neural implicit representations have had a significant impact on scene reconstruction by enabling robots to build continuous, differentiable, and high-fidelity 3D maps. Most existing works focus on geometric reconstruction and lack semantic information for high-level spatial understanding and task planning. Also, as the scale and complexity of the environment increase, neural representations face the challenge of maintaining computational efficiency in back-end optimization. To resolve these two challenges, we introduce a hierarchical neural field that leverages multiresolution submaps to achieve an efficient and scalable implicit representation, and a unified query and decoding mechanism to support both geometric and semantic features. More specifically, the learnable map features can be converted to the output with the query and decoding process for both training and inference. For large-scale representation, we decompose a scene into overlapping submaps and do hierarchical optimization within each local submap, thus enabling scalable computation. To further improve efficiency, we design feature encoders that predict initial hierarchical grid features to substantially reduce the time needed to optimize the submap features from scratch. To correct estimation drift among submaps, we align and fuse them entirely within the implicit feature space, leading to substantial acceleration by avoiding the need to decode the final output. Building upon this efficient hierarchical representation, we embed both geometric features and vision-language latent features into the map, and demonstrate it on both Signed Distance Field (SDF) construction and open-vocabulary object grounding. Our approach significantly improves computation and memory efficiency, maintains high estimation accuracy, and endows the robot with spatial awareness on large-scale real-world benchmarks.

Explore similar work

CardsList
  1. Hierarchical Object Representation for Spatial Robot Perception: Points, Meshes, and Superquadrics

    Jun 1, 2026Ceng Zhang, Wan Su, Mohamed Samshad +23D Scene GraphsObject Geometry

  2. IVGT: Implicit Visual Geometry Transformer for Neural Scene Representation

    May 15, 2026Yuqi Wu, Tianyu Hu, Wenzhao Zheng +4Visual Geometry Grounded Transformer3D Geometry

  3. Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning

    May 3, 2026Sixian Zhang, Yiyao Wang, Xinhang Song +3Semantic MappingVision-Language Navigation