cs.ROSep 24, 2026

OREN-X: Octree Residual Network for Real-Time Multi-Modal Mapping

Authors: Zhirui Dai, Qihao Qian, Dinh Minh Nguyen, Quan-Dung Pham, Kiana Bronder, Carlos Nieto-Granda, Yiyu Chen, Quan Nguyen, +1 more

Organizations: Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, CA 92093, USA · VinMotion · Parsons · U.S. DEVCOM Army Research Laboratory, Adelphi, MD 20783, USA · University of Southern California, Los Angeles, CA 90089, USA

Abstract

To achieve general-purpose autonomy over long horizons, a robot needs to maintain spatial environment information that supports a variety of tasks: geometry for planning and control, radiance for rendering and relocalization, and vision-language features for open-vocabulary grounding. Existing methods represent and estimate each modality separately, multiplying memory and compute cost while forgoing potential synergy among the representations. We develop OREN-X, an online mapping method that uses an octree in 3D space as a shared data structure for indexing and storing a multi-modal field, capturing geometric, radiance, and vision-language information. OREN-X provides efficient unified storage and retrieval of these data in explicit/implicit and full/compressed form. Our unified representation yields cross-modality synergy: SDF estimates are sharpened by occupancy and radiance, while GPU-based ray-octree traversal and octree query enable real-time rendering. We also use online dictionary learning to compress the vision-language features, shrinking them 3.7x below full per-vertex storage while raising the query accuracy. On Replica, OREN-X maps in real time (80+ fps for SDF and 30+ fps for all four modalities), improves near-surface SDF accuracy by 33% over single-modality baselines, and improves mean open-vocabulary 3D mIoU by 71% and mean accuracy by 61% over the best prior method.

Figures & tables

Explore similar work

CardsList
  1. OREN: Octree Residual Network for Real-Time Euclidean Signed Distance Mapping

    Oct 21, 2025Zhirui Dai, Qihao Qian, Tianxing Fan +1Continuous Distance FieldPoint Clouds

  2. VLEM: Real-Time 3D Vision-Language Embedding Mapping

    Aug 8, 2025Christian Rauch, Björn Ellensohn, Linus Nwankwo +2Semantic Scene UnderstandingRobotic Perception

  3. CrossMaps: Confidence-Aware Open-Vocabulary Semantic Mapping for Rover Navigation

    Jun 15, 2026Jan-Niklas Klein, Sona Ghahremani, Christian Medeiros Adriano +1Semantic MappingSimultaneous Localization And Mapping