OREN-X: Octree Residual Network for Real-Time Multi-Modal Mapping
Authors: Zhirui Dai, Qihao Qian, Dinh Minh Nguyen, Quan-Dung Pham, Kiana Bronder, Carlos Nieto-Granda, Yiyu Chen, Quan Nguyen, +1 more
Organizations: Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, CA 92093, USA · VinMotion · Parsons · U.S. DEVCOM Army Research Laboratory, Adelphi, MD 20783, USA · University of Southern California, Los Angeles, CA 90089, USA
To achieve general-purpose autonomy over long horizons, a robot needs to maintain spatial environment information that supports a variety of tasks: geometry for planning and control, radiance for rendering and relocalization, and vision-language features for open-vocabulary grounding. Existing methods represent and estimate each modality separately, multiplying memory and compute cost while forgoing potential synergy among the representations. We develop OREN-X, an online mapping method that uses an octree in 3D space as a shared data structure for indexing and storing a multi-modal field, capturing geometric, radiance, and vision-language information. OREN-X provides efficient unified storage and retrieval of these data in explicit/implicit and full/compressed form. Our unified representation yields cross-modality synergy: SDF estimates are sharpened by occupancy and radiance, while GPU-based ray-octree traversal and octree query enable real-time rendering. We also use online dictionary learning to compress the vision-language features, shrinking them 3.7x below full per-vertex storage while raising the query accuracy. On Replica, OREN-X maps in real time (80+ fps for SDF and 30+ fps for all four modalities), improves near-surface SDF accuracy by 33% over single-modality baselines, and improves mean open-vocabulary 3D mIoU by 71% and mean accuracy by 61% over the best prior method.
Figures & tables
Fig. 1 : OREN-X at a glance. From RGB-D data, OREN-X builds a single octree-based multi-modal field that jointly represents geometry, radiance and vision-language features (shown here for two backbones, CLIP and TIPS), supporting real-time open-vocabulary queries (e.g. “blanket”) localized directly in 3D.
Fig. 5 : Comparison on radiance and vision-language. OREN-X vs. LangSplatV2, OnlineLangSplat and LatentAM. Left: RGB rendering of Replica room0 . Remaining columns: open-vocabulary query of room1 (“blanket”), room2 (“table”) and TUM 3 (“monitor”).
Metric
nvblox
OREN [ 6 ]
OREN-X (ours)
Accuracy [cm] ↓
1.00
2.50
2.49
Completion [cm] ↓
2.17
2.12
2.14
Chamfer [cm] ↓
1.58
2.31
2.31
F-score [%] ↑
93.27
90.73
90.68
SDF MAE, All [cm] ↓
4.44
2.25
1.92
SDF MAE, Near [cm] ↓
5.47
1.70
1.14
TABLE I : Geometry on Replica (8-scene mean). Mesh distances at a 5 cm threshold, matching OREN’s protocol. Best in bold and shaded green, second yellow.
PSNR ↑
SSIM ↑
LPIPS ↓
Depth L1 (mm) ↓
Method
Replica
TUM
Replica
TUM
Replica
TUM
Replica
TUM
LangSplatV2 [ 30 ]
36.00
19.74
0.96
0.78
0.0683
0.2281
38.50
77.72
OLS [ 35 ]
32.38
15.61
0.91
0.62
0.2359
0.4377
17.52
36.21
LatentAM [ 11 ]
23.77
12.84
0.76
0.49
0.3789
0.6057
6.29
54.98
OREN-X (ours)
28.01
14.21
0.88
0.56
0.2749
0.5038
10.02
17.12
TABLE II : Photometric quality on Replica and TUM (mean over each dataset). Best in bold and shaded green, second yellow.
Method
Metric ↑
Room0
Room1
Room2
Office0
Office1
Office2
Office3
Office4
TUM (mean)
3D
2D
3D
2D
3D
2D
3D
2D
3D
2D
3D
2D
3D
2D
3D
2D
3D
2D
LangSplatV2 1pt
Cos-Sim
0.858
0.918
0.882
0.924
0.865
0.923
0.885
0.917
0.893
0.929
0.864
0.914
0.852
0.904
0.868
0.921
0.783
0.903
mIoU
19.10
28.00
18.40
26.80
18.30
25.00
13.30
23.60
10.00
12.30
21.60
34.70
15.30
26.70
23.40
29.80
9.57
42.10
mAcc
29.80
36.00
32.80
37.70
31.90
40.30
22.60
35.90
24.10
27.60
30.70
42.90
26.00
35.90
41.80
47.20
28.30
61.77
mAP
24.10
32.70
26.80
36.00
25.30
31.40
25.70
31.20
19.20
27.70
26.60
32.70
21.80
29.50
31.70
42.00
16.83
50.13
OLS 1pt
Cos-Sim
0.738
0.945
0.854
0.952
0.768
0.951
0.830
0.939
0.854
0.938
0.765
0.934
0.790
0.939
0.814
0.942
0.864
0.901
TABLE III : Open-vocabulary understanding on Replica and TUM (mean). Best per column shaded green and bold, second yellow.
OREN-X
Baselines
Metric
SDF
SDF+OCC
SDF+Rad
SDF+OCC+Rad
SDF+OCC+Rad+VL
LatentAM
OnlineLangSplat
LangSplatV2
nvblox
FPS ↑
80.06
71.72
60.89
54.07
32.37
34.11
1.79
0.42
73.66
GPU Peak (GiB) ↓
0.28
0.34
0.78
0.87
1.01
8.66
11.32
13.18
1.00
TABLE IV : Efficiency on Replica room0 . Throughput and peak GPU memory of OREN-X as modalities are added, and of the baselines.
Gaussian σ
Axial σ
Variant
No Noise
0.01
0.02
0.05
0.0011
0.0023
0.0057
SDF
15.59
13.67
13.55
20.07
14.66
13.88
24.32
SDF+OCC
15.83
13.78
13.38
19.70
14.90
14.06
23.91
SDF+Rad
12.06
12.04
12.95
18.66
11.87
12.97
24.19
SDF+OCC+Rad
11.40
12.94
13.47
19.23
12.20
13.47
23.81
TABLE V : Depth-noise robustness. Replica room0 . Near-surface SDF MAE (mm) across injected depth-noise levels, for Gaussian noise and axial noise. Best per column in bold , shaded green.
Variant
PSNR ↑
SSIM ↑
LPIPS ↓
FPS ↑
SDF zero crossing
28.35
0.87
0.2978
64.84
Segment midpoint
28.36
0.88
0.3209
66.28
TABLE VI : SDF-induced surface location. Replica room0 . Rendering with the sample placed at the SDF zero crossing vs. the segment midpoint. Best per column in bold , shaded green.
Repr.
PSNR ↑
SSIM ↑
LPIPS ↓
Mem. (MiB) ↓
FPS ↑
RGB
28.32
0.8742
0.2981
242.8
32.75
SH1
28.30
0.8737
0.2953
313.8
32.07
SH2
28.40
0.8743
0.2947
433.5
31.92
SH3
28.46
0.8746
0.2942
608.4
31.45
TABLE VII : RGB vs. spherical-harmonic (SH) color. Replica room0 . Mem. is peak GPU memory per training step, and FPS is training throughput. Best per column in bold , shaded green.
Train/Eval
PSNR ↑
SSIM ↑
LPIPS ↓
Train FPS ↑
Eval FPS ↑
Mem. (MiB) ↓
ray/ray
28.32
0.8742
0.2981
31.80
12.0
341.1
ray/tile
27.68
0.8703
0.3036
1513.8
tile/ray
26.77
0.8751
0.3154
29.25
12.0
262.5
tile/tile
27.00
0.8751
0.3138
1456.7
TABLE VIII : Train/eval rasterizer. Replica room0 . Ray- vs. tile-based rasterizer for training and evaluation. Train FPS and Mem. (peak GPU memory) depend only on the training rasterizer, and Eval FPS is measured at 1200×680 . Best per column in bold .
Variant
Gmax
∣D∣
Cosine ↑
mAP ↑
Mem. ↓
FPS ↑
Full features
–
–
0.907
45.87%
256.2
31.00
Implicit MLP
–
–
0.902
41.59%
9.5
31.06
Offline (uncentered)
256
256
0.905
46.66%
69.3
32.37
Offline (centered)
256
256
0.736
44.54%
69.3
32.37
Online, dense coeff.
64
64
0.845
40.96%
20.5
–
Online, dense coeff.
128
128
0.884
45.38%
36.8
–
TABLE IX : Online dictionary. Replica room0 . VL field stored as full per-vertex features, as implicit features decoded by an MLP, as codes over an offline PCA dictionary (with/without centering), and as codes from our online dictionary, with dense coefficients for varying Gmax or sparse top- K coefficients at Gmax=256 . ∣D∣ is the number of atoms the growth criterion activated. It equals Gmax for the fixed-size offline dictionaries. Dense rows pad each coefficient vector to Gmax . Top- K rows store K coefficients and their indices. Mem. is the storage of the VL field (MB). Best per column in bold , shaded green.
Backbone
Cos-Sim (3D/2D) ↑
mIoU (3D/2D) ↑
mAcc (3D/2D) ↑
mAP (3D/2D) ↑
CLIP ViT-B/16
0.768/ 0.764
31.05 / 33.06
52.84 / 53.84
41.06/ 37.69
DINOv3.txt
0.684/0.694
27.96/30.59
43.72/45.66
39.40/35.25
Talk2DINOv3-ViT-L
0.771 /0.748
33.50 / 35.43
46.87/48.01
42.42 / 40.46
TIPSv2-L
0.884 / 0.881
28.50/30.57
48.83 / 50.02
42.40 /37.54
TABLE X : VL backbone comparison. Replica (8-scene mean). VL reconstruction and open-vocabulary query quality of OREN-X with four VL backbones, each cell 3D/2D. Best per 3D or 2D half in bold , shaded green, second yellow.
Reconstructing signed distance functions (SDFs) from point cloud data benefits many robot autonomy capabilities, including localization, mapping, motion planning, and control. Methods that support online and large-scale SDF reconstruction often rely on discrete volumetric data structures, which affects the continuity and differentiability of the SDF estimates. Neural network methods have demonstrated high-fidelity differentiable SDF reconstruction but they tend to be less efficient, experience catastrophic forgetting and memory limitations in large environments, and are often restricted to truncated SDF. This work proposes OREN, a hybrid method that combines an explicit prior from octree interpolation with an implicit residual from neural network regression. Our method achieves non-truncated (Euclidean) SDF reconstruction with computational and memory efficiency comparable to volumetric methods and differentiability and accuracy comparable to neural network methods. Extensive experiments demonstrate that OREN outperforms the state of the art in terms of accuracy and efficiency, providing a scalable solution for downstream tasks in robotics and computer vision.
Zhirui Dai, Qihao Qian, Tianxing Fan +1
Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, CA 92093, USA
Semantic scene understanding in robotics requires representations that are both metric-accurate and queryable via natural language in real-time. While recent Vision-Language Models enable powerful 2D image-text alignment, their integration into real-time 3D mapping systems remains challenging due to their requirements on ground truth poses, computational cost, and memory constraints. We present VLEM (Vision-Language Embedding Mapping), a real-time framework for integrating pixel-aligned 2D vision-language embeddings from various backends into a globally consistent, metric-accurate 3D representation, requiring only a raw RGB-D stream. Compared to ConceptFusion, Open-Fusion, and RayFronts, VLEM provides better open-set segmentation performance and a more compact representation. We further demonstrate VLEM's versatility in interactive real-time robotic manipulation tasks and mobile mapping scenarios.
Christian Rauch, Björn Ellensohn, Linus Nwankwo +2
Technical University of Leoben, Franz Josef-Straße 18, 8700 Leoben, Austria
Rovers rely on perception to maintain spatial maps that encode both objects and sensor quality (e.g., range reliability, lighting artifacts, data density), guiding data fusion, embedding updates, and navigation under partial observability. To study these coupled perception-navigation processes, we present CrossMaps, a real-time confidence-aware open-vocabulary semantic mapping pipeline that constructs language-queryable maps from RGB-D data. Building on VLMaps-style approaches, CrossMaps integrates multi-scale CLIP embeddings with confidence-aware fusion and a dual-memory architecture consisting of Short-Term Memory (STM) and Long-Term Memory (LTM). The STM aggregates noisy visual observations using geometric, semantic, and temporal confidence cues, while confident and coherent cells are promoted to the LTM as persistent semantic landmarks. Designed for deployment with a Jetson Orin-powered UGV alongside SLAM, CrossMaps runs in real time and produces semantic heatmaps that can be queried with natural language to guide rover navigation.
Jan-Niklas Klein, Sona Ghahremani, Christian Medeiros Adriano +1
Hasso Plattner Institute for Digital Engineering, Potsdam, Germany.