OREN-X: Octree Residual Network for Real-Time Multi-Modal Mapping
Authors: Zhirui Dai, Qihao Qian, Dinh Minh Nguyen, Quan-Dung Pham, Kiana Bronder, Carlos Nieto-Granda, Yiyu Chen, Quan Nguyen, +1 more
Organizations: Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, CA 92093, USA · VinMotion · Parsons · U.S. DEVCOM Army Research Laboratory, Adelphi, MD 20783, USA · University of Southern California, Los Angeles, CA 90089, USA
To achieve general-purpose autonomy over long horizons, a robot needs to maintain spatial environment information that supports a variety of tasks: geometry for planning and control, radiance for rendering and relocalization, and vision-language features for open-vocabulary grounding. Existing methods represent and estimate each modality separately, multiplying memory and compute cost while forgoing potential synergy among the representations. We develop OREN-X, an online mapping method that uses an octree in 3D space as a shared data structure for indexing and storing a multi-modal field, capturing geometric, radiance, and vision-language information. OREN-X provides efficient unified storage and retrieval of these data in explicit/implicit and full/compressed form. Our unified representation yields cross-modality synergy: SDF estimates are sharpened by occupancy and radiance, while GPU-based ray-octree traversal and octree query enable real-time rendering. We also use online dictionary learning to compress the vision-language features, shrinking them 3.7x below full per-vertex storage while raising the query accuracy. On Replica, OREN-X maps in real time (80+ fps for SDF and 30+ fps for all four modalities), improves near-surface SDF accuracy by 33% over single-modality baselines, and improves mean open-vocabulary 3D mIoU by 71% and mean accuracy by 61% over the best prior method.
Figures & tables
Fig. 1 : OREN-X at a glance. From RGB-D data, OREN-X builds a single octree-based multi-modal field that jointly represents geometry, radiance and vision-language features (shown here for two backbones, CLIP and TIPS), supporting real-time open-vocabulary queries (e.g. “blanket”) localized directly in 3D.
Fig. 5 : Comparison on radiance and vision-language. OREN-X vs. LangSplatV2, OnlineLangSplat and LatentAM. Left: RGB rendering of Replica room0 . Remaining columns: open-vocabulary query of room1 (“blanket”), room2 (“table”) and TUM 3 (“monitor”).
Metric
nvblox
OREN [ 6 ]
OREN-X (ours)
Accuracy [cm] ↓
1.00
2.50
2.49
Completion [cm] ↓
2.17
2.12
2.14
Chamfer [cm] ↓
1.58
2.31
2.31
F-score [%] ↑
93.27
90.73
90.68
SDF MAE, All [cm] ↓
4.44
2.25
1.92
SDF MAE, Near [cm] ↓
5.47
1.70
1.14
TABLE I : Geometry on Replica (8-scene mean). Mesh distances at a 5 cm threshold, matching OREN’s protocol. Best in bold and shaded green, second yellow.
PSNR ↑
SSIM ↑
LPIPS ↓
Depth L1 (mm) ↓
Method
Replica
TUM
Replica
TUM
Replica
TUM
Replica
TUM
LangSplatV2 [ 30 ]
36.00
19.74
0.96
0.78
0.0683
0.2281
38.50
77.72
OLS [ 35 ]
32.38
15.61
0.91
0.62
0.2359
0.4377
17.52
36.21
LatentAM [ 11 ]
23.77
12.84
0.76
0.49
0.3789
0.6057
6.29
54.98
OREN-X (ours)
28.01
14.21
0.88
0.56
0.2749
0.5038
10.02
17.12
TABLE II : Photometric quality on Replica and TUM (mean over each dataset). Best in bold and shaded green, second yellow.
Method
Metric ↑
Room0
Room1
Room2
Office0
Office1
Office2
Office3
Office4
TUM (mean)
3D
2D
3D
2D
3D
2D
3D
2D
3D
2D
3D
2D
3D
2D
3D
2D
3D
2D
LangSplatV2 1pt
Cos-Sim
0.858
0.918
0.882
0.924
0.865
0.923
0.885
0.917
0.893
0.929
0.864
0.914
0.852
0.904
0.868
0.921
0.783
0.903
mIoU
19.10
28.00
18.40
26.80
18.30
25.00
13.30
23.60
10.00
12.30
21.60
34.70
15.30
26.70
23.40
29.80
9.57
42.10
mAcc
29.80
36.00
32.80
37.70
31.90
40.30
22.60
35.90
24.10
27.60
30.70
42.90
26.00
35.90
41.80
47.20
28.30
61.77
mAP
24.10
32.70
26.80
36.00
25.30
31.40
25.70
31.20
19.20
27.70
26.60
32.70
21.80
29.50
31.70
42.00
16.83
50.13
OLS 1pt
Cos-Sim
0.738
0.945
0.854
0.952
0.768
0.951
0.830
0.939
0.854
0.938
0.765
0.934
0.790
0.939
0.814
0.942
0.864
0.901
TABLE III : Open-vocabulary understanding on Replica and TUM (mean). Best per column shaded green and bold, second yellow.
OREN-X
Baselines
Metric
SDF
SDF+OCC
SDF+Rad
SDF+OCC+Rad
SDF+OCC+Rad+VL
LatentAM
OnlineLangSplat
LangSplatV2
nvblox
FPS ↑
80.06
71.72
60.89
54.07
32.37
34.11
1.79
0.42
73.66
GPU Peak (GiB) ↓
0.28
0.34
0.78
0.87
1.01
8.66
11.32
13.18
1.00
TABLE IV : Efficiency on Replica room0 . Throughput and peak GPU memory of OREN-X as modalities are added, and of the baselines.
Gaussian σ
Axial σ
Variant
No Noise
0.01
0.02
0.05
0.0011
0.0023
0.0057
SDF
15.59
13.67
13.55
20.07
14.66
13.88
24.32
SDF+OCC
15.83
13.78
13.38
19.70
14.90
14.06
23.91
SDF+Rad
12.06
12.04
12.95
18.66
11.87
12.97
24.19
SDF+OCC+Rad
11.40
12.94
13.47
19.23
12.20
13.47
23.81
TABLE V : Depth-noise robustness. Replica room0 . Near-surface SDF MAE (mm) across injected depth-noise levels, for Gaussian noise and axial noise. Best per column in bold , shaded green.
Variant
PSNR ↑
SSIM ↑
LPIPS ↓
FPS ↑
SDF zero crossing
28.35
0.87
0.2978
64.84
Segment midpoint
28.36
0.88
0.3209
66.28
TABLE VI : SDF-induced surface location. Replica room0 . Rendering with the sample placed at the SDF zero crossing vs. the segment midpoint. Best per column in bold , shaded green.
Repr.
PSNR ↑
SSIM ↑
LPIPS ↓
Mem. (MiB) ↓
FPS ↑
RGB
28.32
0.8742
0.2981
242.8
32.75
SH1
28.30
0.8737
0.2953
313.8
32.07
SH2
28.40
0.8743
0.2947
433.5
31.92
SH3
28.46
0.8746
0.2942
608.4
31.45
TABLE VII : RGB vs. spherical-harmonic (SH) color. Replica room0 . Mem. is peak GPU memory per training step, and FPS is training throughput. Best per column in bold , shaded green.
Train/Eval
PSNR ↑
SSIM ↑
LPIPS ↓
Train FPS ↑
Eval FPS ↑
Mem. (MiB) ↓
ray/ray
28.32
0.8742
0.2981
31.80
12.0
341.1
ray/tile
27.68
0.8703
0.3036
1513.8
tile/ray
26.77
0.8751
0.3154
29.25
12.0
262.5
tile/tile
27.00
0.8751
0.3138
1456.7
TABLE VIII : Train/eval rasterizer. Replica room0 . Ray- vs. tile-based rasterizer for training and evaluation. Train FPS and Mem. (peak GPU memory) depend only on the training rasterizer, and Eval FPS is measured at 1200×680 . Best per column in bold .
Variant
Gmax
∣D∣
Cosine ↑
mAP ↑
Mem. ↓
FPS ↑
Full features
–
–
0.907
45.87%
256.2
31.00
Implicit MLP
–
–
0.902
41.59%
9.5
31.06
Offline (uncentered)
256
256
0.905
46.66%
69.3
32.37
Offline (centered)
256
256
0.736
44.54%
69.3
32.37
Online, dense coeff.
64
64
0.845
40.96%
20.5
–
Online, dense coeff.
128
128
0.884
45.38%
36.8
–
TABLE IX : Online dictionary. Replica room0 . VL field stored as full per-vertex features, as implicit features decoded by an MLP, as codes over an offline PCA dictionary (with/without centering), and as codes from our online dictionary, with dense coefficients for varying Gmax or sparse top- K coefficients at Gmax=256 . ∣D∣ is the number of atoms the growth criterion activated. It equals Gmax for the fixed-size offline dictionaries. Dense rows pad each coefficient vector to Gmax . Top- K rows store K coefficients and their indices. Mem. is the storage of the VL field (MB). Best per column in bold , shaded green.
Backbone
Cos-Sim (3D/2D) ↑
mIoU (3D/2D) ↑
mAcc (3D/2D) ↑
mAP (3D/2D) ↑
CLIP ViT-B/16
0.768/ 0.764
31.05 / 33.06
52.84 / 53.84
41.06/ 37.69
DINOv3.txt
0.684/0.694
27.96/30.59
43.72/45.66
39.40/35.25
Talk2DINOv3-ViT-L
0.771 /0.748
33.50 / 35.43
46.87/48.01
42.42 / 40.46
TIPSv2-L
0.884 / 0.881
28.50/30.57
48.83 / 50.02
42.40 /37.54
TABLE X : VL backbone comparison. Replica (8-scene mean). VL reconstruction and open-vocabulary query quality of OREN-X with four VL backbones, each cell 3D/2D. Best per 3D or 2D half in bold , shaded green, second yellow.