AdaOcc: Adaptive 3D Occupancy Prediction for Embodied Tasks
Authors: Jinglong Wang, Yunjie Wang, Zhiyang Zhang, Jiawei He, Ye Yuan, Bo Qiu, Jing Zhang
Organizations: Beihang University · Beijing Academy of Artificial Intelligence · XYZ Embodied AI · Hebei University of Technology · ShanghaiTech University · University of Science and Technology Beijing
Embodied tasks demand accurate, flexible, and semantically rich 3D scene representations. 3D semantic occupancy is well suited to this requirement, as it can model holistic 3D spaces by encoding geometric occupancy along with semantic categories. However, existing occupancy prediction methods struggle to meet practical deployment requirements, such as adapting to varying computing budgets, sensor setups, and observation views. In this paper, we propose a point-based Adaptive 3D Occupancy Prediction method, called AdaOcc, tailored for embodied scenarios. To accommodate heterogeneous sensor inputs, AdaOcc uses an adaptive geometry-guided dual-branch encoder that can support RGB images in various numbers of views with (estimated) depth maps or LiDAR scans. AdaOcc represents occupied regions via sparse semantic points trained with a progressive query learning strategy, allowing the prediction computational budget to be flexibly adjusted through query point numbers and decoder layers. To facilitate high-fidelity geometric modeling for lightweight point-based occupancy learning, we further propose a novel containment loss that regularizes predicted points to reside within valid occupied regions. Extensive experiments show that our method achieves a new state-of-the-art on Occ-ScanNet with considerable performance improvements over previous methods. Moreover, our framework demonstrates strong practical applicability as an adaptive 3D perception module in real-world embodied systems.
Figures & tables
Figure 1: Overview of AdaOcc. AdaOcc targets adaptive 3D occupancy prediction for embodied scene understanding. It accepts heterogeneous visual and geometric inputs from different embodied platforms and adapts to different computational budgets through adjustable query numbers and decoder depths. AdaOcc predicts sparse semantic points, which can be converted into semantic occupancy to provide structured spatial cues for downstream embodied tasks.
Figure 2: Overview of AdaOcc. AdaOcc is an adaptive 3D occupancy prediction framework for embodied scene understanding. It takes RGB observations as the primary input and can optionally incorporate geometric cues from estimated depth maps, depth cameras, or LiDAR scans. AdaOcc also adjusts its prediction budget under different computational constraints. The predicted semantic occupancy provides structured spatial cues for downstream embodied tasks.
Figure 3: Illustration of Containment-Guided Point Optimization (C.P.O.), which regularizes predicted outlier points in supervised empty space to reside inside nearby occupied voxels, reducing surface-floating artifacts.
Dataset
Method
Rep.
mIoU
ceiling
floor
wall
window
chair
bed
sofa
table
tvs
furniture
objects
IoU
Occ-ScanNet
TPVFormer
T
24.94
6.96
32.97
14.41
9.10
24.01
41.49
45.44
28.61
10.66
35.37
25.31
33.39
MonoScene
V
24.62
15.17
44.71
22.41
12.55
26.11
27.03
35.91
28.32
6.57
32.16
19.84
41.60
ISO
V
28.71
19.88
41.88
22.37
16.98
29.09
42.43
42.00
29.60
10.62
36.36
24.61
42.16
SurroundOcc
V
30.83
18.90
49.30
24.80
18.00
26.80
42.00
44.10
32.90
18.60
36.80
26.90
42.52
GaussianFormer
G
29.93
20.70
42.00
23.40
17.40
27.00
44.30
44.80
32.70
15.30
36.70
25.00
40.91
OPUS
P
38.96
39.06
45.04
34.97
28.63
35.92
49.27
54.39
37.93
23.93
45.04
34.42
45.62
Table 1: Local prediction performance on the Occ-ScanNet dataset. Rep. denotes the main scene representation: V for dense voxels, T for TPV features, G for Gaussians, and P for points or point queries. † indicates the AdaOcc variant with an EfficientNet image encoder. Bold and underline indicate the best and second-best results.
Figure 4: Qualitative comparison on Occ-ScanNet. AdaOcc produces cleaner and more complete semantic occupancy compared to the existing approaches.
Figure 5: Qualitative adaptability analysis on TartanGround. AdaOcc accepts geometric inputs from depth maps or LiDAR scans, and processes RGB observations with varying numbers of views jointly in a single forward pass.
Image Encoder
Geometric Initialization
3D Encoder
C.P.O.
Progressive Training
mIoU
IoU
RADIO
–
–
–
–
49.65
56.29
–
–
✓
✓
55.50
62.92
✓
–
✓
✓
57.06
64.72
✓
✓
–
✓
54.54
61.31
✓
✓
✓
–
54.63
62.35
✓
✓
✓
✓
58.74
65.38
Table 2: Ablation study of key components in AdaOcc.
Parameter
Setting
IoU
mIoU
Margin m
0.00 / 0.04 / 0.08
64.25 / 65.38 / 65.32
57.47 / 58.74 / 58.12
C.P.O. weight
0.5× / 1 × / 2×
65.19 / 65.38 / 65.32
58.16 / 58.74 / 58.17
Decay γ
0.80 / 0.90 / 0.95
65.51 / 65.38 / 65.35
58.62 / 58.74 / 58.37
Query schedule
– / 50/20 / 100/40
62.35 / 65.17 / 65.38
54.63 / 58.15 / 58.74
Table 3: Sensitivity to training hyperparameters on Occ-ScanNet-mini. All other settings are fixed within each study, and bold marks the default configuration. For the query schedule, X/Y denotes adding X queries every Y epochs, from X up to 500.
Query
Time ↓ (ms)
Mem. ↓ (MiB)
mIoU
IoU
200
63.2
1710.7
55.13
62.36
500
64.4
1710.6
57.97
65.16
800
73.3
1798.8
58.32
64.58
1000
75.9
1863.7
58.25
64.29
Table 4: Efficiency and adaptability analysis on the Occ-ScanNet-mini dataset using a single NVIDIA H20 GPU. All results use the same EfficientNet image encoder as SplatSSC for a controlled comparison. (a) Query-number ablation for AdaOcc † . (b) Anchor-number ablation for SplatSSC. (c) Decoder-layer output ablation for AdaOcc † . Within each panel only the listed factor is varied while all other settings are fixed.
Figure 6: Real-world embodied applications of AdaOcc. AdaOcc provides semantic 3D occupancy for navigation, manipulation, and mobile manipulation. The semantic prompts are plant for navigation, bottle and basket for manipulation, and trash and basket for mobile manipulation.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Open-vocabulary semantic activation in reconstructed 3D occupancy. Semantically related text queries activate consistent object regions, demonstrating that the learned CLIP-aligned voxel features can respond to free-form language.
Variant
Input Source
IoU (%)
mIoU (%)
GT-depth version
Depth-derived points
67.55
26.13
LiDAR version
LiDAR points
55.65
19.40
Appendix
Table 5: Comparison of the GT-depth and LiDAR input variants on the TartanGround test set.
Input views
Front
Left
Back
Right
Current-view union
Four-view union
Front
62.87 / 23.24
–
–
–
62.87 / 23.24
17.24 / 6.93
Front + Left
65.98 / 25.56
64.43 / 23.95
–
–
65.29 / 25.06
32.34 / 13.29
Front + Left + Back
65.88 / 25.15
64.99 / 24.37
67.24 / 25.26
–
66.14 / 25.04
52.14 / 19.18
Four views
67.57 / 26.55
66.33 / 25.18
68.63 / 26.49
67.29 / 27.03
67.55 / 26.14
67.55 / 26.14
Appendix
Table 6: Occupancy IoU / mIoU on the TartanGround test set under one to four input views, using the model trained with randomly sampled view subsets and evaluated without retraining. “Current-view union” evaluates the union of the regions covered by the views currently provided as input, whereas “Four-view union” uses the fixed union covered by all four cameras.
Methods
RayIoU 1m
RayIoU 2m
RayIoU 4m
RayIoU
RenderOcc ( Pan et al., 2024 )
13.1
19.6
25.5
19.5
BEVFormer ( Li et al., 2024b )
26.1
32.9
38.0
32.4
BEVDet-Occ
23.6
30.0
35.1
29.6
BEVDet-Occ (8f)
26.6
33.1
38.2
32.6
FB-Occ (16f) ( Li et al., 2023b )
26.7
34.1
39.7
33.5
SparseOcc (8f) ( Tang et al., 2024 )
28.0
34.7
39.4
34.0
Appendix
Table 7: Comparison of RayIoU results on the Occ3D-nuScenes dataset.
Study
Setting
IoU
mIoU
Precision
Recall
F1
Margin m
0.00
64.25
57.47
70.50
87.88
78.24
0.04
65.38
58.74
75.32
83.41
79.16
0.08
65.32
58.12
75.36
83.06
79.02
C.P.O. weight
0.05→0.15
65.19
58.16
74.32
84.14
–
0.10→0.30
65.38
58.74
75.32
83.41
–
0.20→0.60
65.32
58.17
75.17
83.29
–
Appendix
Table 8: Effect of the C.P.O. offset margin m and of the C.P.O. loss weight schedule on Occ-ScanNet-mini. Bold marks the default configuration.
Study
Setting
IoU
mIoU
Decay γ
0.80
65.51
58.62
0.90
65.38
58.74
0.95
65.35
58.37
Query schedule
none
62.35
54.63
50/20
65.17
58.15
100/40
65.38
58.74
Appendix
Table 9: Effect of the decoder loss decay factor γ and of the progressive query schedule on Occ-ScanNet-mini. The query schedule is written as queries added per epoch interval; bold marks the default configuration.
Depth input
AbsRel ↓
RMSE (cm) ↓
IoU ↑
mIoU ↑
Ground truth
0.0000
0.00
69.61
62.67
DAv2
0.0441
13.82
65.50
58.35
MoGe-2
0.0782
23.43
62.44
55.76
Appendix
Table 10: Depth quality and the resulting occupancy accuracy on Occ-ScanNet-mini. All three entries are produced in a single controlled run with identical training settings.
Figure 8: The photo of devices used in real-world embodied tasks.
Method
NE ↓
OSR ↑
SR ↑
SPL ↑
FP-Nav
9.40
39.80
28.80
24.00
AdaOcc + value-map stopping
5.38
63.66
43.53
25.04
Appendix
Table 11: Closed-loop navigation on the FloorPlan-R2R val-unseen split. Both methods share the same grounding, planning, and stopping backend; FP-Nav additionally uses privileged floorplan alignment, whereas our pipeline estimates pose and scale online.
Figure 9: Costmap-based navigation from AdaOcc predictions. The semantic PLY prediction is converted into a 2.5D traversability costmap, where free space, inflated obstacles, and unknown regions are represented in different gray levels. Given an open-semantic target activation, A* plans a collision-aware path from the robot start position to the target region and exports map-frame waypoints for the Unitree Go2 robot.
Device Name
Type
Embodied Tasks
Unitree Go2
Quadruped robot
Navigation/Mobile manipulation
ARX X5
Robotic arm
Maniulation/Mobile manipulation
Intel RealSense D455
Long-range stereo depth camera
Navigation
Intel RealSense D405
Short-range depth camera
Mobile manipulation
MRDVS M4
ToF RGB-D sensor
Maniulation
Appendix
Table 12: Equipment Used in Our Real-world Experiments
Seed
mIoU
ceiling
floor
wall
window
chair
bed
sofa
table
tvs
furniture
objects
IoU
S0
58.74
48.73
57.61
56.25
48.18
58.90
75.30
75.07
57.58
46.52
64.37
57.61
65.38
S1
58.26
47.13
57.39
56.12
47.64
58.69
75.19
74.99
57.63
43.60
64.47
58.05
65.24
S2
58.54
48.69
57.54
56.40
48.15
58.91
75.53
75.27
57.93
42.91
64.82
57.81
65.49
S3
58.46
48.03
57.57
56.17
47.94
58.94
75.54
75.03
57.85
43.18
65.19
57.59
65.46
S4
58.29
46.71
57.59
56.51
47.52
58.88
74.71
75.28
57.30
43.93
64.92
57.85
65.46
Mean
58.46
47.86
57.54
56.29
47.89
58.86
75.25
75.13
57.66
44.03
64.75
57.78
65.41
Appendix
Table 13: Standard deviation over five random seeds on Occ-ScanNet-mini. S0–S4 correspond to seeds 0, 1785784487, 1602478221, 1383549716, and 977375669, respectively. All results are reported at epoch 200. Bold and underline indicate the best and second-best results among the five runs.
Data and Input Configuration
Dataset
Occ-ScanNet-mini
Training / validation samples
4,639 / 2,007
RGB observations I
Single-view RGB, current frame only
Geometric observations G
Depth-derived 3D points for benchmark experiments
Image resolution
960×720 after resizing
Occupancy range Ω
[−3.2,−4.8,−5.6,7.2,4.8,5.6]
Appendix
Table 14: Main implementation details for AdaOcc on Occ-ScanNet-mini.
Datasets
RGB Input
LiDAR Input
Open Vocabulary
Depth Map Type
#Views
Occ-ScanNet
✓
Pseudo
1
Occ3D-nuScenes (setting1)
✓
✓
6
Occ3D-nuScenes (setting2)
✓
Pseudo
6
TartanGround (setting1)
✓
✓
✓
1-4
TartanGround (setting2)
✓
✓
GT
1-4
Appendix
Table 15: Experimental Settings on Different Datasets
Item
Value
Hardware
8 × NVIDIA H20 GPUs
GPU memory
97,871 MiB per GPU
Training framework
PyTorch DDP with NCCL backend
CUDA / cuDNN
CUDA 12.1 / cuDNN 8.9.2
PyTorch / TorchVision
2.2.0 / 0.17.0
Training time
8.0 h on 8 GPUs
Appendix
Table 16: Compute resources for the Occ-ScanNet-mini AdaOcc run.