AdaOcc: Adaptive 3D Occupancy Prediction for Embodied Tasks
Authors: Jinglong Wang, Yunjie Wang, Zhiyang Zhang, Jiawei He, Ye Yuan, Bo Qiu, Jing Zhang
Organizations: Beihang University · Beijing Academy of Artificial Intelligence · XYZ Embodied AI · Hebei University of Technology · ShanghaiTech University · University of Science and Technology Beijing
Embodied tasks demand accurate, flexible, and semantically rich 3D scene representations. 3D semantic occupancy is well suited to this requirement, as it can model holistic 3D spaces by encoding geometric occupancy along with semantic categories. However, existing occupancy prediction methods struggle to meet practical deployment requirements, such as adapting to varying computing budgets, sensor setups, and observation views. In this paper, we propose a point-based Adaptive 3D Occupancy Prediction method, called AdaOcc, tailored for embodied scenarios. To accommodate heterogeneous sensor inputs, AdaOcc uses an adaptive geometry-guided dual-branch encoder that can support RGB images in various numbers of views with (estimated) depth maps or LiDAR scans. AdaOcc represents occupied regions via sparse semantic points trained with a progressive query learning strategy, allowing the prediction computational budget to be flexibly adjusted through query point numbers and decoder layers. To facilitate high-fidelity geometric modeling for lightweight point-based occupancy learning, we further propose a novel containment loss that regularizes predicted points to reside within valid occupied regions. Extensive experiments show that our method achieves a new state-of-the-art on Occ-ScanNet with considerable performance improvements over previous methods. Moreover, our framework demonstrates strong practical applicability as an adaptive 3D perception module in real-world embodied systems.
Figures & tables
Figure 1: Overview of AdaOcc. AdaOcc targets adaptive 3D occupancy prediction for embodied scene understanding. It accepts heterogeneous visual and geometric inputs from different embodied platforms and adapts to different computational budgets through adjustable query numbers and decoder depths. AdaOcc predicts sparse semantic points, which can be converted into semantic occupancy to provide structured spatial cues for downstream embodied tasks.
Figure 2: Overview of AdaOcc. AdaOcc is an adaptive 3D occupancy prediction framework for embodied scene understanding. It takes RGB observations as the primary input and can optionally incorporate geometric cues from estimated depth maps, depth cameras, or LiDAR scans. AdaOcc also adjusts its prediction budget under different computational constraints. The predicted semantic occupancy provides structured spatial cues for downstream embodied tasks.
Figure 3: Illustration of Containment-Guided Point Optimization (C.P.O.), which regularizes predicted outlier points in supervised empty space to reside inside nearby occupied voxels, reducing surface-floating artifacts.
Dataset
Method
Rep.
mIoU
ceiling
floor
wall
window
chair
bed
sofa
table
tvs
furniture
objects
IoU
Occ-ScanNet
TPVFormer
T
24.94
6.96
32.97
14.41
9.10
24.01
41.49
45.44
28.61
10.66
35.37
25.31
33.39
MonoScene
V
24.62
15.17
44.71
22.41
12.55
26.11
27.03
35.91
28.32
6.57
32.16
19.84
41.60
ISO
V
28.71
19.88
41.88
22.37
16.98
29.09
42.43
42.00
29.60
10.62
36.36
24.61
42.16
SurroundOcc
V
30.83
18.90
49.30
24.80
18.00
26.80
42.00
44.10
32.90
18.60
36.80
26.90
42.52
GaussianFormer
G
29.93
20.70
42.00
23.40
17.40
27.00
44.30
44.80
32.70
15.30
36.70
25.00
40.91
OPUS
P
38.96
39.06
45.04
34.97
28.63
35.92
49.27
54.39
37.93
23.93
45.04
34.42
45.62
Table 1: Local prediction performance on the Occ-ScanNet dataset. Rep. denotes the main scene representation: V for dense voxels, T for TPV features, G for Gaussians, and P for points or point queries. † indicates the AdaOcc variant with an EfficientNet image encoder. Bold and underline indicate the best and second-best results.
Figure 4: Qualitative comparison on Occ-ScanNet. AdaOcc produces cleaner and more complete semantic occupancy compared to the existing approaches.
Figure 5: Qualitative adaptability analysis on TartanGround. AdaOcc accepts geometric inputs from depth maps or LiDAR scans, and processes RGB observations with varying numbers of views jointly in a single forward pass.
Image Encoder
Geometric Initialization
3D Encoder
C.P.O.
Progressive Training
mIoU
IoU
RADIO
–
–
–
–
49.65
56.29
–
–
✓
✓
55.50
62.92
✓
–
✓
✓
57.06
64.72
✓
✓
–
✓
54.54
61.31
✓
✓
✓
–
54.63
62.35
✓
✓
✓
✓
58.74
65.38
Table 2: Ablation study of key components in AdaOcc.
Parameter
Setting
IoU
mIoU
Margin m
0.00 / 0.04 / 0.08
64.25 / 65.38 / 65.32
57.47 / 58.74 / 58.12
C.P.O. weight
0.5× / 1 × / 2×
65.19 / 65.38 / 65.32
58.16 / 58.74 / 58.17
Decay γ
0.80 / 0.90 / 0.95
65.51 / 65.38 / 65.35
58.62 / 58.74 / 58.37
Query schedule
– / 50/20 / 100/40
62.35 / 65.17 / 65.38
54.63 / 58.15 / 58.74
Table 3: Sensitivity to training hyperparameters on Occ-ScanNet-mini. All other settings are fixed within each study, and bold marks the default configuration. For the query schedule, X/Y denotes adding X queries every Y epochs, from X up to 500.
Query
Time ↓ (ms)
Mem. ↓ (MiB)
mIoU
IoU
200
63.2
1710.7
55.13
62.36
500
64.4
1710.6
57.97
65.16
800
73.3
1798.8
58.32
64.58
1000
75.9
1863.7
58.25
64.29
Table 4: Efficiency and adaptability analysis on the Occ-ScanNet-mini dataset using a single NVIDIA H20 GPU. All results use the same EfficientNet image encoder as SplatSSC for a controlled comparison. (a) Query-number ablation for AdaOcc † . (b) Anchor-number ablation for SplatSSC. (c) Decoder-layer output ablation for AdaOcc † . Within each panel only the listed factor is varied while all other settings are fixed.
Figure 6: Real-world embodied applications of AdaOcc. AdaOcc provides semantic 3D occupancy for navigation, manipulation, and mobile manipulation. The semantic prompts are plant for navigation, bottle and basket for manipulation, and trash and basket for mobile manipulation.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Open-vocabulary semantic activation in reconstructed 3D occupancy. Semantically related text queries activate consistent object regions, demonstrating that the learned CLIP-aligned voxel features can respond to free-form language.
Variant
Input Source
IoU (%)
mIoU (%)
GT-depth version
Depth-derived points
67.55
26.13
LiDAR version
LiDAR points
55.65
19.40
Appendix
Table 5: Comparison of the GT-depth and LiDAR input variants on the TartanGround test set.
Input views
Front
Left
Back
Right
Current-view union
Four-view union
Front
62.87 / 23.24
–
–
–
62.87 / 23.24
17.24 / 6.93
Front + Left
65.98 / 25.56
64.43 / 23.95
–
–
65.29 / 25.06
32.34 / 13.29
Front + Left + Back
65.88 / 25.15
64.99 / 24.37
67.24 / 25.26
–
66.14 / 25.04
52.14 / 19.18
Four views
67.57 / 26.55
66.33 / 25.18
68.63 / 26.49
67.29 / 27.03
67.55 / 26.14
67.55 / 26.14
Appendix
Table 6: Occupancy IoU / mIoU on the TartanGround test set under one to four input views, using the model trained with randomly sampled view subsets and evaluated without retraining. “Current-view union” evaluates the union of the regions covered by the views currently provided as input, whereas “Four-view union” uses the fixed union covered by all four cameras.
Methods
RayIoU 1m
RayIoU 2m
RayIoU 4m
RayIoU
RenderOcc ( Pan et al., 2024 )
13.1
19.6
25.5
19.5
BEVFormer ( Li et al., 2024b )
26.1
32.9
38.0
32.4
BEVDet-Occ
23.6
30.0
35.1
29.6
BEVDet-Occ (8f)
26.6
33.1
38.2
32.6
FB-Occ (16f) ( Li et al., 2023b )
26.7
34.1
39.7
33.5
SparseOcc (8f) ( Tang et al., 2024 )
28.0
34.7
39.4
34.0
Appendix
Table 7: Comparison of RayIoU results on the Occ3D-nuScenes dataset.
Study
Setting
IoU
mIoU
Precision
Recall
F1
Margin m
0.00
64.25
57.47
70.50
87.88
78.24
0.04
65.38
58.74
75.32
83.41
79.16
0.08
65.32
58.12
75.36
83.06
79.02
C.P.O. weight
0.05→0.15
65.19
58.16
74.32
84.14
–
0.10→0.30
65.38
58.74
75.32
83.41
–
0.20→0.60
65.32
58.17
75.17
83.29
–
Appendix
Table 8: Effect of the C.P.O. offset margin m and of the C.P.O. loss weight schedule on Occ-ScanNet-mini. Bold marks the default configuration.
Study
Setting
IoU
mIoU
Decay γ
0.80
65.51
58.62
0.90
65.38
58.74
0.95
65.35
58.37
Query schedule
none
62.35
54.63
50/20
65.17
58.15
100/40
65.38
58.74
Appendix
Table 9: Effect of the decoder loss decay factor γ and of the progressive query schedule on Occ-ScanNet-mini. The query schedule is written as queries added per epoch interval; bold marks the default configuration.
Depth input
AbsRel ↓
RMSE (cm) ↓
IoU ↑
mIoU ↑
Ground truth
0.0000
0.00
69.61
62.67
DAv2
0.0441
13.82
65.50
58.35
MoGe-2
0.0782
23.43
62.44
55.76
Appendix
Table 10: Depth quality and the resulting occupancy accuracy on Occ-ScanNet-mini. All three entries are produced in a single controlled run with identical training settings.
Figure 8: The photo of devices used in real-world embodied tasks.
Method
NE ↓
OSR ↑
SR ↑
SPL ↑
FP-Nav
9.40
39.80
28.80
24.00
AdaOcc + value-map stopping
5.38
63.66
43.53
25.04
Appendix
Table 11: Closed-loop navigation on the FloorPlan-R2R val-unseen split. Both methods share the same grounding, planning, and stopping backend; FP-Nav additionally uses privileged floorplan alignment, whereas our pipeline estimates pose and scale online.
Figure 9: Costmap-based navigation from AdaOcc predictions. The semantic PLY prediction is converted into a 2.5D traversability costmap, where free space, inflated obstacles, and unknown regions are represented in different gray levels. Given an open-semantic target activation, A* plans a collision-aware path from the robot start position to the target region and exports map-frame waypoints for the Unitree Go2 robot.
Device Name
Type
Embodied Tasks
Unitree Go2
Quadruped robot
Navigation/Mobile manipulation
ARX X5
Robotic arm
Maniulation/Mobile manipulation
Intel RealSense D455
Long-range stereo depth camera
Navigation
Intel RealSense D405
Short-range depth camera
Mobile manipulation
MRDVS M4
ToF RGB-D sensor
Maniulation
Appendix
Table 12: Equipment Used in Our Real-world Experiments
Seed
mIoU
ceiling
floor
wall
window
chair
bed
sofa
table
tvs
furniture
objects
IoU
S0
58.74
48.73
57.61
56.25
48.18
58.90
75.30
75.07
57.58
46.52
64.37
57.61
65.38
S1
58.26
47.13
57.39
56.12
47.64
58.69
75.19
74.99
57.63
43.60
64.47
58.05
65.24
S2
58.54
48.69
57.54
56.40
48.15
58.91
75.53
75.27
57.93
42.91
64.82
57.81
65.49
S3
58.46
48.03
57.57
56.17
47.94
58.94
75.54
75.03
57.85
43.18
65.19
57.59
65.46
S4
58.29
46.71
57.59
56.51
47.52
58.88
74.71
75.28
57.30
43.93
64.92
57.85
65.46
Mean
58.46
47.86
57.54
56.29
47.89
58.86
75.25
75.13
57.66
44.03
64.75
57.78
65.41
Appendix
Table 13: Standard deviation over five random seeds on Occ-ScanNet-mini. S0–S4 correspond to seeds 0, 1785784487, 1602478221, 1383549716, and 977375669, respectively. All results are reported at epoch 200. Bold and underline indicate the best and second-best results among the five runs.
Data and Input Configuration
Dataset
Occ-ScanNet-mini
Training / validation samples
4,639 / 2,007
RGB observations I
Single-view RGB, current frame only
Geometric observations G
Depth-derived 3D points for benchmark experiments
Image resolution
960×720 after resizing
Occupancy range Ω
[−3.2,−4.8,−5.6,7.2,4.8,5.6]
Appendix
Table 14: Main implementation details for AdaOcc on Occ-ScanNet-mini.
Datasets
RGB Input
LiDAR Input
Open Vocabulary
Depth Map Type
#Views
Occ-ScanNet
✓
Pseudo
1
Occ3D-nuScenes (setting1)
✓
✓
6
Occ3D-nuScenes (setting2)
✓
Pseudo
6
TartanGround (setting1)
✓
✓
✓
1-4
TartanGround (setting2)
✓
✓
GT
1-4
Appendix
Table 15: Experimental Settings on Different Datasets
Item
Value
Hardware
8 × NVIDIA H20 GPUs
GPU memory
97,871 MiB per GPU
Training framework
PyTorch DDP with NCCL backend
CUDA / cuDNN
CUDA 12.1 / cuDNN 8.9.2
PyTorch / TorchVision
2.2.0 / 0.17.0
Training time
8.0 h on 8 GPUs
Appendix
Table 16: Compute resources for the Occ-ScanNet-mini AdaOcc run.
Existing learning-based occupancy prediction methods rely on large-scale 3D annotations and generalize poorly across environments. We present FreeOcc, a training-free framework for open-vocabulary occupancy prediction from monocular or RGB-D sequences. Unlike prior approaches that require voxel-level supervision and ground-truth camera poses, FreeOcc operates without 3D annotations, pose ground truth, or any learning stage. FreeOcc incrementally builds a globally consistent occupancy map via a four-layer pipeline: a SLAM backbone estimates poses and sparse geometry; a geometrically consistent Gaussian update constructs dense 3D Gaussian maps; open-vocabulary semantics from off-the-shelf vision-language models are associated with Gaussian primitives; and a probabilistic Gaussian-to-occupancy projection produces dense voxel occupancy. Despite being entirely training-free and pose-agnostic, FreeOcc achieves over 2× improvements in IoU and mIoU on EmbodiedOcc-ScanNet compared to prior self-supervised methods. We further introduce ReplicaOcc, a benchmark for indoor open-vocabulary occupancy prediction, and show that FreeOcc transfers zero-shot to novel environments, substantially outperforming both supervised and self-supervised baselines. Project page: https://the-masses.github.io/freeocc-web/.
Zeyu Jiang, Changqing Zhou, Xingxing Zuo +1
The Hong Kong University of Science and Technology (Guangzhou) · MBZUAI
3D occupancy prediction is fundamental to scene understanding, yet existing 3D semantic occupancy methods are typically specialized to fixed scene types and occupancy protocols. We introduce Cross-Scene 3D Semantic Occupancy Prediction, a new task setting which requires a single model to handle heterogeneous indoor and outdoor scenes with varying cameras, spatial ranges, voxel specifications, and semantic taxonomies. This setting poses a fundamental challenge: achieving metric-consistent yet scene-adaptive image-to-3D lifting across varying camera configurations and scene scales. To address this challenge, we propose OccAnyScene, a pixel-frustum-centered Gaussian framework built upon a pretrained depth foundation model. Specifically, the framework employs Pixel-Aligned Frustum Feature Aggregation to construct a camera-aware frustum query for each feature pixel, and Frustum-Parameterized Gaussian Construction to decode each query into multiple Gaussians whose positions and sizes are constrained by the predicted pixel depth and corresponding frustum geometry. OccAnyScene sets new state-of-the-art results, achieving 59.92% mIoU on the indoor Occ-ScanNet and 23.06% mIoU on the outdoor SurroundOcc-nuScenes.
3D semantic occupancy prediction requires accurate 2D-to-3D feature lifting, yet current methods restrict camera geometry to initial projections. Subsequent operations like offset learning, attention weighting, and cross-camera aggregation remain geometry-agnostic, ignoring essential physical constraints. We propose VGGT-Occ, a framework that embeds geometric tokens throughout the entire pipeline. We introduce Projection-Aware Deformable Attention (PA-DA) to inject geometry into all attention stages. PA-DA projects 3D offsets back to image planes and leverages the projection Jacobian as an additive bias to suppress unreliable observations. Features are then integrated through a view-quality semantic gate for cross-view consistency. To optimize both efficiency and performance, we employ a sequential coarse-to-fine decoder with gated fusion, where low-resolution features are refined into higher resolutions, allocating computation by information density while substantially reducing decoder cost. Extensive evaluations demonstrate the effectiveness and accuracy of our approach. On SurroundOcc-nuScenes, VGGT-Occ achieves 33.00% IoU and 21.08% mIoU (T=1), and 33.64% IoU and 21.43% mIoU with T=2 inference, outperforming existing methods, with only ∼41M trainable parameters in the occupancy head. Code will be released publicly.
Xun Chen, Tianchen Deng, Rui Wang +5
Nanyang Technological University · Shanghai Jiao Tong University · ETH Zurich