3D Occupancy Prediction

Latest papers 64

Nov 21, 2025cs.RO

MobileOcc: A Human-Aware Semantic Occupancy Dataset for Mobile Robots

Dense 3D semantic occupancy perception is critical for mobile robots operating in pedestrian-rich environments, yet it remains underexplored compared to its application in autonomous driving. To address this gap, we present MobileOcc, a semantic occupancy dataset for mobile robots operating in crowded human environments. Our dataset is built using an annotation pipeline that incorporates static object occupancy annotations and a novel mesh optimization framework explicitly designed for human occupancy modeling. It reconstructs deformable human geometry from 2D images, then refines and optimizes it using associated LiDAR point data. Using MobileOcc, we establish benchmarks for two tasks: i) Occupancy prediction and ii) Pedestrian velocity prediction, using different methods, including monocular, stereo, and panoptic occupancy, with metrics and baseline implementations for reproducible comparison. Beyond occupancy prediction, we further assess our annotation method on 3D human pose estimation datasets. Results demonstrate that our method exhibits robust performance across different datasets. Our code and dataset are released at https://autonomousrobots.nl/paper_websites/mobileocc
Mar 28, 2025cs.CV

Streaming Dense Voxel Representations for 3D Occupancy Prediction

In this paper, we explore dense voxel streaming for accurate and efficient 3D occupancy prediction. While dense voxel representations offer fine-grained spatial details and streaming paradigm enables efficient temporal processing, naively combining the two introduces key challenges: (i) warping-induced distortions caused by interpolation used for temporal alignment, and (ii) degraded dynamic object representations due to motion misalignment and detail loss in image-to-voxel projection. To address these, we propose StreamOcc, a novel framework that utilizes two aggregation strategies. Specifically, it first refines propagated voxel features to reduce warping artifacts before temporal accumulation, and then selectively injects instance-level query features encoding dynamic-object semantics into the corresponding occupied voxel regions, preserving temporally consistent modeling while strengthening dynamic object representations. Unlocking effective dense voxel streaming, StreamOcc achieves state-of-the-art performance on SurroundOcc-benchmark and Occ3D-nuScenes under real-time constraints, outperforming the prior best methods by +1.3/2.5 and +1.5/2.0 in (overall/dynamic object) mIoU, respectively, while running at 83.3 ms per frame with only 2.8 GB of memory. The project page is available at https://moonseokha.github.io/StreamOcc/.
Dec 8, 2024cs.CV

LightOcc: Lightweight Spatial Embedding for Efficient Vision-based 3D Occupancy Prediction

Occupancy prediction has garnered increasing attention in recent years for its comprehensive fine-grained environmental representation and strong generalization to open-set objects. Nevertheless, mainstream occupancy prediction methods employ cumbersome voxel features as the scene representation, incurring substantial overheads in both memory and computation. When comparing the occupancy distribution in each spatial dimension, we find that the information entropy of the height dimension is much lower than the other two dimensions that constitute the Bird's Eye View (BEV) plane, which indicates that the height distribution of occupancy is easier to learn and predict. Accordingly, we propose Lightweight Spatial Embedding that can represent complete height information in a more compact way than voxel features, thus significantly enhancing its deployability. First, Single-Channel Occupancy is sampled from the multi-view depth distributions, which is then processed by Spatial-to-Channel mechanism to extract Lightweight Spatial Embeddings of different views by 2D convolution. These embeddings will interact with each other through the Lightweight Cross-View Interaction module to obtain the Unified Embedding, which can directly supplement BEV features with height information. Furthermore, we extract Edge-aware Spatial Embedding and apply Geometric Supervision on Spatial Embeddings, aiming to enhance their capability to represent spatial information. We also propose BEV-CutMix, a feature-level data augmentation strategy, to increase the diversity of the driving scenes. We integrate these innovative components into a pure 2D convolutional model, namely LightOcc. Sufficient experimental results show that LightOcc achieves state-of-the-art performance on multiple benchmarks while demonstrating significant efficiency advantages.
Date pendingcs.RO

Diagnosing and Dynamically Filtering Occupancy World Models for Active Mapping

Active mapping requires a robot to select camera viewpoints that efficiently reconstruct an unknown 3D scene. To reason about unobserved regions, recent systems use pretrained occupancy networks as world models that complete missing geometry. The predicted structure contributes to expected coverage gain and constrains feasible robot motion. Consequently, occupancy errors can change both what the robot chooses to explore and where it is able to move. We diagnose these effects by holding the planner fixed and varying only the occupancy representation provided to it. We consider planning without completion, with learned occupancy, with false positives removed by a ground truth oracle, with false negatives restored by an oracle, and with ground truth occupancy. Our experiments show that correcting false positives or false negatives alone does not consistently improve final coverage. This finding reveals a gap between occupancy accuracy and downstream planning performance. Ground truth occupancy provides a much larger improvement in coverage efficiency than in endpoint coverage, suggesting that planning and reachability remain important bottlenecks even when the geometric world model is accurate. Based on these findings, we introduce a dynamic filtering strategy that preserves predictions in unexplored space while suppressing repeatedly unsupported occupancy using online observations. Preliminary examples show that this strategy can redirect viewpoint selection toward reachable surfaces that would otherwise remain unobserved.