Organizations: State Key Laboratory of Virtual Reality Technology and Systems, Beihang University. · School of Computer Science and Information Engineering, Hefei University of Technology. · Hangzhou Innovation Institute, Beihang University.
Abstract
Occupancy prediction has garnered increasing attention in recent years for its comprehensive fine-grained environmental representation and strong generalization to open-set objects. Nevertheless, mainstream occupancy prediction methods employ cumbersome voxel features as the scene representation, incurring substantial overheads in both memory and computation. When comparing the occupancy distribution in each spatial dimension, we find that the information entropy of the height dimension is much lower than the other two dimensions that constitute the Bird's Eye View (BEV) plane, which indicates that the height distribution of occupancy is easier to learn and predict. Accordingly, we propose Lightweight Spatial Embedding that can represent complete height information in a more compact way than voxel features, thus significantly enhancing its deployability. First, Single-Channel Occupancy is sampled from the multi-view depth distributions, which is then processed by Spatial-to-Channel mechanism to extract Lightweight Spatial Embeddings of different views by 2D convolution. These embeddings will interact with each other through the Lightweight Cross-View Interaction module to obtain the Unified Embedding, which can directly supplement BEV features with height information. Furthermore, we extract Edge-aware Spatial Embedding and apply Geometric Supervision on Spatial Embeddings, aiming to enhance their capability to represent spatial information. We also propose BEV-CutMix, a feature-level data augmentation strategy, to increase the diversity of the driving scenes. We integrate these innovative components into a pure 2D convolutional model, namely LightOcc. Sufficient experimental results show that LightOcc achieves state-of-the-art performance on multiple benchmarks while demonstrating significant efficiency advantages.
Vision-based 3D semantic occupancy prediction is essential for autonomous driving, yet dense voxel representations waste computation on largely empty space, while BEV and TPV projections compromise fine-grained 3D structure. Fully sparse representations offer an attractive alternative, but existing methods, including SparseOcc, entangle scene completion with semantic prediction by indiscriminately propagating high-dimensional features into empty regions and applying voxel-wise classification. This creates excessive activations, computational overhead, and geometric ambiguity. We present SparseOcc++, a geometry-aware sparse framework that explicitly decouples scene completion from semantic segmentation. SparseOcc++ reformulates completion as signed-distance regression on sparse anchor voxels through a scene completion field (SCF). To model complex outdoor geometry robustly, it combines orthogonal decomposition with discretized distance learning. A geometry-guided propagation module then converts the SCF into a complete volumetric scene and restricts semantic segmentation to geometrically verified regions. Experiments establish new state of the art: SparseOcc++ improves IoU by 2.3 points and is 3.9x faster than SparseOcc on nuScenes, while achieving a 5.9x speedup over OccFormer on SemanticKITTI.
In this paper, we explore dense voxel streaming for accurate and efficient 3D occupancy prediction. While dense voxel representations offer fine-grained spatial details and streaming paradigm enables efficient temporal processing, naively combining the two introduces key challenges: (i) warping-induced distortions caused by interpolation used for temporal alignment, and (ii) degraded dynamic object representations due to motion misalignment and detail loss in image-to-voxel projection. To address these, we propose StreamOcc, a novel framework that utilizes two aggregation strategies. Specifically, it first refines propagated voxel features to reduce warping artifacts before temporal accumulation, and then selectively injects instance-level query features encoding dynamic-object semantics into the corresponding occupied voxel regions, preserving temporally consistent modeling while strengthening dynamic object representations. Unlocking effective dense voxel streaming, StreamOcc achieves state-of-the-art performance on SurroundOcc-benchmark and Occ3D-nuScenes under real-time constraints, outperforming the prior best methods by +1.3/2.5 and +1.5/2.0 in (overall/dynamic object) mIoU, respectively, while running at 83.3 ms per frame with only 2.8 GB of memory. The project page is available at https://moonseokha.github.io/StreamOcc/.
3D semantic occupancy prediction requires accurate 2D-to-3D feature lifting, yet current methods restrict camera geometry to initial projections. Subsequent operations like offset learning, attention weighting, and cross-camera aggregation remain geometry-agnostic, ignoring essential physical constraints. We propose VGGT-Occ, a framework that embeds geometric tokens throughout the entire pipeline. We introduce Projection-Aware Deformable Attention (PA-DA) to inject geometry into all attention stages. PA-DA projects 3D offsets back to image planes and leverages the projection Jacobian as an additive bias to suppress unreliable observations. Features are then integrated through a view-quality semantic gate for cross-view consistency. To optimize both efficiency and performance, we employ a sequential coarse-to-fine decoder with gated fusion, where low-resolution features are refined into higher resolutions, allocating computation by information density while substantially reducing decoder cost. Extensive evaluations demonstrate the effectiveness and accuracy of our approach. On SurroundOcc-nuScenes, VGGT-Occ achieves 33.00% IoU and 21.08% mIoU (T=1), and 33.64% IoU and 21.43% mIoU with T=2 inference, outperforming existing methods, with only ∼41M trainable parameters in the occupancy head. Code will be released publicly.