-Occ: Resilient 3D Semantic Occupancy Prediction for Autonomous Driving with Incomplete Camera Inputs
Abstract
Semantic occupancy prediction enables dense 3D geometric and semantic understanding for autonomous driving. However, existing camera-based approaches implicitly assume complete surround-view observations, an assumption that rarely holds in real-world deployment due to occlusion, hardware malfunction, or communication failures. We study semantic occupancy prediction under incomplete multi-camera inputs and introduce -Occ, a framework designed to preserve geometric structure and semantic coherence when views are missing. -Occ addresses two complementary challenges. First, a Multi-view Masked Reconstruction (MMR) module leverages the spatial overlap among neighboring cameras to recover missing-view representations directly in the feature space. Second, a Feature Memory Module (FMM) introduces a learnable memory bank that stores class-level semantic prototypes. By retrieving and integrating these global priors, the FMM refines ambiguous voxel features, ensuring semantic consistency even when observational evidence is incomplete. We introduce a systematic missing-view evaluation protocol on the nuScenes-based SurroundOcc benchmark, encompassing both deterministic single-view failures and stochastic multi-view dropout scenarios. Under the safety-critical missing back-view setting, -Occ improves the IoU by 4.36%. As the number of missing cameras increases, the robustness gap further widens; for instance, under the setting with five missing views, our method boosts the IoU by 6.67%. These gains are achieved without compromising full-view performance. The source code will be publicly released at https://github.com/qixi7up/M2-Occ.
Figures & tables
| Setting | Method | IoU | mIoU | barrier | bicycle | bus | car | const.veh. | motorcycle | pedestrian | traffic cone | trailer | truck | drive.surf. | other flat | sidewalk | terrain | manmade | vegetation |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Standard | SurroundOcc (Wei et al. 2023 [ 7 ] ) | 31.18 | 17.23 | 16.16 | 8.15 | 19.07 | 29.28 | 10.22 | 9.11 | 10.59 | 8.19 | 11.09 | 18.54 | 38.04 | 19.20 | 23.35 | 21.57 | 12.68 | 20.44 |
| Ours (SurroundOcc) | 31.17 | 18.03 | 17.56 | 7.65 | 23.86 | 29.54 | 9.92 | 12.63 | 11.31 | 7.96 | 12.03 | 20.31 | 38.78 | 18.29 | 24.10 | 21.42 | 12.63 | 20.56 | |
| Front | SurroundOcc (Wei et al. 2023 [ 7 ] ) | 23.96 | 13.35 | 13.90 | 7.83 | 15.91 | 25.02 | 9.90 | 6.01 | 8.93 | 8.04 | 9.55 | 16.77 | 14.56 | 14.12 | 16.41 | 17.39 | 11.34 | 17.96 |
| Ours (SurroundOcc) | 29.30 | 15.19 | 14.64 | 4.03 | 18.64 | 25.26 | 9.40 | 7.48 | 7.41 | 6.92 | 10.27 | 17.56 | 35.35 | 14.09 | 21.57 | 19.24 | 11.91 | 19.33 | |
| Front Right | SurroundOcc (Wei et al. 2023 [ 7 ] ) | 29.89 | 15.92 | 15.12 | 7.28 | 17.66 | 27.23 | 8.65 | 8.21 | 9.40 | 7.49 | 10.29 | 17.43 | 36.58 | 17.82 | 21.45 | 19.29 | 11.80 | 19.01 |
| Ours (SurroundOcc) | 30.34 | 17.02 | 16.32 | 6.99 | 23.64 | 27.58 | 8.71 | 11.75 | 9.83 | 7.38 | 11.03 | 19.26 | 37.79 | 17.31 | 23.03 | 20.38 | 11.82 | 19.54 |
| Setting | Method | mIoU (vox) | mIoU (pts) | barrier | bicycle | bus | car | const.veh. | motorcycle | pedestrian | traffic cone | trailer | truck | drive.surf. | other flat | sidewalk | terrain | manmade | vegetation |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Standard | TPVFormer (Huang et al. 2023 [ 8 ] ) | 47.84 | 26.92 | 35.07 | 12.71 | 48.70 | 48.51 | 29.52 | 20.04 | 19.61 | 8.90 | 38.32 | 46.09 | 15.45 | 13.82 | 19.82 | 20.83 | 27.29 | 25.99 |
| Ours (TPVFormer) | 50.26 | 30.11 | 36.42 | 11.30 | 50.71 | 53.86 | 35.98 | 32.90 | 28.90 | 10.62 | 41.51 | 53.28 | 17.41 | 12.75 | 19.18 | 21.45 | 28.49 | 27.06 | |
| Front | TPVFormer (Huang et al. 2023 [ 8 ] ) | 43.82 | 25.45 | 34.17 | 12.04 | 39.19 | 47.20 | 29.71 | 17.60 | 17.41 | 8.50 | 32.28 | 44.52 | 18.04 | 13.42 | 19.66 | 21.73 | 26.94 | 24.80 |
| Ours (TPVFormer) | 46.89 | 27.94 | 34.43 | 10.92 | 40.95 | 52.07 | 34.78 | 32.18 | 27.38 | 10.33 | 36.86 | 48.97 | 16.44 | 11.32 | 17.39 | 20.69 | 27.41 | 24.98 | |
| Front Right | TPVFormer (Huang et al. 2023 [ 8 ] ) | 42.00 | 23.48 | 32.16 | 11.96 | 43.43 | 45.64 | 20.79 | 17.72 | 13.84 | 8.36 | 33.24 | 40.37 | 14.72 | 11.28 | 16.55 | 19.07 | 23.48 | 23.06 |
| Ours (TPVFormer) | 44.36 | 26.09 | 31.40 | 10.33 | 45.91 | 49.06 | 26.16 | 23.61 | 26.88 | 9.35 | 35.34 | 45.71 | 16.48 | 10.25 | 17.21 | 20.45 | 25.12 | 24.15 |
| Missing Views | Method | IoU (%) | mIoU (%) |
|---|---|---|---|
| 0 | Baseline | 31.18 | 17.23 |
| 1 | Baseline | 27.70 | 14.95 |
| + MMR | 29.70 | 15.75 | |
| + MMR + FMM | 29.77 | 16.25 | |
| 3 | Baseline | 20.00 | 10.04 |
| + MMR | 26.09 | 11.11 |
| Missing | MMR | FMM | IoU | mIoU |
|---|---|---|---|---|
| 31.18 | 17.23 | |||
| ✓ | 27.70 | 14.95 | ||
| ✓ | ✓ | 29.70 | 15.75 | |
| ✓ | ✓ | ✓ | 29.77 | 16.25 |
| Method | Number of Missing Views | Latency(s) | Memory(G) |
|---|---|---|---|
| Baseline | - | 0.50 | 5.927 |
| Ours | 1 | 0.77 | 6.077 |
| 2 | 0.91 | ||
| 3 | 1.00 | ||
| 4 | 1.11 | ||
| 5 | 1.25 |