HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End-to-end Autonomous Driving
Organizations: Wangxuan Institute of Computer Technology, Peking University, Beijing, China. · University of California, Merced, USA.
Abstract
Three-dimensional feature extraction and multi-task perception are fundamental components of modern autonomous driving systems. Although large image encoders, high-resolution inputs, and long temporal contexts can substantially improve representation quality and overall performance, jointly leveraging these strategies remains challenging due to prohibitive computational costs during both training and inference. Furthermore, different perception tasks often require distinct feature representations, making it difficult for a unified architecture to achieve end-to-end multi-task performance comparable to specialized single-task systems. To address these challenges, we propose HENet++, a unified framework for multi-task 3D perception and end-to-end autonomous driving that employs a hybrid image encoding strategy, using a large encoder for short-term frames and a lightweight encoder for long-term temporal context to strike a favorable balance between accuracy and efficiency. The framework jointly extracts dense background features and sparse foreground features, enabling task-specific representations that reduce cumulative errors and provide richer information for downstream prediction and planning modules. HENet++ is compatible with diverse 3D feature extraction pipelines and supports multi-modal inputs, including camera and radar data. Extensive experiments demonstrate state-of-the-art performance on the nuScenes multi-task 3D perception benchmark, achieving the lowest collision rate on the nuScenes planning benchmark and higher PDMS on the NAVSIM benchmark.
Figures & tables
| Methods | Backbone | Frames | Time/E | Detection | BEV Segmentation | Occupancy | |||||
| NDS | mAP | ||||||||||
| VPN | ResNet50 | 1 | - | 33.4 | 25.7 | 43.8 | 37.3 | 76.0 | 18.0 | - | - |
| LSS | ResNet50 | 1 | - | 41.0 | 34.4 | 45.0 | 42.8 | 73.9 | 18.3 | - | 11.4 |
| BEVFormer-S | ResNet101 | 1 | - | 45.3 | 38.0 | 47.3 | 44.4 | 77.6 | 19.8 | - | - |
| BEVFormer | ResNet101 | 5 | 213min | 52.0 | 41.2 | 49.4 | 46.7 | 77.5 | 23.9 | - | 30.5 |
| OccNet | ResNet101 | 5 | 230min | 52.0 | 41.2 | 24.6 | 12.9 | 47.2 | 13.8 | 41.1 | 27.0 |
| Method | Input | Backbone | UniAD Metrics | VAD/STP3 Metrics | |||||||||||||||
| L2 (m) | Collision Rate (%) | CCR | L2 (m) | Collision Rate (%) | |||||||||||||||
| 1s | 2s | 3s | Avg. | 1s | 2s | 3s | Avg. | Avg. | 1s | 2s | 3s | Avg. | 1s | 2s | 3s | Avg. | |||
| ST-P3 | C | EfficientNet-b4 | 1.72 | 3.26 | 4.86 | 3.28 | 0.44 | 1.08 | 3.01 | 1.51 | 8.37 | 1.33 | 2.11 | 2.90 | 2.11 | 0.23 | 0.62 | 1.27 | 0.71 |
| OccNet | C | ResNet101-DCN | 1.29 | 2.13 | 2.99 | 2.14 | 0.21 | 0.59 | 1.37 | 0.72 | - | - | - | - | - | - | - | - | - |
| UniAD | C | ResNet101 | 0.48 | 0.96 | 1.65 | 1.03 | 0.05 | 0.17 | 0.71 | 0.31 | 1.59 | 0.45 | 0.70 | 1.04 | 0.73 | 0.62 | 0.58 | 0.63 | 0.61 |
| Ego-MLP | C | ResNet50 | 0.15 | 0.32 | 0.59 | 0.35 | 0.00 | 0.27 | 0.85 | 0.37 | 2.93 | - | - | - | - | - | - | - | - |
| Method | Mod. | NC | DAC | TTC | Comf. | EP | PDMS |
| UniAD | C | 97.8 | 91.9 | 92.9 | 100.0 | 78.8 | 83.4 |
| PARA-Drive | C | 97.9 | 92.4 | 93.0 | 99.8 | 79.3 | 84.0 |
| iPad | C | 98.6 | 98.3 | 94.9 | 100.0 | 88.0 | 91.7 |
| HENet++ | C | 98.9 | 98.7 | 94.2 | 100.0 | 88.3 | 92.6 |
| Methods | Backbone | NDS | mAP | mATE | mASE | mAOE | mAVE | mAAE |
| BEVDet | ResNet50 | 37.9 | 29.8 | 0.725 | 0.279 | 0.589 | 0.860 | 0.245 |
| BEVDet4D | ResNet50 | 45.7 | 32.2 | 0.703 | 0.278 | 0.495 | 0.354 | 0.206 |
| PETRv2 | ResNet50 | 45.6 | 34.9 | 0.700 | 0.275 | 0.580 | 0.437 | 0.187 |
| BEVStereo | ResNet50 | 50.0 | 37.2 | 0.598 | 0.270 | 0.438 | 0.367 | 0.190 |
| SOLOFusion | ResNet50 | 53.4 | 42.7 | 0.567 | 0.274 | 0.511 | 0.252 | 0.181 |
| Sparse4Dv2 | ResNet50 | 53.9 | 43.9 | 0.598 | 0.270 | 0.475 | 0.282 | 0.179 |
| Methods | Backbone | NDS | mAP | mATE | mASE | mAOE | mAVE | mAAE |
| BEVDet4D | Swin-B | 56.9 | 45.1 | 0.511 | 0.241 | 0.386 | 0.301 | 0.121 |
| PolarFormer | V2-99 | 57.2 | 49.3 | 0.556 | 0.256 | 0.364 | 0.439 | 0.127 |
| PETRv2 | V2-99 | 58.2 | 49.0 | 0.561 | 0.243 | 0.361 | 0.343 | 0.120 |
| HoP-BEVFormer | V2-99 | 60.3 | 51.7 | 0.501 | 0.245 | 0.346 | 0.362 | 0.105 |
| BEVDepth | ConvNeXt-B | 60.9 | 52.0 | 0.445 | 0.243 | 0.352 | 0.347 | 0.127 |
| BEVStereo | V2-99 | 61.0 | 52.5 | 0.431 | 0.246 | 0.358 | 0.357 | 0.138 |
| Methods | Backbone | ||||
| VPN | ResNet50 | 42.7 | 31.8 | 76.9 | 19.4 |
| LSS | ResNet50 | 46.5 | 41.7 | 77.7 | 20.0 |
| BEVFormer | ResNet101-DCN | 50.2 | 44.8 | 80.1 | 25.7 |
| FIERY | ResNet-101 | - | 38.2 | - | - |
| M2BEV | ResNeXt-101 | - | - | 77.2 | 40.5 |
| PETRv2 | V2-99 | 60.3 | 46.3 | 85.6 | 49.0 |
| Method | Backbone | Visible Mask | RayIoU | mIoU | others | barrier | bicycle | bus | car | const. veh. | motorcycle | pedestrian | traffic cone | trailer | truck | drive. suf. | other flat | sidewalk | terrain | manmade | vegetation |
| BEVFormer | R-50 | ✔ | 32.4 | 23.7 | 5.0 | 38.8 | 10.0 | 34.4 | 41.1 | 13.2 | 16.5 | 18.2 | 17.8 | 18.7 | 27.7 | 49.0 | 27.7 | 29.1 | 25.4 | 15.4 | 14.5 |
| SparseOcc | R-50 | ✔ | 36.1 | 30.9 | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - |
| OPUS | R-50 | ✔ | 38.4 | 33.2 | 10.7 | 39.8 | 21.3 | 39.8 | 45.3 | 23.4 | 21.8 | 17.8 | 19.3 | 27.5 | 33.2 | 71.6 | 37.1 | 45.1 | 43.6 | 33.8 | 33.2 |
| TPVFormer | R-50 | ✔ | - | 34.2 | 7.7 | 44.0 | 17.7 | 40.9 | 47.0 | 15.1 | 20.5 | 24.7 | 24.7 | 24.3 | 29.3 | 79.3 | 40.7 | 48.5 | 49.4 | 32.6 | 29.8 |
| ODG | R-50 | ✔ | 39.2 | 35.5 | 13.7 | 39.0 | 23.0 | 46.8 | 49.3 | 25.8 | 23.6 | 20.7 | 18.5 | 30.0 | 35.6 | 76.8 | 39.3 | 45.0 | 46.8 | 37.5 | 32.2 |
| OccFormer | R-50 | ✔ | - | 37.4 | 9.2 | 45.8 | 18.2 | 42.8 | 50.3 | 24.0 | 20.8 | 22.9 | 21.0 | 31.9 | 38.1 | 80.1 | 38.2 | 50.8 | 54.3 | 46.4 | 40.2 |
| Model | Backbone | Input | Frames | NDS | mAP | FPS | GPU memory | Time | |
| A | BEVDepth4D | R18 | 640 1152 | 2 frames | 48.6 | 34.8 | 14.3 | 31.2G / BS=64 | 14h |
| B | BEVDepth4D | R18 | 256 704 | 9 frames | 48.6 | 34.9 | 21.7 | 22.9G / BS=64 | 14h |
| C | BEVDepth4D | R18 | 640 1152 | 9 frames | 51.2 | 39.2 | 14.1 | 43.5G / BS=64 | 44h |
| A+B | Model Ensemble | - | - | - | 48.9 | 35.2 | 7.93 | - | - |
| A+B | Dense Hybrid Encoding | R18 & R18 | 640 1152 & 256 704 | 2 + 7 frames | 52.1 | 39.8 | 8.91 | 41.3G / BS=64 | 13h |
| D | BEVDepth4D | R50 | 256 704 | 9 frames | 53.8 | 40.9 | 19.1 | 14.1G / BS=16 | 16h |
| Model | Backbone | Input | Frames | NDS | |||
| A | BEVDepth4D | R50 | 256 704 | 9 frames | 53.0 | 51.5 | - |
| B | BEVStereo | V2-99 | 640 1152 | 2 frames | 58.0 | 55.8 | - |
| C | BEVStereo | V2-99 | 896 1600 | 2 frames | 58.9 | 56.7 | - |
| A+B | Dense Hybrid Encoding w/o AFFM | V2-99 & R50 | 640 1152 & 256 704 | 2 + 7 frames | 59.0 | 56.9 | - |
| A+B | Dense Hybrid Encoding | V2-99 & R50 | 640 1152 & 256 704 | 2 + 7 frames | 59.9 | 58.0 | - |
| D | HENet++ w/o Hybrid Encoding | R50 | 256 704 | 9 frames | 54.1 | 51.9 | 39.4 |
| Load from | NDS | ||
| Detection | 53.6 0.2 | 54.3 0.5 | 36.2 0.3 |
| Segmentation | 52.8 0.4 | 55.0 0.3 | 34.9 0.6 |
| Occupancy | 53.2 0.2 | 54.9 0.5 | 37.3 0.2 |
| Model Merge | 53.8 0.3 | 55.3 0.3 | 37.6 0.2 |
| UniAD Metrics | VAD/STP3 Metrics | |||||
| mean L2 | mean Col. | mean L2 | mean Col. | |||
| ✓ | 1.69 | 0.52 | 0.71 | 0.17 | ||
| ✓ | ✓ | 1.36 | 0.20 | 0.58 | 0.10 | |
| ✓ | ✓ | ✓ | 1.29 | 0.13 | 0.55 | 0.05 |