cs.CVNov 10, 2025

HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End-to-end Autonomous Driving

Authors: Zhongyu Xia, Zhiwei Lin, Yongtao Wang, Ming-Hsuan Yang

Organizations: Wangxuan Institute of Computer Technology, Peking University, Beijing, China. · University of California, Merced, USA.

Abstract

Three-dimensional feature extraction and multi-task perception are fundamental components of modern autonomous driving systems. Although large image encoders, high-resolution inputs, and long temporal contexts can substantially improve representation quality and overall performance, jointly leveraging these strategies remains challenging due to prohibitive computational costs during both training and inference. Furthermore, different perception tasks often require distinct feature representations, making it difficult for a unified architecture to achieve end-to-end multi-task performance comparable to specialized single-task systems. To address these challenges, we propose HENet++, a unified framework for multi-task 3D perception and end-to-end autonomous driving that employs a hybrid image encoding strategy, using a large encoder for short-term frames and a lightweight encoder for long-term temporal context to strike a favorable balance between accuracy and efficiency. The framework jointly extracts dense background features and sparse foreground features, enabling task-specific representations that reduce cumulative errors and provide richer information for downstream prediction and planning modules. HENet++ is compatible with diverse 3D feature extraction pipelines and supports multi-modal inputs, including camera and radar data. Extensive experiments demonstrate state-of-the-art performance on the nuScenes multi-task 3D perception benchmark, achieving the lowest collision rate on the nuScenes planning benchmark and higher PDMS on the NAVSIM benchmark.

Figures & tables

Explore similar work

CardsList
  1. MATS: A novel multi-modality multi-task learning framework for 3D perception in autonomous driving

    Jul 27, 2026Junchen Huo, Wanming Hao, Song Wang +3Mixture of ExpertsCamera-LiDAR Fusion

  2. Geometry-Grounded Unified 3D Perception for Autonomous Driving

    Aug 13, 2026Longfei Xu, Xiaohui Wang, Zehao Huang +4Depth EstimationAutonomous Driving Perception

  3. Towards Compact Autonomous Driving Perception with Balanced Learning and Multi-sensor Fusion

    Jun 2, 2026Oskar Natan, Jun MiuraDepth EstimationAutonomous Driving Perception