cs.AIOct 8, 2026

The Operator Mismatch Problem: Deploying BEV Perception with Portable GPU Compute

Authors: Rohit Verma, Anand V Bodas

Organizations: Intel

Abstract

Modern autonomous driving systems rely on bird's-eye-view (BEV) perception models that fuse camera and LiDAR inputs to detect objects in 3D space. These models are accurate, but they cannot be deployed through standard inference runtimes. The reason is an operator mismatch between dense convolutions (which runtimes handle well), sparse 3D convolutions (which runtimes cannot represent), and geometric scatter operations (which runtimes have no vocabulary for). Today, every sparse convolution library is CUDA-only and PyTorch-coupled, locking BEV deployment to a single vendor's hardware and a single execution framework. We present BEVPIPE, a framework for deploying multimodal BEV perception pipelines using portable GPU compute APIs and integrating them with production inference runtimes. BEVPIPE partitions the model into runtime-managed dense subgraphs and three external operator extensions (voxelizer, sparse encoder, BEV projector), connected through a shared GPU memory space. BEVPIPE achieves a 19.5x end-to-end speedup over conventional deployments while retaining 98.5% of reference mAP. We also showcase that BEVPIPE is portable across different GPU backends.

Figures & tables

Explore similar work

CardsList
  1. FlashBEV: Fast and Memory-Efficient Exact BEV Transformation with IO-Awareness

    Jul 11, 2026Shunsuke Yokokawa, Hironori KasaharaGPU Kernel OptimizationAutonomous Driving Perception

  2. Fast-BEV++: Fast by Algorithm, Deployable by Design

    Dec 9, 2025Yuanpeng Chen, Hui Song, Sheng Yang +5Autonomous Driving PerceptionBird's-Eye-View Representation

  3. Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints

    Jun 23, 2026Hojun Choi, Seulbin Hwang, Dae Jung Kim +3VLMs for Autonomous DrivingOpen-Vocabulary Semantic Segmentation