cs.LGOct 31, 2025

MLPerf Automotive

Authors: Radoyeh Shojaei, Predrag Djurdjevic, Mostafa El-Khamy, James Goel, Kasper Mecklenburg, Pınar Muyan-Özçelik, John Owens, Tom St. John, +2 more

Organizations: University of California, Davis · Arm · Samsung · Qualcomm · California State University, Sacramento · Gilmet Labs · NVIDIA · AMD

Abstract

We present MLPerf Automotive, the first standardized public performance benchmark for evaluating Machine Learning systems that are deployed for AI acceleration in automotive systems. Developed through a collaborative partnership within MLCommons, this benchmark addresses the need for standardized performance evaluation methodologies in automotive machine learning systems. Existing benchmark suites cannot be utilized for these systems since automotive workloads have unique constraints including sensor suites, safety, and real-time processing that distinguish them from the domains that previously introduced benchmarks target. Our implemented and adopted MLPerf Automotive benchmark is a framework for evaluation and methodology for benchmarking automotive systems with reproducible performance metrics. The benchmark consists of automotive perception tasks in 2D object detection, 2D semantic segmentation, 3D object detection, end-to-end driving, and an infotainment system application. We carefully curated and customized models for automotive use cases. We describe the methodology behind the benchmark design including the task selection, reference models, and submission rules. We also discuss the challenges involved in acquiring the datasets and the engineering efforts to develop the reference implementations. Our benchmark code is available at https://github.com/mlcommons/mlperf_automotive.

Figures & tables

Explore similar work

CardsList
  1. Bench2Drive-Robust: Benchmarking Closed-Loop Autonomous Driving under Deployment Perturbations

    May 18, 2026Zhiyuan Zhang, Zhenghao Jin, Yanlun Peng +8Autonomous Driving SimulationAutonomous Driving

  2. Automated Benchmark Auditing for AI Agents and Large Language Models

    May 25, 2026Junlin Wang, Federico Bianchi, Shang Zhu +4Artificial Intelligence BenchmarksAgentic Benchmarks

  3. Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models

    May 12, 2026Mohammed Saidul Islam, Negin Baghbanzadeh, Farnaz Kohankhaki +5Ground TruthFoundation Model