We present MLPerf Automotive, the first standardized public performance benchmark for evaluating Machine Learning systems that are deployed for AI acceleration in automotive systems. Developed through a collaborative partnership within MLCommons, this benchmark addresses the need for standardized performance evaluation methodologies in automotive machine learning systems. Existing benchmark suites cannot be utilized for these systems since automotive workloads have unique constraints including sensor suites, safety, and real-time processing that distinguish them from the domains that previously introduced benchmarks target. Our implemented and adopted MLPerf Automotive benchmark is a framework for evaluation and methodology for benchmarking automotive systems with reproducible performance metrics. The benchmark consists of automotive perception tasks in 2D object detection, 2D semantic segmentation, 3D object detection, end-to-end driving, and an infotainment system application. We carefully curated and customized models for automotive use cases. We describe the methodology behind the benchmark design including the task selection, reference models, and submission rules. We also discuss the challenges involved in acquiring the datasets and the engineering efforts to develop the reference implementations. Our benchmark code is available at https://github.com/mlcommons/mlperf_automotive.
Figures & tables
Level 0
Level 1
Level 2
Level 3
Level 4
Level 5
Name
No Automation
Driver Assistance
Partial Driving Automation
Conditional Driving Automation
High Driving Automation
Full Driving Automation
Definition
Human driver in full control even when enhanced by active safety systems.
System performs lateral or longitudinal vehicle motion, but not both.
System performs lateral and longitudinal vehicle motion.
System drives under limited conditions. Human driver is ready to take over if needed.
System drives under limited conditions. Human driver is not expected to take over.
System drives under all conditions.
Human Driver Engagement
Human driver is constantly engaged when support features are active.
Human driver takes over when requested.
Human driver is not required to take over.
Table 1: SAE International levels of driving automation ( SAE International 2021 ) . Blue indicates where a human driver is actively engaged in driving. Yellow indicates an automated system is driving.
MLPerf Automotive
MLPerf Inference
Domain
Automotive
General inference with datacenter focus
Tasks
Perception for driving
Language, vision, speech, etc.
Datasets
Driving scenes
Varied text, speech, and image
Table 2: Differences between the scopes of MLPerf Automotive and MLPerf Inference.
Figure 1: The goal of standardizing the benchmarking process for automotive system suppliers. On the left is the complicated individual benchmarking process and on the right is the standardized use of MLPerf Automotive.
Figure 2: A system under test (SUT) during an inference run. (1) Setup benchmark, model, dataset, pre/post processing. (2) LoadGen creates queries of Sample IDs from the dataset for SUT. (3) Load samples into memory. (4) SUT is ready. (5) Issue request to SUT. (6) SUT return results and results are post-processed. (7) Logs output for latency and accuracy analysis.
Model
Tail latency (%)
Accuracy constraint (%)
Params
Images/ query
Image resolution
Target SAE
Tail latency (ms)
BEVFormer-tiny
99.9
99
45M
6
800 × 450
≥ 3
100
SSD
99.9
99.9
14M
1
3840 × 2160
< 3
592
DeepLabv3+
99.9
99.9
40M
1
3840 × 2160
≤ 3
1728
UniAD
99.9
99
85M
6
800 × 450
≥ 4
614
Llama-3.1
90
95
8B
—
—
—
58
Table 3: Overview of the automotive benchmark suite. The tail latency and accuracy are both expressed as percentiles. The accuracy constraint is a percentage of the reference model accuracy that needs to be met. The first three models were introduced in v0.5 and the last two added after the v0.5 submissions. The reference model tail latency was measured using a V100 GPU for all models except Llama-3.1, which was done on an H100.
Figure 3: Benchmark scenarios
Model
Dataset
Mean Latency (ms)
99% Tail Latency (ms)
99.9% Tail Latency (ms)
SSD
ZOD
73–76
96–101
102–105
SSD
Cognata
76–78
99–102
103–107
DeepLabv3+
ZOD
125–126
131–137
148–151
DeepLabv3+
Cognata
127–129
133–144
149–155
Table 4: Latency comparison across models and datasets. Ranges are min and max values for each metric achieved across three benchmark runs.
Figure 4: Sample images from MLCommons Cognata dataset (top row) and nuScenes (bottom row)
Model
mAP
BSS (baseline)
0.6483
+scales
0.6648
+scales+Feature map
0.6943
+scales+5 × 5Detection Head
0.6767
+scales+both
0.7141
Table 5: mAP of different SSD variants. Scales refers to improving anchor box scaling to match the MLCommons Cognata dataset.
Figure 5: SSD trained on the MLCommons Cognata dataset for 60 epochs. The variant with the best accuracy showed immediate benefit in the first epoch and maintained better accuracy until accuracy plateaued for all variants.
Robustness is a critical requirement for deploying autonomous driving systems in the real world. Existing robustness benchmarks for autonomous driving have made important progress in studying the effects of image-level corruptions, such as adverse weather or camera degradation, on perception modules and open-loop planning outputs. However, deployment can also involve system-level imperfections, such as inference latency and ego-state estimation errors, which remain less studied in closed-loop E2E-AD evaluation. These imperfections can accumulate through the feedback loop and destabilize control. In this work, we present Bench2Drive-Robust, to our knowledge the first device-centric robustness benchmark for closed-loop end-to-end autonomous driving under realistic deployment perturbations. We systematically evaluate deployment-oriented perturbations arising from three major sources: camera-stream failures (frame drop, partial observation), ego-state estimation errors (GPS noise, and speed or odometry errors), and compute-induced control delay (model inference delay). We evaluate representative end-to-end driving methods and analyze their robustness under different perturbation severities. Our results show that these deployment-related perturbations can substantially degrade closed-loop driving performance, revealing robustness challenges that are not fully captured by conventional image-level corruption evaluations. By establishing a closed-loop evaluation protocol and demonstrating the substantial impact of these deployment-oriented perturbations, Bench2Drive-Robust defines practical robustness problems for end-to-end autonomous driving and encourages further research on deployment-aware robust driving systems.
Zhiyuan Zhang, Zhenghao Jin, Yanlun Peng +8
Sch. of Computer Science & Sch. of Artificial Intelligence, Shanghai Jiao Tong University · Institute of Trustworthy Embodied AI (TEAI), Fudan University · Great Wall Motor +2
Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit assumptions, incomplete environment specifications, and brittle evaluation logic that human annotation cannot reliably catch. We introduce Auto Benchmark Audit (ABA), an agentic framework that systematically audits individual benchmark tasks, uncovering issues such as hidden environment dependencies, specification gaps, and limited grading logic. We run ABA on a collection of frontier LLM benchmarks and previous NeurIPS publications, totaling 168 benchmarks across nine domains. Across this corpus, ABA identifies critical issues including ambiguous task design, execution environment conflicts, and incorrect ground truths in over 25.7% of the evaluated tasks. The precision of these automated audits is validated by expert review and independent third-party reports such as upstream PRs. Crucially, we demonstrate that these problematic tasks severely distorts capability assessments for agents and LLMs: filtering out these tasks with issues shifts model rankings and increases average performance on SWE-bench Verified and Terminal-Bench 2 by 9.9% and 9.6%, respectively. We release the agentic tool and all task annotations to support the future development of frontier benchmarks.
Junlin Wang, Federico Bianchi, Shang Zhu +4
Duke University · Together AI · Stanford University
Evaluation of foundation models often rely on aggregate scores from benchmarks that lack comprehensive coverage and metadata for a fine-grained evaluation. We introduce a framework for automated benchmark generation. Our framework generates evaluation problems grounded in reference material, such as textbooks, producing benchmarks with broad coverage, rich metadata, and robustness to contamination. The pipeline employs a multi-agent architecture for problem generation and a solution-graph-driven strategy that significantly improves the reliability of ground truth solutions. Using the framework, we generate three benchmarks in Machine Learning, Corporate Finance, and Personal Finance. Expert review finds a significantly lower ground-truth error rate than previous benchmarks such as MMLU and GSM8K. Evaluation of 12 commercial and open-source models shows that our benchmarks achieve near-uniform competency coverage and surface performance differences across models that existing benchmarks fail to capture. We will open-source the framework and our curated benchmarks soon.
Mohammed Saidul Islam, Negin Baghbanzadeh, Farnaz Kohankhaki +5