We present MLPerf Automotive, the first standardized public performance benchmark for evaluating Machine Learning systems that are deployed for AI acceleration in automotive systems. Developed through a collaborative partnership within MLCommons, this benchmark addresses the need for standardized performance evaluation methodologies in automotive machine learning systems. Existing benchmark suites cannot be utilized for these systems since automotive workloads have unique constraints including sensor suites, safety, and real-time processing that distinguish them from the domains that previously introduced benchmarks target. Our implemented and adopted MLPerf Automotive benchmark is a framework for evaluation and methodology for benchmarking automotive systems with reproducible performance metrics. The benchmark consists of automotive perception tasks in 2D object detection, 2D semantic segmentation, 3D object detection, end-to-end driving, and an infotainment system application. We carefully curated and customized models for automotive use cases. We describe the methodology behind the benchmark design including the task selection, reference models, and submission rules. We also discuss the challenges involved in acquiring the datasets and the engineering efforts to develop the reference implementations. Our benchmark code is available at https://github.com/mlcommons/mlperf_automotive.
Figures & tables
Level 0
Level 1
Level 2
Level 3
Level 4
Level 5
Name
No Automation
Driver Assistance
Partial Driving Automation
Conditional Driving Automation
High Driving Automation
Full Driving Automation
Definition
Human driver in full control even when enhanced by active safety systems.
System performs lateral or longitudinal vehicle motion, but not both.
System performs lateral and longitudinal vehicle motion.
System drives under limited conditions. Human driver is ready to take over if needed.
System drives under limited conditions. Human driver is not expected to take over.
System drives under all conditions.
Human Driver Engagement
Human driver is constantly engaged when support features are active.
Human driver takes over when requested.
Human driver is not required to take over.
Table 1: SAE International levels of driving automation ( SAE International 2021 ) . Blue indicates where a human driver is actively engaged in driving. Yellow indicates an automated system is driving.
MLPerf Automotive
MLPerf Inference
Domain
Automotive
General inference with datacenter focus
Tasks
Perception for driving
Language, vision, speech, etc.
Datasets
Driving scenes
Varied text, speech, and image
Table 2: Differences between the scopes of MLPerf Automotive and MLPerf Inference.
Figure 1: The goal of standardizing the benchmarking process for automotive system suppliers. On the left is the complicated individual benchmarking process and on the right is the standardized use of MLPerf Automotive.
Figure 2: A system under test (SUT) during an inference run. (1) Setup benchmark, model, dataset, pre/post processing. (2) LoadGen creates queries of Sample IDs from the dataset for SUT. (3) Load samples into memory. (4) SUT is ready. (5) Issue request to SUT. (6) SUT return results and results are post-processed. (7) Logs output for latency and accuracy analysis.
Model
Tail latency (%)
Accuracy constraint (%)
Params
Images/ query
Image resolution
Target SAE
Tail latency (ms)
BEVFormer-tiny
99.9
99
45M
6
800 × 450
≥ 3
100
SSD
99.9
99.9
14M
1
3840 × 2160
< 3
592
DeepLabv3+
99.9
99.9
40M
1
3840 × 2160
≤ 3
1728
UniAD
99.9
99
85M
6
800 × 450
≥ 4
614
Llama-3.1
90
95
8B
—
—
—
58
Table 3: Overview of the automotive benchmark suite. The tail latency and accuracy are both expressed as percentiles. The accuracy constraint is a percentage of the reference model accuracy that needs to be met. The first three models were introduced in v0.5 and the last two added after the v0.5 submissions. The reference model tail latency was measured using a V100 GPU for all models except Llama-3.1, which was done on an H100.
Figure 3: Benchmark scenarios
Model
Dataset
Mean Latency (ms)
99% Tail Latency (ms)
99.9% Tail Latency (ms)
SSD
ZOD
73–76
96–101
102–105
SSD
Cognata
76–78
99–102
103–107
DeepLabv3+
ZOD
125–126
131–137
148–151
DeepLabv3+
Cognata
127–129
133–144
149–155
Table 4: Latency comparison across models and datasets. Ranges are min and max values for each metric achieved across three benchmark runs.
Figure 4: Sample images from MLCommons Cognata dataset (top row) and nuScenes (bottom row)
Model
mAP
BSS (baseline)
0.6483
+scales
0.6648
+scales+Feature map
0.6943
+scales+5 × 5Detection Head
0.6767
+scales+both
0.7141
Table 5: mAP of different SSD variants. Scales refers to improving anchor box scaling to match the MLCommons Cognata dataset.
Figure 5: SSD trained on the MLCommons Cognata dataset for 60 epochs. The variant with the best accuracy showed immediate benefit in the first epoch and maintained better accuracy until accuracy plateaued for all variants.
Sch. of Computer Science & Sch. of Artificial Intelligence, Shanghai Jiao Tong University · Institute of Trustworthy Embodied AI (TEAI), Fudan University · Great Wall Motor +2