Research
MLPerf Automotive
Overview Research area: Machine learning benchmarking and performance evaluation for automotive AI systems (ADAS, autonomous driving, and in-vehicle infotainment). Technical level: Intermediate. The p

- arXiv
- 2510.27065
- Published
- 2025-10-31
- Authors
- Radoyeh Shojaei, Predrag Djurdjevic, Mostafa El-Khamy, James Goel, Kasper Mecklenburg, John Owens, Pınar Muyan-Özçelik, Tom St. John, Jinho Suh, Arjun Suresh
AI summary
Overview
Research area: Machine learning benchmarking and performance evaluation for automotive AI systems (ADAS, autonomous driving, and in-vehicle infotainment).
Technical level: Intermediate. The paper is a benchmark design and methodology document rather than a novel algorithmic contribution; it assumes familiarity with ML inference concepts (latency, quantization, reference models) but explains automotive constraints in accessible terms.
Scope: The paper describes MLPerf Automotive v1.0, the first standardized public benchmark suite for measuring the inference performance of ML systems deployed in automotive AI accelerators, covering five tasks across perception, end-to-end driving, and infotainment.
What This Paper Is About
Automotive machine learning systems have unique constraints — real-time processing, safety, specialized sensor suites, long hardware lifespans, and wide ranges of compute and power — that existing benchmark suites for datacenter, mobile, IoT, or edge systems do not address. Today, automotive suppliers each evaluate systems individually, producing results that cannot be fairly compared. The goal of this work is to provide a collaborative, standardized benchmark framework with reproducible metrics so that automotive ML systems can be evaluated and compared on equal footing.
Key Contributions
-
First standardized automotive ML benchmark. MLPerf Automotive is presented as the first standardized public performance benchmark for evaluating ML systems deployed for AI acceleration in automotive systems, developed within MLCommons with participation across the automotive supply chain (IP and SoC suppliers, OS vendors, system integrators, Tier 1s, and OEMs).
-
A five-task benchmark suite spanning automation levels. The suite covers 2D object detection (SSD), 2D semantic segmentation (DeepLabv3+), 3D object detection (BEVFormer-tiny), end-to-end driving (UniAD), and an in-vehicle infotainment application (Llama-3.1). The first three were introduced in v0.5; UniAD and Llama-3.1 were added in the second iteration, the first major release (v1.0).
-
Automotive-specific methodology and inference rules. The paper defines two scenarios (Single Stream and Constant Stream, e.g., at a fixed 15 FPS), a 99.9% tail latency metric for safety-critical tasks (90% for Llama-3.1), accuracy constraints, closed and open submission divisions, and compliance tests, with prohibitions on retraining, caching results, and benchmark-aware preprocessing in the closed division.
-
Public reference implementations and open benchmark code. Reference models are provided in Python using PyTorch and/or ONNX, designed to run without accelerators and to be portable across the wide variety of automotive hardware architectures. The benchmark code is available at https://github.com/mlcommons/mlperf_automotive.
Main Findings
-
Automotive compute and power span extremely wide ranges: theoretical peak performance of automotive SoCs ranges from five TOPS to one thousand TOPS, and power consumption ranges from tens to hundreds of watts — higher than mobile devices and lower than high-end servers.
-
Automotive hardware must last far longer than typical computing hardware: passenger vehicle lifespans range from 9 to 23 years, requiring automotive SoCs to meet mechanical stress testing and functional safety requirements beyond other computing environments.
-
Existing public datasets cannot fully satisfy benchmark needs: no public dataset meets all requirements of geographic diversity, multimodal sensing, and labeling across 2D/3D perception, planning/prediction, and end-to-end driving. Licensing ambiguity (e.g., non-commercial public dataset licenses) further limits their use.
-
Three datasets were used: the public nuScenes dataset for BEVFormer and UniAD, the MMLU dataset for Llama-3.1-8B, and a synthetic dataset from Cognata for SSD and DeepLabv3+.
-
Synthetic Cognata data produced latency metrics close to real data: using a single H100 GPU over three benchmark runs, SSD mean latency was 73–76 ms on ZOD versus 76–78 ms on Cognata; DeepLabv3+ mean latency was 125–126 ms on ZOD versus 127–129 ms on Cognata. Tail latencies varied more, as expected on a shared, non-automotive-grade server.
-
Accuracy constraints differ by task criticality: 99% of the reference model accuracy for BEVFormer and UniAD, 99.9% for DeepLabv3+ and SSD, and a relaxed 95% for Llama-3.1-8B, which is not safety critical. Post-Training Quantization is permitted in the closed division; Quantization-Aware Training is prohibited.
-
Training SSD for 8 MP Cognata data was costly: training for 60 epochs on eight NVIDIA H100s took about 1.5 days, and model tuning was required to reach acceptable accuracy.
-
Model tuning improved SSD accuracy: starting from a baseline with anchor box feature sizes and steps (BSS) matched to Cognata's resolution at mAP 0.6483, anchor box scaling raised mAP by 0.0165 to 0.6648; combining scaling with an added feature map and a 5×5 detection head yielded the best mAP of 0.7141.
-
A proof-of-concept demo did not attract much interest: engagement and industry traction increased only after official results were published.
Methodology in Plain English
The authors followed the general structure of MLPerf Inference but adapted it to automotive needs. A reference implementation defines exactly what operations must be performed, while submitters are responsible for the system under test (SUT). The dataset, the Load Generator (LoadGen), and the accuracy scripts are supplied by MLPerf. LoadGen issues queries of sample IDs to the SUT, the SUT loads samples into memory, returns results, and the results are logged for latency and accuracy analysis.
Submitters produce two runs per benchmark: a performance run that measures tail latencies, and an accuracy run that verifies the model's accuracy meets a task-specific constraint expressed as a percentage of the reference model's accuracy. Compliance tests must also be run to confirm the rules were followed. In the closed division, retraining, caching, and benchmark-aware preprocessing are prohibited; the open division relaxes these restrictions.
Tasks and models were chosen to represent different SAE automation levels: SSD and DeepLabv3+ represent classical, smaller CNN architectures relevant to lower levels of autonomy (below SAE level 3), BEVFormer-tiny targets SAE level 3 and above, and UniAD targets SAE level 4 and above. Llama-3.1-8B was chosen for infotainment because the working group expected in-vehicle model sizes to grow and wanted a forward-looking choice. Because no public dataset satisfied all requirements — particularly 8 MP image resolution, matching commercially available automotive cameras — the group acquired a synthetic dataset from Cognata and validated that its latency behavior tracked closely to real 8 MP ZOD imagery.
Why This Matters
Impact on research: The paper establishes a shared methodology for measuring automotive inference performance, giving researchers and engineers a reproducible baseline for comparing architectures. It shifts automotive benchmarking from ad hoc, supplier-specific evaluation toward a common framework with defined tail-latency metrics, accuracy tolerances, and compliance checks.
Real-world applications:
- Advanced Driver Assistance Systems (ADAS) that must meet hard real-time latency deadlines for safety-critical perception.
- Autonomous driving stacks at higher SAE levels, including 3D perception and end-to-end driving pipelines.
- In-vehicle infotainment (IVI), such as navigation assistance powered by a large language model.
- Procurement and system selection for automotive ML hardware, since the benchmark is designed to be usable in the Request for Information/Quotation (RFI/RFQ) process.
Industry relevance: The working group spans the entire automotive supply chain, from IP and SoC suppliers to OS vendors, system integrators, carmakers, Tier 1s, and OEMs. Because producers and consumers of ML performance data participate in the same forum, the suite is kept aligned with representative workloads. The paper reports strong indications that the benchmarks are already being adopted into the RFI/RFQ process, which is a key phase in evaluating and ultimately selecting automotive ML systems that go into vehicles.
Future Directions
-
Incorporating vision-language-action (VLA) models. The authors identify the shift toward multi-modal and VLA models for end-to-end autonomous driving as the likely future direction, and plan to add a VLA model along with classical safety-oriented counterpart models.
-
A pre-silicon submission category. This would allow performance assessment during earlier stages of hardware development, before silicon is available.
-
Standardized power measurement protocols. These would provide power metrics alongside performance benchmarks, addressing the critical power constraints in automotive systems.
-
More sophisticated safety-centric accuracy metrics. Proposed additions include difficult objects, rare objects, zero-shot evaluation, and temporal accuracy clustering.
-
Continued suite expansion. The paper states that ongoing development will add additional models and scenarios to keep pace with the fast-moving automotive ML field.
Target Audience
This paper is most valuable to automotive ML engineers and hardware architects who select, design, or evaluate AI accelerators for vehicles; benchmark engineers and performance analysts who need to interpret or submit MLPerf Automotive results; and researchers working on perception, end-to-end driving, or in-vehicle infotainment who need to understand what performance targets are considered representative. Procurement and technical decision-makers in Tier 1 suppliers, OEMs, and SoC vendors will also benefit, since the benchmark is designed to feed into RFI/RFQ evaluations. Readers seeking novel model architectures or state-of-the-art accuracy results will not find them here; the contribution is methodological.
Authors’ abstract
We present MLPerf Automotive, the first standardized public benchmark for evaluating Machine Learning systems that are deployed for AI acceleration in automotive systems. Developed through a collaborative partnership between MLCommons and the Autonomous Vehicle Computing Consortium, this benchmark addresses the need for standardized performance evaluation methodologies in automotive machine learning systems. Existing benchmark suites cannot be utilized for these systems since automotive workloads have unique constraints including safety and real-time processing that distinguish them from the domains that previously introduced benchmarks target. Our benchmarking framework provides latency and accuracy metrics along with evaluation protocols that enable consistent and reproducible performance comparisons across different hardware platforms and software implementations. The first iteration of the benchmark consists of automotive perception tasks in 2D object detection, 2D semantic segmentation, and 3D object detection. We describe the methodology behind the benchmark design including the task selection, reference models, and submission rules. We also discuss the first round of benchmark submissions and the challenges involved in acquiring the datasets and the engineering efforts to develop the reference implementations. Our benchmark code is available at https://github.com/mlcommons/mlperf_automotive.