Behavioral Safety Assessment towards Large-scale Deployment of Autonomous Vehicles, Part II: Assessment Results
Authors: Henry X. Liu, Tinghan Wang, Xintao Yan, Haowei Sun, Zhijie Qiao, Kenneth Boyd, Shuo Feng, Greg Stevens, +1 more
Organizations: University of Michigan Transportation Research Institute, Ann Arbor, MI 48109 USA · Department of Civil and Environmental Engineering, University of Michigan, Ann Arbor, MI 48109 USA · Department of Civil Engineering, The University of Hong Kong, Hong Kong 999077, China · Laplace Intelligence, Ann Arbor, MI 48109 USA · Department of Automation, Tsinghua University, Beijing 100084, China
Third-party evaluations of autonomous vehicle (AV) safety can play a vital role in improving public acceptance, building consumer confidence, and establishing effective safety standards. In Part I of this study, we propose a dedicated third-party testing initiative for systematically evaluating AV behavioral safety. In this paper, we validate our proposed framework using Autoware.Universe, an open-source Level 4 Automated Driving System (ADS), tested both in simulated environments and on the physical test track at the University of Michigan's Mcity Testing Facility. The results indicate that Autoware.Universe possesses 6 out of 14 behavioral competencies and exhibited a crash rate of 3.01x10^-3 crashes per mile, approximately 1,000 times higher than the average human driver crash rate. During the tests, we also uncovered a number of unknown unsafe scenarios for Autoware.Universe. These findings underscore the necessity of behavioral safety evaluations for improving AV safety performance prior to widespread public deployment.
Figures & tables
Fig. 1 : Sketches of the behavioral competencies tested in this study
Fig. 2 : Demonstration of the cut-in scenario development pipeline
Fig. 3 : Demonstration of the simulated Mcity environment for the Driving Intelligence Test
Behavioral Competency
Case number
Result
Collision
Phase
Behavioral category (%)
(%)
Aggressive
Assertive
Normal
Conservative
Ultra-conservative
(a) Detect and respond to cut-in vehicle
36
Pass
0.00
Pre
2.78
13.89
1.85
2.78
78.70
Post
0.00
0.00
100.00
0.00
0.00
(b) Detect and follow the leading vehicle
40
Pass
0.00
Acc
0.00
0.00
13.33
86.67
0.00
Cruise
0.00
0.00
0.00
0.00
100.00
Dec
0.00
0.00
0.00
0.00
100.00
TABLE I : Overall evaluation results of BCT
Fig. 4 : Analysis of pre-encroachment phase
Fig. 5 : Analysis of post-encroachment phase
Fig. 6 : Behavior diagnosis
Fig. 7 : The Driving Intelligence Test results of Autoware.Universe. a, Demonstration of the Mcity testing environment and testing route of AV. b, Examples of different crash types experienced by Autoware.Universe: head-on (b.1), sideswipe (b.2), angle (b.3), and rear-end (b.4) collisions. The red vehicle represents the AV under test, while the blue vehicles represent background traffic. c, Crash rate estimation. d, Crash type estimation. e, Crash severity estimation. f-h, Demonstration of identified issues in Autoware.Universe: obstacle avoidance replanning error (f), right-of-way yielding issue (g), and trajectory prediction error (h). The white vehicle represents the AV, blue vehicles indicate background traffic, and the vehicle circled in red highlights the BV that is in conflict with the AV. A video demonstration can be found in Supplementary Movie 4.
Fig. 8 : The field experiment of the BCT. a.1-a.3, The test system. a.1, The vehicle under test equipped with the RTK system. a.2, The Humanetics UFO Pro platform installed with the dummy vehicle. a.3, The Humanetics UFO Nano platform installed with the dummy child. b.1-b.2, Testing process for the left turn (AV goes straight) scenario. b.1, Images captured by the roadside camera and Tesla’s front-facing camera at the scenario start moment, defined as the point when the relative distance and speed satisfy the conditions specified in the test case. b.2, Images captured by the roadside camera and Tesla’s front-facing camera when the AV comes to a stop. c.1-c.2, Testing process for the VRU crossing the street without the crosswalk scenario. c.1, Images captured by the roadside camera and Tesla’s front-facing camera at the scenario start moment. c.2, Images captured by the roadside camera and Tesla’s front-facing camera when the AV comes to a stop.
Fig. 9 : A snapshot of the field experiment of the Driving Intelligence Test. a, Testing environment view showing the complete traffic environment of Mcity, with the white vehicle representing the AV and all yellow vehicles representing the BVs. The red circled vehicle denotes the conflicting BV. b, Autoware.Universe’s view, with the white vehicle representing the AV and surrounding BVs indicated by blue boxes. c, Raw forward-facing camera image from the AV. d, Augmented forward-facing camera image, with BVs integrated into the scene.
Autonomous vehicles (AVs) have significantly advanced in real-world deployment in recent years, yet safety continues to be a critical barrier to widespread adoption. Traditional functional safety approaches, which primarily verify the reliability, robustness, and adequacy of AV hardware and software systems from a vehicle-centric perspective, do not sufficiently address the AV's broader interactions and behavioral impact on the surrounding traffic environment. To overcome this limitation, we propose a paradigm shift toward behavioral safety, a comprehensive approach focused on evaluating AV responses and interactions within the traffic environment. To systematically assess behavioral safety, we introduce a third-party AV safety assessment framework comprising two complementary evaluation components: the Behavioral Competency Test and the Driving Intelligence Test. The Behavioral Competency Test evaluates the AV's reactive behaviors under controlled scenarios, ensuring basic behavioral competency. In contrast, the Driving Intelligence Test assesses the AV's interactive behaviors within naturalistic traffic conditions, quantifying the frequency of safety-critical events to deliver statistically meaningful safety metrics before large-scale deployment. In Part II of this study, an open-source Level 4 Automated Driving System (ADS) is tested to demonstrate the effectiveness of the proposed method.
Henry X. Liu, Tinghan Wang, Xintao Yan +6
University of Michigan Transportation Research Institute, Ann Arbor, MI 48109 USA · Department of Civil and Environmental Engineering, University of Michigan, Ann Arbor, MI 48109 USA · Department of Civil Engineering, The University of Hong Kong, Hong Kong 999077, China +2
Long-tail scenarios remain a major bottleneck for autonomous driving evaluation, even as datasets grow by orders of magnitude. Existing evaluation pipelines are rarely human-aligned, safety-aware, verifiable, and explainable at the same time: closed-loop metrics often saturate among strong planners, while unstructured human ratings can be noisy without a carefully designed protocol. We formulate planning evaluation as additional-threat detection: given a planner trajectory and an expert reference, does the planner's displacement introduce new unsafe driving behavior? We propose FluidTest, an evaluation pipeline with three components: a pairwise WebUI protocol for reliable human annotation; a taxonomy of 32 semantic threats with evidence-grounded decision graphs; and a three-agent verification system with reflection for precision and auditability. Experiments on the WOD-E2E dataset show that FluidTest produces consistent labels among trained annotators and identifies additional threats in 65% of Poutine trajectories and 51% of RAP trajectories. These results show that state-of-the-art planners can still exhibit substantial safety-relevant failures despite high Rater Feedback Scores (RFS) and low Average Displacement Error (ADE). Additional details, guidance, and code are available at https://fluidtest.web.app.
Qiao Sun, Weicheng Zheng, Yixin Huang +1
Shanghai Qi Zhi Institute · Tongji University · Tsinghua University
Recent end-to-end (E2E) autonomous driving policies achieve high driving scores in closed-loop simulations. Yet it remains unclear whether these policies handle common safety-critical scenarios. We present Safe2Drive (S2D), a set of Bench2Drive-aligned scenario extensions focused on three frequent families of road hazards: work zones, pedestrian jaywalking, and occluded vulnerable road users (VRUs). Safe2Drive adds 100 common but challenging scenarios and introduces SafeDriving Score (SDS), a safety-centric metric that augments prior evaluators with pre-crash braking, work zone-object contact, lane centering, and smoothness checks. Evaluating two state-of-the-art policies (LEAD and SimLingo) on S2D, we find that their driving scores drop sharply relative to their reported Bench2Drive baselines (LEAD: from 94.70 DS on Bench2Drive to 39.95 DS on S2D; SimLingo: from 85.07 DS on Bench2Drive to 41.00 DS on S2D) and that SDS on S2D is low (11.85 for LEAD and 15.27 for Sim-Lingo). These results are consistent with brittle safe-driving behaviors such as poor work-zone understanding, red-light violations, and late or absent braking for pedestrians. This study highlights a lack of safe behavioral reasoning in E2E models even when tested on CARLA towns that are part of the training set. We plan to release the code and videos for all 100 S2D scenarios.
Nishad Sahu, Kalpana Panda, Congyuan Yu +3
1Carnegie Mellon University · 2Birla Institute of Technology and Science Pilani