Generating corner-case scenarios with appropriate adversity in a simulation environment is critical for testing an autonomous vehicle (AV) software stack's safety performance before deployment. Existing autonomous-driving scenario generators can enforce specific behavior, adversity, or feasibility conditions, but they provide limited control over how extreme a generated scenario is relative to plausible futures in the same traffic context. This study represents the adversity of a generated scenario as its percentile in the conditional distribution of future risk given the observed history. This view supports calibrated answers to two questions: how "corner" a generated corner-case scenario is and how its "cornerness" can be fine-tuned. To this end, we formulate history-conditioned risk-percentile requests and learn a reference risk distribution that maps each requested percentile to a physical risk target. We then use a percentile-conditioned joint diffusion model with sampling-time risk guidance to generate multi-agent futures, together with a reference-based criterion for evaluating percentile realization. Experiments use the minimum post-encroachment time (PET) between the ego and its surrounding vehicles as the risk surrogate on highD. On the primary evaluation set, our method realizes 1,422 of 1,440 requests within a 0.05 percentile tolerance (98.75%), with mean percentile error 0.00673 and PET-target error 0.00991 seconds. The resulting interface connects context-relative risk specification, physical realization, and evaluation through a common risk scale. Project website and videos of generated scenarios are available at https://hhj233.github.io/CornerPercentile/.
Figures & tables
Figure 1 : Why adversity percentiles matter. Existing scenario-generation methods offer control variables that are usually uncalibrated to context-conditional adversity.
Figure 2 : Three-layer workflow: data preprocessing, training and testing. The percentile predictor denotes the reference risk distribution, trained by CRPS on observed scene-risk measurements. Leave-recording-out references supply percentile labels for generator training, and the calibrated reference fitted on all training recordings translates a test-time request into a physical risk target. The request enters the joint generator directly and guides its sampling path, together with road and background-separation constraints.
Figure 3 : Empirical PET CDFs of 11,649 natural clips for histories with 3–5, 6–8 and at least nine vehicles, with group sizes in the legend. The dotted line marks PET =1 second. Group-level outcome variation motivates the contextual scale.
CRPS ↓
Reference
All
N=3 –5
N=6 –8
N≥9
PIT tail↓
Cap Brier ↓
Unconditional ECDF
0.60387
0.75103
0.51421
0.42501
0.03461
0.09969
kNN conditional CDF ( Stone, 1977 )
0.40331
0.46997
0.38083
0.28574
0.02059
0.04256
Quantile regression forest ( Meinshausen, 2006 )
0.33817
0.37500
0.33036
0.26391
0.00402
0.02735
Static-query reference
0.11160
0.15139
0.08866
0.06061
0.04318
0.02003
Range-query reference (ours)
0.10550
0.14041
0.08504
0.06144
0.01051
0.01722
Table 1 : Reference evaluation on 11,649 natural clips. Lower is better. CRPS is in seconds and is given overall and for histories with N=3 –5, 6–8 and at least nine vehicles (5,214, 4,301 and 2,134 clips). PIT tail : largest deviation of the expected CDF of the randomized probability integral transform (PIT) from the diagonal at levels 0.05–0.30, where the targets of the critical requests p≥0.7 lie. Cap Brier: Brier score for PET reaching the four-second cap. Classical estimators use pooled history features without calibration. The static-query reference is our dual-stream architecture with range-independent queries. Bold marks the best value in each column.
Method
Fine (%) ↑
P-MAE ↓
PET-MAE ↓
BG (%) ↓
Ego (%) ↓
Road (%) ↓
External methods without a percentile input
TrafficGen motion ( Feng et al., 2023a )
16.04
0.28665
0.14083
2.08
1.04
52.08
CTG++ unguided prior ( Zhong et al., 2023a )
14.00
0.33789
0.19553
0.21
0.00
21.25
STRIVE traffic prior ( Rempe et al., 2022 )
13.06
0.31757
0.26149
6.67
1.67
75.49
External method with a physical risk input
RADE ( Wang et al., 2025 )
14.31
0.31555
0.20461
7.50
3.89
34.17
Table 2 : Percentile-request realization and scene geometry of the compared generators for 1,440 requests (96 histories, five percentiles, three noise draws). Methods without a percentile input are scored against all five requests of each history. TrafficGen gives one output per history by top-1 decoding, and CTG++ and STRIVE give 15. RADE receives the physical risk target of each request, converted to its risk level. Fine: fine-control rate, the percentage of requests realized within 0.05 of the requested percentile (interval error at most 0.05). P-MAE: mean percentile error. PET-MAE: mean absolute PET-target error (s). BG, Ego and Road: percentages of requests with background–background overlap, ego-related overlap and a strict road violation. Bold marks the best value in each column.
Figure 4 : Where the outputs land. Each column shows, for one requested percentile, the share of outputs whose realized percentile falls in each 0.05-wide bin, on a square-root color scale. An output at a probability mass is placed at the point of its percentile interval nearest the request, as in the interval criterion. Black marks show the request. Natural priors receive no request, so each of their outputs is placed against all five requests.
Figure 5 : Precision across tolerances and selection among samples. (a) Share of the 1,440 requests realized within a percentile tolerance τ . The band around our method is its 95% cluster-bootstrap interval over the 96 histories, the gray band spans the three natural priors, and the dotted line marks the 0.05 tolerance used throughout. (b) Share of requests realized within 0.05 when the sample whose realized percentile is closest to the request is chosen among K samples of the same history, for the two priors with 15 samples per history. Values are exact averages over all subsets of K samples, with 95% cluster-bootstrap bands. The other generators give one output per request, except TrafficGen with one output per history, and appear at K=1 .
Configuration
Fine (%) ↑
P-MAE ↓
PET-MAE ↓
BG (%) ↓
Ego (%) ↓
Road (%) ↓
Our method
98.75
0.00673
0.00991
0.00
0.00
5.76
Condition only
24.79
0.16554
0.11612
0.00
0.49
6.11
Guidance only
95.07
0.01469
0.01688
0.00
0.00
6.18
No P path
16.18
0.29792
0.16246
0.00
0.69
6.25
Table 3 : P-path ablation of our method on the same 1,440 requests. All configurations share the trained weights, reference, histories and noise draws, and road and background constraints are active in every row. Condition only keeps the learned risk-conditioning pathway, Guidance only keeps sampling-time risk guidance, and No P path removes both. Without either pathway, the output does not depend on p , and one future per history and noise draw is scored against all five requests. Metrics as in Table 2 . Bold marks the best value in each column in which the configurations differ.
Figure 6 : P-path ablation over the 1,440 requests. (a) Realized percentile against requested percentile for Our method, Condition only, Guidance only and No P path. Shaded bands span the mean lower and upper ends of the compatible percentile interval, and lines mark the center of each band. (b) Cumulative distribution of the interval error, with a dotted line at the 0.05 tolerance. (c) Fine, the percentage of requests realized within the 0.05 tolerance. Together, the panels show the dominant precision gain from guidance and the further improvement from learned conditioning.
Figure 7 : Physical targets and realized percentiles of the four configurations of Table 3 over the 1,440 requests. (a) At each request, the distribution of the physical targets qH(p) (lines) and of the realized minimum ego–SV PET (filled), pooled over the 96 histories and three noise draws. The last bin collects outcomes at the 4 s cap. (b) The realized percentile of each output at each request in 0.05-wide bins, placed at the point of its percentile interval nearest the request as in Fig. 4 . Dotted lines mark the request. No P path generates one future per history and noise draw and scores it against all five requests.
Figure 8 : Lane changes and car following generated by our method. Each block shows the observed history (the last 0.96 seconds, at a larger scale) and the futures generated from it with one noise draw at p=0.1,0.5,0.9 , and each row reports the realized PET, its target and the percentile error. (a) A seven-vehicle history in which the ego changes lanes at all three requests. (b) An eight-vehicle history in which the PET witness at p=0.9 changes lanes. (c) A ten-vehicle history in which the ego follows its PET witness, the lead vehicle in the same lane, at all three requests. Orange marks the ego, blue the PET witness of each future and gray the other vehicles. Observed histories are dashed and generated futures solid. Vehicle bodies are drawn at the last observed state in the history views and at 6.96 seconds in the future rows, and dots mark positions at the start of the future. All vehicles are shown with their measured dimensions, and road views use equal x and y scales. The examples were selected to illustrate these interactions.
Figure 9 : Direct percentile conditioning and PET-based selection from external priors on the car-following history of Fig. 8 (c). (a) Our method at p=0.1,0.5,0.9 , the futures shown in Fig. 8 (c). (b) TrafficGen motion: top-1 output, generated without a percentile input. (c) CTG++ unguided prior and (d) STRIVE traffic prior: for each target, the sample whose PET is closest to that target among the prior’s 15 samples, selected after generation for display. A sample that is closest to several targets appears in each of their rows. (e) PET of all samples, with filled markers for the displayed samples and dotted lines at the targets. The top panel shows the observed history at a larger scale. The other road views share one extent, and all road views use equal x and y scales. Vehicle colors follow Fig. 8 .
Futures
W1 speed
W1 ax
W1 ay
Harsh braking (%)
Implausible (%)
Observed futures
–
–
–
0.12
0.01
TrafficGen motion
0.323
0.445
0.097
0.00
0.00
CTG++ unguided prior
1.129
0.375
0.130
0.00
0.00
STRIVE traffic prior
0.105
0.583
0.185
0.00
0.00
RADE
0.225
0.311
0.086
0.05
0.00
P-CVAE
0.189
0.357
0.103
0.00
0.00
Table 4 : Kinematic realism of generated futures on the 96 evaluation histories. Speed and longitudinal and lateral acceleration are finite differences of positions on a 0.2 s grid, pooled over all vehicles and times. W1 is the Wasserstein-1 distance to the observed futures of the same histories (m/s for speed, m/s 2 for acceleration). Harsh braking: longitudinal acceleration below −4 m/s 2 . Implausible: acceleration magnitude above 8 m/s 2 . Bold marks the smallest distance in each column.
Configuration
Fine (%)
P-MAE
TTC-MAE
Cont. Fine (%)
Atom Fine (%)
Generator trained on all clips
No guidance
17.22
0.30670
3.479
10.56
34.07
TTC guidance
98.61
0.00532
0.057
98.55
98.77
Generator trained on car-following clips
Full
98.82
0.00499
0.060
98.64
99.26
Condition only
21.60
0.27589
3.184
11.24
47.79
Table 5 : Percentile requests on minimum TTC for pure car-following histories: 96 evaluation histories (32 per vehicle-count group), five percentiles and three noise draws (1,440 requests, 408 with a target at an atom of the bounded TTC reference, 330 at TTC 0 and 78 at the 20 s cap). The upper block uses the generator of the main experiments without retraining. The lower block uses a generator trained only on car-following clips with TTC percentile labels, where Full combines conditioning and guidance. Guidance steers sampling toward the target of the bounded TTC reference. TTC-MAE: mean absolute TTC-target error (s). Cont. and Atom: Fine for requests whose target lies in the continuous part of the reference or at an atom. Bold marks the best value in each column.
Scenarios
Request
Rear coll. (%)
Hard brake (%)
PET <1 s (%)
Median PET (s)
Rear TET (s)
Observed futures
–
1.04
3.12
44.79
1.128
0.048
Our method
p=0.1
0.00
3.82
39.58
1.210
0.008
p=0.3
0.69
3.47
42.71
1.178
0.031
p=0.5
1.04
3.82
44.44
1.161
0.083
p=0.7
1.39
4.51
45.14
1.144
0.098
p=0.9
14.58
12.85
49.65
1.030
0.578
Table 6 : An IDM and MOBIL planner tested against generated SV behavior on the 96 evaluation histories. The planner replaces the ego from its observed state at the start of the future and drives for 6.96 s, while the SVs replay their futures without reacting. Our method receives percentile requests, and RADE receives three fixed physical risk levels, given by their target PET. Rear coll.: an SV reaches the planner from behind. Hard brake: planner deceleration beyond 4 m/s 2 . PET: minimum ego–SV occupancy PET of the planner future. Rear TET: mean time the planner spends with a TTC below 3 s to the SV behind it in its lane ( Minderhoud and Bovy, 2001 ) . Each generated row covers 288 scenarios (96 histories, three noise draws).
Figure 10 : How difficulty unfolds in the planner test, for the 288 scenarios of each request and the 96 observed futures. (a) Share of scenarios in which an SV behind the planner in its lane approaches with a TTC below 3 s. (b) Share of scenarios in which the planner has braked harder than 4 m/s 2 so far. Hard braking at the start of the window comes from the shared histories.
Figure 11 : Planner test on two histories, drawn as in Fig. 9 . At p=0.9 , the encounter that sets the planner’s minimum PET is (a) a follower closing in and (b) a lane change behind the planner. The IDM and MOBIL planner replaces the ego, and the SVs replay their generated futures without reacting. Each block shows the history and noise draw whose planner PET falls the most from p=0.1 to p=0.9 among the collision-free scenarios with this encounter type at p=0.9 . Top left: simulation view behind the planner at p=0.9 when the second of the two vehicles reaches the shared point, with the vehicles colored as in the road views. Top right: planner PET against the request for the three noise draws of the history, with the observed future as a dotted line. Below: road views at p=0.1 , 0.5 and 0.9. Each row gives the planner’s minimum PET and the type of the encounter that sets it, and the cross marks its shared point. Vehicle bodies are drawn at the end of the window and dots at the start of the future. Each block has its own extent, with equal x and y scales.
Figure 12 : Planner test on two further histories, shown as in Fig. 11 . At p=0.9 , the encounter that sets the planner’s minimum PET is (c) a cut-in ahead of the planner and (d) a lane change of the planner.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Model
CRPS
PIT tail
PIT max
Wide MLP
0.10846
0.03490
0.08911
GRU
0.11375
0.03767
0.11007
Absolute histories only
0.10349
0.01996
0.03342
Relative histories only
0.11631
0.03294
0.05560
Static-query reference
0.11160
0.04318
0.05024
Range-query reference (ours)
0.10550
0.01051
0.07915
Appendix
Table 7: Reference encoders on the primary evaluation set (11,649 clips from 16 recordings). CRPS (seconds) of the calibrated models, with PIT tail and PIT max defined in the text. The range-query reference is the one used in all experiments (ours), and the static-query reference reads the same histories with one static query. All other encoders use range queries. Lower is better in every column. Bold marks the lowest value in each column.
Evaluation set
Outputs
Ego center
Ego complete
Witness center
Witness complete
Primary
1440
7
0
74
10
Additional
1440
31
17
70
17
Single recording
1050
25
8
63
9
Appendix
Table 8: Lane changes of the ego and of the PET witness in futures generated by our method. Center: the vehicle’s center crosses an internal lane boundary. Complete: a completed lane change as defined in the text. The PET witness is determined separately for each output.
Figure 13 : External priors on the car-following history of Fig. 8 (c), without post-selection. The top view shows the observed history at a larger scale. (a) Our method at the requested p=0.1,0.5,0.9 , as in Fig. 9 . (b) TrafficGen motion, top-1 output without a percentile input. (c) CTG++ unguided prior and (d) STRIVE traffic prior, each showing the first three of its 15 samples. Each row reports the scene PET with the target and percentile error (our method) or the reference percentile of the sample (priors). (e) PET of all samples (Our method 3, TrafficGen 1, CTG++ 15, STRIVE 15), with filled markers for the displayed samples and dotted lines at the targets.
Safety-critical scenarios are essential for the development of autonomous vehicles (AVs) but are rare in real-world driving data. While simulation offers a way to generate such scenarios, manually designed test cases lack scalability, and adversarial optimization often produces unrealistic behaviors. In this work, we introduce a conditional latent flow matching approach for scalable and realistic safety-critical scenario generation. Our method uses distribution matching to transform nominal scenes into safety-critical rollouts. Furthermore, we demonstrate that incorporating both simulation and real-world data enables our framework to efficiently generate diverse, data-driven scenarios. Experimental results highlight that our approach is able to more consistently and realistically generate novel safety-critical scenarios, making it a valuable tool for training and benchmarking AV systems.
Zimu Gong, Brian Zhaoning Zhang, Chris Zhang +2
University of Michigan-Ann Arbor · University of Waterloo · 1Waabi Innovation Inc +1
Safety-critical scenarios are central to evaluating autonomous driving systems, yet their rarity in naturalistic logs makes simulation-based stress testing indispensable. Most scenario generation methods treat surrounding agents as adversaries, but they either (i) induce failures without explicitly modeling vehicle-road physical limits, yielding visually extreme yet physically unsolvable crashes, or (ii) enforce physical feasibility or policy feasibility in isolation, which can over-focus on aggressive maneuvers or remain tied to a controller-dependent capability boundary. We propose ScenePilot, a feasibility-guided, boundary-driven framework that targets the boundary band: scenarios that are physically solvable in principle yet still cause the deployed autonomy stack to fail. We formulate generation as constrained multi-objective reinforcement learning, combining an RSS-derived physical-feasibility score σ with an online-learned AV-risk predictor Φ, and introduce step-level feasibility-aware shielding to keep exploration near the feasibility boundary while avoiding infeasible artifacts. Experiments on SafeBench with multiple planners show that ScenePilot yields substantially higher collision rates (+6.2 percentage points) while preserving physical validity, and that adversarial fine-tuning on these boundary-band scenarios consistently reduces downstream crash rates. The code is available at https://github.com/QiyuRuan/ScenePilot.
Qiyu Ruan, Yuxuan Wang, He Li +2
State Key Laboratory of Internet of Things for Smart City (SKL-IOTSC), University of Macau, Macau, China · Faculty of Science and Technology, University of Macau, Macau, China
Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize corner cases with photorealistic observations. Corner-case generation is inherently a multi-source problem spanning visual representation, scene reasoning, and vehicle trajectory generation and control. Prior knowledge- and model-based approaches typically focus on scene or trajectory components in isolation, while diffusion-based methods attempt end-to-end generation but still struggle to ensure spatiotemporal consistency and physical realism. To unify these aspects within a single framework, we propose CARLA-GS, a modular corner-case synthesis pipeline that decouples visual representation, semantic reasoning, and physics-based execution while maintaining tight cross-module coupling. Starting from real driving data, we reconstruct an editable gaussian scene with additional geometry-consistent constraints. A multi-agent LLM then performs scene-level reasoning to identify risky interactions and generate intent-level waypoint trajectories, while the low-level motion control is delegated to CARLA, where a PID controller ensures kinematic and dynamic feasibility. The simulated vehicle states are finally re-projected into the gaussian scene for ego-centric rendering. This design enables high-level semantic reasoning, low-level physically executable motion, and photorealistic corner-case generation within a unified pipeline. Experiments on the Waymo Open Dataset show, both quantitatively and qualitatively, that our framework enables controllable corner-case generation and produces photorealistic, spatiotemporally consistent videos aligned with semantic intent and physically feasible motion.
Kaicong Huang, Meng Ma, Ruimin Ke
Department of Civil and Environmental Engineering, Rensselaer Polytechnic Institute, 110 Eighth Street, Troy, NY USA 12180.