Pickup trucks account for 14% of new light-duty vehicles produced in the United States, yet are among the least aerodynamic. Their open cargo bed adds a flow absent from existing automotive aerodynamics datasets such as DrivAerML and SHIFT-SUV: the shear layer leaving the cab roof passes over a recirculating bed flow before separating again at the tailgate. The resulting drag lowers fuel efficiency, raises emissions and limits the range of electric trucks. Scale-resolved Computational Fluid Dynamics (CFD) is too costly for broad design exploration; neural surrogates can predict flow features at a fraction of that cost, provided they are trained on large-scale, high-fidelity, domain-specific data. We introduce SHIFT-Truck, the first such dataset for pickup trucks. It comprises 1,000 Spalart-Allmaras delayed detached-eddy simulations (SA-DDES) of a reference pickup geometry morphed across 17 shape parameters. Each case is run on a mesh of about 100 million cells at a Reynolds number of 1.4×107 and released with time-averaged surface pressure, wall shear stress, volumetric pressure and velocity. The setup is verified by grid refinement and repeated runs, and checked against wind-tunnel measurements. We define geometry-grouped splits and benchmark four neural surrogates, DoMINO, GeoTransolver, AB-UPT and SMART, on surface and volume tracks. SHIFT-Truck also introduces controlled distribution shifts in the operating point, the input surface discretization and the vehicle archetype. Models with strong in-distribution performance can degrade substantially under these shifts: operating-condition changes expose failures to infer speed dependence, while tessellation and cross-vehicle shifts reveal markedly different robustness across architectures. SHIFT-Truck is thus a benchmark not only for surrogate accuracy but also for generalization across physical and numerical distributions.
Figures & tables
Figure 1: Design parameters of SHIFT-Truck on the GTU baseline. (a) Side, (b) top and (c) front views. Blue arrows mark translational morphs. Dotted line pairs with an arrow across the open end mark angle morphs. Dotted outlines in (b) show the nose at the two extremes of front planview. Hatched regions are the two on/off parts. Ranges are given in Appendix A .
Figure 2: The baseline GTU on the centreplane y=0 , the vertical plane through the middle of the truck from nose to tail (inset). (a) Mesh. (b) Instantaneous velocity magnitude. (c) Time-averaged velocity magnitude.
Figure 3: The design space of SHIFT-Truck. (a) Drag distribution by bed state over the 712 cases of the benchmark set. (b) The 218 base trucks carrying both bed states with every other parameter held fixed, each plotted as its closed-bed against its open-bed drag, with the 1:1 line. (c) Drag against overall height, coloured by bed state.
Surface track
Volume track
Model
Params
GPU-h
Inf. (s)
Cp MAE
τwεL2
CD MAE
CDR2
Params
GPU-h
Cp MAE
uεL2
GeoTransolver
21.7
91
630
0.021
0.173
0.0069
0.958±0.002
22.2
144
0.029
0.149
SMART
12.6
61
267
0.023
0.175
0.0047
0.971±0.010
12.8
69
0.028
0.114
AB-UPT
16.4
54
883
0.025
0.210
0.0077
0.953±0.010
41.3
89
0.027
0.132
DoMINO
10.1
68
102
0.051
0.605
0.0345
0.034±0.034
18.5
132
0.046
0.163
Table 1: Benchmark results on the 90 test cases. Error metrics are means over cases and seeds 42, 43 and 44. Drag R2 entries give the mean ± the standard deviation across seeds. Params is the trainable parameter count in millions and GPU-h the training cost in H100 hours for 500 epochs in FP32. Inf. is the wall-clock time to predict one case on a surface of 8.8 million faces, on one NVIDIA T4 GPU in float32.
Figure 4: Surface pressure coefficient and wall-shear-stress magnitude for a matched pair of test trucks, the same base geometry with the bed closed (top two rows) and open (bottom two rows). Columns are ground truth and the four surface-track models.
Figure 5: Time-averaged velocity magnitude on the centerplane y=0 for the pair of Fig. 4 , with the bed closed (left) and open (right). Rows are ground truth and the four volume-track models.
Figure 6: Absolute drag error of the four truck-trained models against freestream speed (left) and the fraction of input surface faces retained (right). The shaded column is the training condition, a training-set truck at 140 km/h on its full surface. One count is 10−3 in CD .
Model
Truck
SUV, rescaled
SUV, raw
SMART
1.8
251
214
GeoTransolver
7.1
67
81
AB-UPT
6.1
241
199
DoMINO
34.8
147
54
SHIFT-SUV reference
–
1.5
1.5
Table 2: Absolute drag error on ten SHIFT-SUV cases in counts ( 10−3 in CD ), averaged over the cases. Truck-trained predictions are rescaled by (U/Utrain)2 to the SUV speed in the rescaled column and left as predicted in the raw column. The truck column is each model’s error on the training-set truck at 140 km/h. The reference model is trained on SHIFT-SUV.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Direction
Slider range
Displacement (m)
Baseline
Front overhang
x , forward
−0.50 to +0.75
−0.154 to +0.231
−0.20
Hood height
z
−0.5 to +1.0
−0.046 to +0.092
−0.33
Hood inclination
z at hood front (rotation about the cowl)
−0.5 to +0.5
−0.192 to +0.192
+0.00
Front fascia height
z
−0.5 to +0.5
−0.125 to +0.125
+0.00
Front fascia angle
x at fascia top (rotation about its base)
−0.5 to +1.0
−0.188 to +0.375
−0.33
Front planview
y , corners inboard
−0.5 to +1.0
−0.225 to +0.451
−0.33
Appendix
Table 3: The fifteen continuous design parameters. Each is a shape key on the deformation cage. The sampler draws a normalized value in [−1,1] and maps it linearly onto the slider range. Displacement is the largest cage-vertex displacement at each end of the range, and the last column gives the normalized value of the baseline geometry.
Cases at the freeze
712
closed bed / open bed
492 / 220
Distinct base trucks
248
appearing three, two, one times
218 / 28 / 2
carrying both bed states
219
carrying both bed states, chin spoiler held fixed
218
carrying both spoiler states, bed closed
245
Appendix
Table 4: Composition of the benchmark set, the 712 cases complete at the training freeze, and of the release.
Figure 7: Wall y+ on the baseline GTU with the dataset setup, (a) side and (b) underside. The body panels are wall-modelled at y+ of 20 to 35. The mirrors, the chin spoiler and the wheel edges resolve below 5.
Quantity
Symbol
Value
Freestream speed
U∞
38.89 m s -1
Freestream static temperature
T∞
293.15 K
Freestream density (derived)
ρ∞
1.204 kg m -3
Reynolds number U∞L/ν (derived)
Re
1.4×107
Reference length (body length)
L
5.25 m
Reference area (frontal)
Aref
2.65 m 2
Appendix
Table 5: Operating point, reference quantities and domain, common to every case.
Absolute CD
ΔCD from baseline
Configuration
Ours
Experiment
Ours
Experiment
Reference CFD
4×2 baseline
0.446
0.393
0
0
0
No mirrors
0.435
0.381
−0.0106
−0.012
−0.010
No airdam
0.481
0.421
+0.0349
+0.028
+0.035
Appendix
Table 6: Validation against the wind-tunnel experiment and the CFD of Woodiga et al. (2020) on the GTU baseline and two single-part removals. Deltas are referenced to each source’s own baseline. The reference CFD is reported as deltas only.
Figure 8: Grid refinement on the baseline GTU. (a) Drag coefficient and (b) configuration deltas against cell count for the five-mesh family. Dashed and dotted lines are the experiment and the CFD of Woodiga et al. (2020) . (c)–(e) Time-averaged velocity magnitude on the centreplane for the 80-, 113- and 304-million-cell meshes.
Mesh
Cells (M)
CD
ΔCD mirrors
ΔCD airdam
L1
26
0.4431
+0.0056
+0.0481
L2
80
0.4329
−0.0052
+0.0445
L3
107
0.4287
−0.0070
+0.0445
L4
113
0.4295
−0.0082
+0.0450
L5
304
0.4269
−0.0071
+0.0445
Experiment ( Woodiga et al., 2020 )
—
0.3930
−0.0120
+0.0280
Appendix
Table 7: Grid refinement study on the baseline GTU. Deltas are referenced to each mesh’s own baseline.
Figure 9: Validation on the baseline GTU. (a) Absolute drag coefficient for the three configurations, from the experiment of Woodiga et al. (2020) and the present DDES. (b) Drag change on removing the mirrors and the airdam, from the experiment and CFD of Woodiga et al. (2020) and the present DDES.
Figure 10: Time-averaged pressure coefficient on the centreplane y=0 of the baseline GTU, the same plane and case as Fig. 2 .
Table 8: Training configurations. All models use 500 epochs in FP32. Query counts are per case during training; volume evaluation uses the full mesh for every model. The volume-track rows give the surface and volume query counts used together in that track.
Figure 11: Surface pressure coefficient for the matched pair of Fig. 4 , with the bed closed (top) and open (bottom), in the front-high and side views. Each prediction is shown beside its error, the prediction minus the ground truth. Both bed states share one error scale.
Figure 12: Wall-shear-stress magnitude for the matched pair of Fig. 4 , laid out as in Fig. 11 .
Figure 13: Velocity magnitude (top) and pressure coefficient (bottom) on the centreplane y=0 for the closed-bed truck of the matched pair of Fig. 4 . Each prediction is shown beside its error, the prediction minus the ground truth. Figure 14 uses the same colour scales.
Figure 14: Velocity magnitude (top) and pressure coefficient (bottom) on the centreplane y=0 for the open-bed truck of the matched pair, laid out as in Fig. 13 .
File
Format
Contents
Size
merged_surfaces.stl
binary STL
vehicle surface, triangulated
386 MB
merged_surfaces.vtp
VTK PolyData
time-averaged pressure and wall shear stress, point and cell data
395 MB
merged_volumes.vtu
VTK UnstructuredGrid
time-averaged pressure and velocity on the full mesh
11.1 GB
forces.json
JSON
CD , CL , length of the averaging window
—
metadata.json
JSON
solver run provenance, iteration count, physical time, removed patches
—
params.json
JSON
reference values of Table 5
—
Appendix
Table 9: Files per case, with sizes for variant_0000 .
Motivation
Training and benchmarking of neural surrogates for the external aerodynamics of pickup trucks, a vehicle class absent from open datasets.
Composition
1,000 cases, each a morphed GTU pickup at one operating point, with 15 continuous shape parameters and two parts switched on and off. Each case has a surface with time-averaged pressure and wall shear stress, a volume with time-averaged pressure and velocity, and its drag and lift coefficients. The out-of-distribution cases of § 4 are not included. The data contain no personal information.
Collection
Geometries morphed from GTU configuration 7 with a deformation cage (Appendix A ), then meshed and simulated with DDES (Appendix B ) on 4 H100 GPUs per case. Appendix C validates the setup.
Preprocessing
Fields and forces averaged over the final 10,000 steps. Surface restricted to the vehicle. Instantaneous data discarded. Forces of two cases recovered by integration.
Uses
Surrogate training and benchmarking, studies of shape and drag, and transfer between vehicle classes. Absolute drag carries an offset against the wind tunnel (Appendix C ). The data do not replace simulation or testing in safety-relevant decisions.
Distribution
Hugging Face, https://huggingface.co/datasets/luminary-shift/Truck , with gated access on request. CC-BY-NC-4.0.
Maintenance
Corrections and additions recorded in the change log of the dataset card.
Appendix
Table 10: Datasheet.
Dataset
Class
Cases
Method
Surface
Volume
Experimental validation
AhmedML
bluff body
500
SA-DDES
yes
yes
CD 4.7%, CL 0.6% (baseline)
WindsorML
bluff body
355
WMLES
yes
yes
CD 0.34 vs. 0.33 (baseline); PIV
DrivAerML
passenger car
500
SA-DDES
yes
yes
CD 7.5% (baseline)
DrivAerNet
passenger car
4,000
RANS, k – ω SST
yes
yes
CD 0.81% (baseline)
DrivAerNet++
passenger car
8,000
RANS, k – ω SST
yes
yes
CD 2.2–4.9% (4 baselines)
DrivAerStar
passenger car
12,000
RANS, k – ω SST
yes
yes
CD 1.04% (mean, 3 baselines)
Appendix
Table 11: Open high-fidelity datasets for external vehicle aerodynamics. n/s : not stated in the source. Case counts are those of each release as of September 2026. Validation gives the drag error against the wind tunnel on the baseline geometries; where a source gives coefficients only, the error is computed from them.
Deploying Scientific Machine Learning surrogates in industrial CFD workflows requires adapting pretrained models to new vehicle families without large datasets; yet whether geometric representations learned by a geometry encoder transfer to topologically distinct shapes remains unvalidated. We address this through leave-one-family-out experiments on a 61.47M-parameter Transformer surrogate (AB-UPT) pretrained on four vehicle families (411 external aerodynamics cases) and adapted to the held-out fifth with only 20 samples. Three strategies are compared: Full Fine-Tuning (FFT), Lightweight Fine-Tuning (LFT), and Low-Rank Adaptation (LoRA). The central finding is that pretrained geometry encoders learn transferable representations, but the adaptation mechanism determines whether they can be exploited. FFT destabilizes as 61.47M unconstrained parameters overfit to 20 samples (R^2=0.40); LFT fails because the frozen encoder cannot represent unseen shapes (R^2<0). LoRA resolves both: rank-constrained adapters injected into all layers regularize the loss landscape while preserving pretrained features, achieving R^2=0.85+/-0.02 across all five families with 50% lower force RMSE than FFT and 28% lower pointwise field errors. LoRA also outperforms from-scratch training using 3x more target-family data, eliminating the need for large per-family datasets. These results recast LoRA from a memory-saving convenience into a convergence enabler for geometry transfer: a shared backbone paired with lightweight per-family adapters trainable in hours from minimal data.
Computational Fluid Dynamics (CFD) is central to race-car aerodynamic development, yet its cost -- tens of thousands of core-hours per high-fidelity evaluation -- severely limits the design space exploration feasible within realistic budgets. AI-based surrogate models promise to alleviate this bottleneck, but progress has been constrained by the limited complexity of public datasets, which are dominated by smoothed passenger-car shapes that fail to exercise surrogates on the thin, complex, highly loaded components governing motorsport performance. This work presents three primary contributions. First, we introduce a high-fidelity RANS dataset built on a parametric LMP2-class CAD model and spanning six operating conditions (map points) covering straight-line and cornering regimes, generated and validated by aerodynamics experts at Dallara to preserve features relevant to industrial motorsport. Second, we present the Gauge-Invariant Spectral Transformer (GIST), a graph-based neural operator whose spectral embeddings encode mesh connectivity to enhance predictions on tightly packed, complex geometries. GIST guarantees discretization invariance and scales linearly with mesh size, achieving state-of-the-art accuracy on both public benchmarks and the proposed race-car dataset. Third, we demonstrate that GIST achieves a level of predictive accuracy suitable for early-stage aerodynamic design, providing a first validation of the concept of interactive design-space exploration -- where engineers query a surrogate in place of the CFD solver -- within industrial motorsport workflows.
Nicholas Thumiger, Andrea Bartezzaghi, Mattia Rigotti +5
This paper describes the first-ever open-source high-fidelity CFD dataset of a high-lift aircraft for the purpose of AI surrogate model development. The dataset is composed of 1800 samples, arising from 180 geometry variants and 10 angles of attack for the high-lift NASA Common Research Model (CRM) geometry, used within the AIAA High-Lift Prediction Workshop series. One of the novelties of this dataset is the use of a GPU-accelerated high-fidelity explicit, wall-modeled LES approach for each simulation, using solution-adapted grids between 300M and 500M cells. This ensures the greatest possible accuracy given known challenges in steady-state RANS approaches for these portions of the flight envelope. The entire dataset (geometries, time-averaged volume and surface variables and integral forces) are available, free of charge with a permissive open-source license (CC-BY-4.0). By making this data publicly available, we aim to accelerate the research and development of AI surrogate modeling within the aerospace industry.
Neil Ashton, Adam Clark, Liam Heidt +11
NVIDIA, Santa Clara, US · The Boeing Company, Seattle, US · Cadence Design Systems, Santa Clara, US