Estimating accident mechanics from real-world crashes is important for vehicle-safety analysis, injury modeling, crash-severity prediction, and operational workflows such as insurance claim triage. In standard crash records, key metadata such as impact configuration, principal direction of force, and change in velocity (ΔV) may be missing, delayed, or corrupted, while post-crash photographs are widely available and contain rich visual evidence of deformation. We study how much crash-mechanics information can be recovered directly from vehicle photos when structured signals are absent. We formulate crash understanding as supervised prediction from per-case multi-view photo sets. Targets include six Collision Deformation Classification (CDC) descriptors and the longitudinal and lateral components of reconstructed ΔV. Each photo is encoded by a shared visual backbone, and the resulting view-level features are fused into a case-level representation from which target-specific heads predict crash descriptors. Using 15.2k training cases from the Crash Investigation Sampling System, drawn from about 1.5M photos before filtering, together with 1.15k validation and 1.15k test cases, we define an evaluation protocol for vision-based crash descriptor estimation from incomplete multi-view evidence. Post-crash imagery alone provides usable signal for several non-trivial crash-mechanics descriptors, while weakly observable and long-tailed targets remain challenging. Within the compared training regimes, the selected joint-training recipe reduces mean absolute angular error for principal direction of force from 20.1 to 14.05 degrees and longitudinal ΔV MAE from 8.04 to 7.45 km/h. Our work provides a reference point for future multimodal fusion with structured crash metadata.
Figures & tables
Figure 1: Comparison of the selected joint-training and single-task reference pipelines. Performance is shown as relative change (%) with respect to the corresponding selected single-task result for each target.
Figure 2: Representative crash cases in the fixed multi-view input format. Each case is shown as an ordered subset of available viewpoints together with its available crash descriptor annotations (CDC-derived targets and ΔV when available). The descriptor denoted as ΔV represents the resulting compound vector and is included only for reader interpretation.
Figure 3: Multi-view crash descriptor architecture: a shared SwinV2 encodes nine ordered views, a lightweight fusion transformer aggregates them into a case embedding, and a decoder predicts CDC descriptors and ΔV .
Target
Out
Classes
Miss%
Top1%
Metrics
Impact plane
Classification
4
0.0
64.9
Acc
DoF clock
Circ-Smooth-Cls
12
2.9
46.1
AngErr
Long./lat. zone
Classification
9
0.0
38.3
Acc
Vert./lat. zone
Classification
7
0.0
89.8
Acc
Damage distribution
Classification
8
0.0
75.4
Acc
Deformation extent
Smooth-Cls
9
14.7
43.6
Acc@ ± 1
Table 1: Target summary with missing-label rates (Miss%) and share of the most frequent class among valid labels (Top1%). Circ-Smooth-Cls: circular classification with Gaussian-smoothed targets and KL divergence [ 10 ] . Smooth-Cls: classification with label smoothing. Acc: accuracy. AngErr: mean absolute angular error (degrees). MAE: mean absolute error. Acc@ ± 1: accuracy within one neighboring extent bin of the ground-truth label.
Descriptor
Metric
Single-task
Multi-task
Δ (ST → MT)
Rel. Δ
Impact plane
Acc
0.924
0.922
-0.002
-0.2%
F1macro
0.714
0.720
+0.006
+0.8%
AngErr ( ∘ )
20.10
14.05
-6.05
-30.1%
DoF clock
Acc@ ± 1
0.859
0.877
+0.018
+2.1%
Zone Long./lat.
Acc
0.582
0.602
+0.020
+3.4%
F1macro
0.577
0.532
-0.045
-7.8%
Table 2: Comparison between selected single-task reference models and the selected joint-training recipe. Δ denotes the absolute change from the selected single-task model to the selected joint-training model, and Rel. Δ denotes the corresponding relative change with respect to the single-task result. For AngErr and MAE, lower is better. For Acc/ F1 , higher is better. The comparison should be interpreted as an empirical comparison between selected training regimes, not as an isolated ablation of multi-task supervision.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
ID
Backbone
Views
Fuse
Target/Loss
AngErr ( ∘ ) ↓
P1
Eff-B4
Rand K (rep)
Concat.
sin/cos MSE
44.1∘
P2
Intern-T
Rand K (rep)
Concat.
Circ-smooth KL
44.6∘
P3
Intern-T
Chan. conc.
–
Circ-smooth KL
37.4∘
P4
SwinV2-B
Rand K (perm)
Mean pool
sin/cos MSE
32.9∘
P5
SwinV2-B
Rand K (rep)
Concat.
sin/cos MSE
27.3∘
P6
SwinV2-B
Fixed-9 (pad)
Transf.fuse
Circ-smooth KL
21.8∘
Appendix
Table 1: DoF ablations. Eff-B4: EfficientNet-B4. SwinV2-B: SwinV2-base. Intern-T: InternImage-T. Rand K : random subset of available views. Rep: replication if fewer than K views. Perm: random permutation. Fixed-9: fixed ordered 9-slot input. Pad: zero-padding for missing views and edge padding to preserve image aspect ratio. AngErr: mean absolute angular error (degrees) under circular distance.
Figure 1: Overview of the crash case acquisition and curation pipeline. Starting from publicly available NHTSA CISS cases, we applied data cleaning, metadata-based image filtering, YOLOv8s vehicle filtering, and a dedicated RetinaNet-based wheel-detector filtering stage to obtain the final curated crash-image sets used in this study.
Figure 2: Representative examples removed by the wheel-filtering step. The discarded images are dominated by close-up wheel views and contain limited information about global crash deformation.
Slot construction
Input organization
AngErr ( ∘ ) ↓
Random / unordered
Random available views
25.18
Metadata slots
Fixed 9-slot, CISS metadata
20.10
Auto-orientation slots
Fixed 9-slot, predicted orientation
20.10
Appendix
Table 2: Effect of view-slot construction on DoF prediction. Metadata slots use CISS viewpoint labels and are used in the main experiments to isolate crash-descriptor learning from viewpoint-estimation noise. Auto-orientation slots use an image-based vehicle-orientation estimator trained from the Car Full View Dataset [ 1 ] to assign views to the fixed 9-slot representation. Results are averaged over three runs using the same DoF model and evaluation protocol.
Figure 3: Label distributions in our CISS-based dataset. We show the class frequencies of all CDC-derived targets together with the distributions of ΔVlong and ΔVlat (WinSMASH reconstruction). The compound ΔV vector is included for reader interpretation. The categorical targets show strong long-tail behavior with dominant frontal/central configurations.
Single-task
Multi-task
Epochs
50
50
Learning rate
1×10−5
2/6/30×10−6
LR split
–
backbone / fusion+pre-head / task heads
Warmup / anneal
3 / 7
5 / 20
Batch / accum.
4 / –
1 / 4
Effective batch
4
4
Appendix
Table 3: Representative training configurations for the reported single-task and multi-task results. Both regimes use the same case representation, backbone, and fusion module, but differ in initialization, optimization, and loss balancing. Single-task hyperparameters were tuned per target, so that column is a compact summary rather than one identical configuration for all descriptors. Quant.-strat. rare/update: quantile-stratified update-level sampling with at least one rare sample per effective optimizer update. Task-aware WRS: WeightedRandomSampler using the maximum available inverse-frequency weight across categorical targets. Init. strategy: transfer initializes from a related single-task checkpoint, e.g., impact plane for vertical/lateral finetuning; clock-pretrain + random initializes the shared trunk from DoF clock pretraining and the remaining components randomly.
Work
Input
Target
Acc
Silver et al. [ 25 ]
1 img
LOC (f/b)
0.920
Ours (ST)
≤9 images
CDC (f/r/l/b)
0.924
Ours (MT)
≤9 images
CDC (f/r/l/b)
0.922
Appendix
Table 4: Impact plane image prediction. Silver et al. [ 25 ] evaluate in a curated single-image setting where the background is masked to retain only the vehicle, inputs are limited to front or rear views, and the label space is analogously reduced to front/back location. In contrast, our benchmark uses real-world crash cases represented as multi-view photo sets with natural viewpoint missingness and 4-plane coverage (front/right/left/back).
Work
Input
Scope
Metric
Target
Score [km/h]
Silver et al. [ 25 ]
1 image
R1
MAE
ΔVlong
4.19
Hasija et al. [ 8 ]
1 image, meta
R2
RMSE
ΔVlong
6.38
MAE
(ΔVlong,ΔVlat)
8.04 / 5.31
Ours (ST)
≤9 images
R3
RMSE
(ΔVlong,ΔVlat)
11.96 / 8.20
MAE
(ΔVlong,ΔVlat)
7.45 / 5.09
Ours (MT)
≤9 images
R3
RMSE
(ΔVlong,ΔVlat)
11.26 / 7.25
Appendix
Table 5: ΔV estimation results. Scope codes: R1 (Silver et al. [ 25 ] ) uses a manually curated set (combined from simulation and real world) of small passenger vehicles for front/rear collisions ( <96 km/h), images are grayscale and manually cropped to the vehicle, reporting MAE for ΔVlong . R2 (Hasija et al. [ 8 ] ) uses a curated frontal-only subset and incorporates vehicle metadata (weight, body type, stiffness); we report their main RMSE for ΔVlong on the full test set (6.38 km/h). They additionally report 7.50 km/h on the EDR-verified subset. R3 (ours) evaluates on real-world CISS multi-view photo sets with natural missingness and image-based inputs, we predict component-wise (ΔVlong,ΔVlat) and report MAE and RMSE, with missing labels handled by masking and a 4-way impact plane.