Estimating accident mechanics from real-world crashes is important for vehicle-safety analysis, injury modeling, crash-severity prediction, and operational workflows such as insurance claim triage. In standard crash records, key metadata such as impact configuration, principal direction of force, and change in velocity (ΔV) may be missing, delayed, or corrupted, while post-crash photographs are widely available and contain rich visual evidence of deformation. We study how much crash-mechanics information can be recovered directly from vehicle photos when structured signals are absent. We formulate crash understanding as supervised prediction from per-case multi-view photo sets. Targets include six Collision Deformation Classification (CDC) descriptors and the longitudinal and lateral components of reconstructed ΔV. Each photo is encoded by a shared visual backbone, and the resulting view-level features are fused into a case-level representation from which target-specific heads predict crash descriptors. Using 15.2k training cases from the Crash Investigation Sampling System, drawn from about 1.5M photos before filtering, together with 1.15k validation and 1.15k test cases, we define an evaluation protocol for vision-based crash descriptor estimation from incomplete multi-view evidence. Post-crash imagery alone provides usable signal for several non-trivial crash-mechanics descriptors, while weakly observable and long-tailed targets remain challenging. Within the compared training regimes, the selected joint-training recipe reduces mean absolute angular error for principal direction of force from 20.1 to 14.05 degrees and longitudinal ΔV MAE from 8.04 to 7.45 km/h. Our work provides a reference point for future multimodal fusion with structured crash metadata.
Figures & tables
Figure 1: Comparison of the selected joint-training and single-task reference pipelines. Performance is shown as relative change (%) with respect to the corresponding selected single-task result for each target.
Figure 2: Representative crash cases in the fixed multi-view input format. Each case is shown as an ordered subset of available viewpoints together with its available crash descriptor annotations (CDC-derived targets and ΔV when available). The descriptor denoted as ΔV represents the resulting compound vector and is included only for reader interpretation.
Figure 3: Multi-view crash descriptor architecture: a shared SwinV2 encodes nine ordered views, a lightweight fusion transformer aggregates them into a case embedding, and a decoder predicts CDC descriptors and ΔV .
Target
Out
Classes
Miss%
Top1%
Metrics
Impact plane
Classification
4
0.0
64.9
Acc
DoF clock
Circ-Smooth-Cls
12
2.9
46.1
AngErr
Long./lat. zone
Classification
9
0.0
38.3
Acc
Vert./lat. zone
Classification
7
0.0
89.8
Acc
Damage distribution
Classification
8
0.0
75.4
Acc
Deformation extent
Smooth-Cls
9
14.7
43.6
Acc@ ± 1
Table 1: Target summary with missing-label rates (Miss%) and share of the most frequent class among valid labels (Top1%). Circ-Smooth-Cls: circular classification with Gaussian-smoothed targets and KL divergence [ 10 ] . Smooth-Cls: classification with label smoothing. Acc: accuracy. AngErr: mean absolute angular error (degrees). MAE: mean absolute error. Acc@ ± 1: accuracy within one neighboring extent bin of the ground-truth label.
Descriptor
Metric
Single-task
Multi-task
Δ (ST → MT)
Rel. Δ
Impact plane
Acc
0.924
0.922
-0.002
-0.2%
F1macro
0.714
0.720
+0.006
+0.8%
AngErr ( ∘ )
20.10
14.05
-6.05
-30.1%
DoF clock
Acc@ ± 1
0.859
0.877
+0.018
+2.1%
Zone Long./lat.
Acc
0.582
0.602
+0.020
+3.4%
F1macro
0.577
0.532
-0.045
-7.8%
Table 2: Comparison between selected single-task reference models and the selected joint-training recipe. Δ denotes the absolute change from the selected single-task model to the selected joint-training model, and Rel. Δ denotes the corresponding relative change with respect to the single-task result. For AngErr and MAE, lower is better. For Acc/ F1 , higher is better. The comparison should be interpreted as an empirical comparison between selected training regimes, not as an isolated ablation of multi-task supervision.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
ID
Backbone
Views
Fuse
Target/Loss
AngErr ( ∘ ) ↓
P1
Eff-B4
Rand K (rep)
Concat.
sin/cos MSE
44.1∘
P2
Intern-T
Rand K (rep)
Concat.
Circ-smooth KL
44.6∘
P3
Intern-T
Chan. conc.
–
Circ-smooth KL
37.4∘
P4
SwinV2-B
Rand K (perm)
Mean pool
sin/cos MSE
32.9∘
P5
SwinV2-B
Rand K (rep)
Concat.
sin/cos MSE
27.3∘
P6
SwinV2-B
Fixed-9 (pad)
Transf.fuse
Circ-smooth KL
21.8∘
Appendix
Table 1: DoF ablations. Eff-B4: EfficientNet-B4. SwinV2-B: SwinV2-base. Intern-T: InternImage-T. Rand K : random subset of available views. Rep: replication if fewer than K views. Perm: random permutation. Fixed-9: fixed ordered 9-slot input. Pad: zero-padding for missing views and edge padding to preserve image aspect ratio. AngErr: mean absolute angular error (degrees) under circular distance.
Figure 1: Overview of the crash case acquisition and curation pipeline. Starting from publicly available NHTSA CISS cases, we applied data cleaning, metadata-based image filtering, YOLOv8s vehicle filtering, and a dedicated RetinaNet-based wheel-detector filtering stage to obtain the final curated crash-image sets used in this study.
Figure 2: Representative examples removed by the wheel-filtering step. The discarded images are dominated by close-up wheel views and contain limited information about global crash deformation.
Slot construction
Input organization
AngErr ( ∘ ) ↓
Random / unordered
Random available views
25.18
Metadata slots
Fixed 9-slot, CISS metadata
20.10
Auto-orientation slots
Fixed 9-slot, predicted orientation
20.10
Appendix
Table 2: Effect of view-slot construction on DoF prediction. Metadata slots use CISS viewpoint labels and are used in the main experiments to isolate crash-descriptor learning from viewpoint-estimation noise. Auto-orientation slots use an image-based vehicle-orientation estimator trained from the Car Full View Dataset [ 1 ] to assign views to the fixed 9-slot representation. Results are averaged over three runs using the same DoF model and evaluation protocol.
Figure 3: Label distributions in our CISS-based dataset. We show the class frequencies of all CDC-derived targets together with the distributions of ΔVlong and ΔVlat (WinSMASH reconstruction). The compound ΔV vector is included for reader interpretation. The categorical targets show strong long-tail behavior with dominant frontal/central configurations.
Single-task
Multi-task
Epochs
50
50
Learning rate
1×10−5
2/6/30×10−6
LR split
–
backbone / fusion+pre-head / task heads
Warmup / anneal
3 / 7
5 / 20
Batch / accum.
4 / –
1 / 4
Effective batch
4
4
Appendix
Table 3: Representative training configurations for the reported single-task and multi-task results. Both regimes use the same case representation, backbone, and fusion module, but differ in initialization, optimization, and loss balancing. Single-task hyperparameters were tuned per target, so that column is a compact summary rather than one identical configuration for all descriptors. Quant.-strat. rare/update: quantile-stratified update-level sampling with at least one rare sample per effective optimizer update. Task-aware WRS: WeightedRandomSampler using the maximum available inverse-frequency weight across categorical targets. Init. strategy: transfer initializes from a related single-task checkpoint, e.g., impact plane for vertical/lateral finetuning; clock-pretrain + random initializes the shared trunk from DoF clock pretraining and the remaining components randomly.
Work
Input
Target
Acc
Silver et al. [ 25 ]
1 img
LOC (f/b)
0.920
Ours (ST)
≤9 images
CDC (f/r/l/b)
0.924
Ours (MT)
≤9 images
CDC (f/r/l/b)
0.922
Appendix
Table 4: Impact plane image prediction. Silver et al. [ 25 ] evaluate in a curated single-image setting where the background is masked to retain only the vehicle, inputs are limited to front or rear views, and the label space is analogously reduced to front/back location. In contrast, our benchmark uses real-world crash cases represented as multi-view photo sets with natural viewpoint missingness and 4-plane coverage (front/right/left/back).
Work
Input
Scope
Metric
Target
Score [km/h]
Silver et al. [ 25 ]
1 image
R1
MAE
ΔVlong
4.19
Hasija et al. [ 8 ]
1 image, meta
R2
RMSE
ΔVlong
6.38
MAE
(ΔVlong,ΔVlat)
8.04 / 5.31
Ours (ST)
≤9 images
R3
RMSE
(ΔVlong,ΔVlat)
11.96 / 8.20
MAE
(ΔVlong,ΔVlat)
7.45 / 5.09
Ours (MT)
≤9 images
R3
RMSE
(ΔVlong,ΔVlat)
11.26 / 7.25
Appendix
Table 5: ΔV estimation results. Scope codes: R1 (Silver et al. [ 25 ] ) uses a manually curated set (combined from simulation and real world) of small passenger vehicles for front/rear collisions ( <96 km/h), images are grayscale and manually cropped to the vehicle, reporting MAE for ΔVlong . R2 (Hasija et al. [ 8 ] ) uses a curated frontal-only subset and incorporates vehicle metadata (weight, body type, stiffness); we report their main RMSE for ΔVlong on the full test set (6.38 km/h). They additionally report 7.50 km/h on the EDR-verified subset. R3 (ours) evaluates on real-world CISS multi-view photo sets with natural missingness and image-based inputs, we predict component-wise (ΔVlong,ΔVlat) and report MAE and RMSE, with missing labels handled by masking and a 4-way impact plane.
We present SynCrash, a multi-stage pipeline for zero-shot accident detection, spatial localization, and collision-type classification in fixed-view CCTV surveillance video. Our approach addresses the ACCIDENT at CVPR 2026 Challenge, which requires predicting when an accident occurs, where in the frame the impact happens, and what type of collision it is, all without access to labeled real-world training data. The pipeline operates in three decoupled stages: (1) Temporal localization via a VideoMAEv2-giant backbone fine-tuned on CARLA-based synthetic clips with metadata-aware embeddings and dense sliding-window inference; (2) Spatial localization using YOLO for object detection combined with a physics-informed hybrid heuristic that leverages bounding-box overlap and trajectory-based reasoning to predict the impact point; and (3) Collision-type classification using a lightweight rule-based strategy derived from the number and configuration of detected vehicles. The key insight is that temporal understanding benefits from supervised fine-tuning on synthetic data, whereas spatial understanding is better served by pretrained object detectors and physics priors that transfer naturally across domains.
Traffic surveillance cameras capture accidents continuously, yet converting raw CCTV footage into structured event records that pinpoint when, where, and what type of collision occurred remains unsolved at scale. The ACCIDENT @ CVPR benchmark evaluates exactly this joint prediction under a strict constraint: no labeled real-world training data is available. We introduce a training-free, two-pass coarse-to-fine pipeline that pairs a frozen Qwen3-VL-32B-Instruct vision-language model with YOLO11x object detection and BoT-SORT tracking. A first pass sparsely samples the full clip to anchor the collision moment in time; a second pass re-examines a tight window around that estimate using frames annotated with stable vehicle identities and normalized bounding-box coordinates, which gives the model both a visual overlay and an explicit numeric description of the same scene. On the official 2,027-clip real-CCTV test set, our system achieves a three-way harmonic mean score of 0.504, surpassing all organizer-published baselines including the best multi-model ensemble (0.412) by a 22% relative margin.
Dipit Saha, Shah Mohammad Abdul Mannan, Mohammad Raihan Rashid +2
Bangladesh University of Engineering and Technology
In this paper, we address the problem of zero-shot understanding of accidents from surveillance videos by identifying when an impact event occurs, what type of impact it is, and where in the frame it occurs using natural language. We propose a three-stage pipeline that decomposes the accident understanding into when, what, and where. The first stage extracts a short temporal window around the impact using vision-language similarity. In the second stage, we perform metadata-driven multi-prompt reasoning with five complementary views (baseline, motion, geometry, contrast, and tiebreaker) and resolve disagreement via an entropy-gated pairwise adjudicator. Finally, we localize the impact of an open-vocabulary detector queried on the predicted accident type and scene layout, and aggregate detections across keyframes using a score-weighted centroid. Our pipeline achieves a substantial improvement in the harmonic-mean score over a centre-of-frame baseline on the zero-shot ACCIDENT @ CVPR benchmark. We show that decomposing zero-shot video understanding into temporal localization, semantic classification, and spatial grounding enable more reliable reasoning with vision-language models than direct prompting alone.