VehDyn: A Driving World Model Benchmark for Vehicle Dynamics
Abstract
Video world models are emerging as data engines, action planners, and generative simulators for autonomous driving, but existing benchmarks primarily assess visual fidelity and coarse physical plausibility, providing limited evidence on whether generated driving futures obey realistic vehicle kinematics and dynamics. This limitation is further compounded by the lack of datasets in which vehicle, road, maneuver, and speed conditions are independently controlled, and ground-truth vehicle states are recorded in synchrony with videos. We introduce VehDyn, a driving world model benchmark for vehicle dynamics. VehDyn is built on a CARLA-CarSim co-simulation platform where photorealistic rendering is coupled with a validated multi-body dynamics model, and it contains 10,080 configurations from a full factorial design over five vehicle types, four tire-road friction coefficients, three maneuvers, four target speeds, 14 scenes, and three illuminations, each paired with synchronized position, velocity, and attitude sequences. Built on this dataset, VehDyn introduces a hierarchical evaluation framework that measures trajectory alignment, kinematic consistency, and dynamic consistency, and benchmarks 12 state-of-the-art video world models. We further assess the video quality using two established protocols and correlate it with the VehDyn score. Trajectory-level metrics are nearly saturated, with ten of twelve models within 20% of ground truth, while no model reaches 92% of ground truth on dynamic consistency, and visual-quality metrics are only weakly correlated with vehicle-dynamics fidelity. DrivingWorld achieves the highest VehDyn score, followed by Cosmos 3 Nano and LTX-Video 2.5, and the VehDyn score agrees closely with human judgment. VehDyn provides a systematic foundation for developing driving world models that are physically consistent and visually realistic.
Figures & tables
| Benchmark | Venue & Year | Data | World Model Types | Vehicle Motion Evaluation | Human | |||
| General | Driving | Trajectory | Kinematics | Dynamics | ||||
| ACT-Bench | ICLRW ’25 | nuScenes | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ |
| WorldSimBench | ICML ’25 | CARLA | ✓ | ✗ | ✗ | ✗ | ✓ | |
| DrivingGen | ICLR ’26 | Real-World | ✓ | ✓ | ✓ | ✗ | ✓ | |
| WorldLens | CVPR ’26 | nuScenes | ✗ | ✓ | ✓ | ✓ | ||
| What-If World | arXiv ’26 | nuScenes | ✓ | ✗ | ✗ | ✗ | ✓ | |
| VehDyn Metrics Models | Trajectory Alignment | Kinematic Consistency | Dynamic Consistency | Rank | ||||||||
| All | Emergency-Braking | Single-Lane-Changing | All | All | ||||||||
| ADE | FDE | Pitch | Roll | Yaw | FCCR | CDC Dir | CDC Mag | CDC Trend | ||||
| Ground Truth | 0.514 | 0.814 | 0.897 | 2.522 | 0.697 | 0.602 | 3.885 | 0.900 | 0.722 | 0.713 | 0.803 | – |
| General World Models | ||||||||||||
| LTX-Video 2.5 | 1.665 | 2.310 | 3.217 | 2.077 | 0.932 | 0.558 | 3.330 | 0.686 | 0.563 | 0.610 | 0.539 | 3 |
| CogVideoX 1.5 | 2.804 | 5.368 | 5.683 | 3.018 | 1.270 | 0.971 | 3.543 | 0.738 | 0.582 | 0.557 | 0.565 | 6 |
Appendix figures & tables39 assets
Supplementary material from the paper’s appendix.
Appendix
| Factor | Definition | Unit | Levels | Number |
| Dynamic Factors | ||||
| Vehicle type | Ego vehicle model | – | Cooper S; Model 3, Grand Tourer; T2, Cybertruck | 5 |
| Friction coefficient | Tire-road friction coefficient | – | Asphalt: 0.85 (dry), 0.70 (wet); stone: 0.60 (dry), 0.50 (wet) | 4 |
| Driving maneuver | Prescribed ego maneuver | – | Single-lane-changing (left), single-lane-changing (right), emergency-braking | 3 |
| Target speed | Cruise speed before maneuver | km/h | 20, 54, 72, 108 | 4 |
| Visual Factors | ||||
| Parameter | Unit | Cooper S | Model 3 | Grand Tourer | T2 | Cybertruck |
| Mass and Geometry | ||||||
| Sprung mass | kg | 1220 | 1445 | 1370 | 1250 | 1800 |
| Length | m | 3.81 | 4.79 | 4.61 | 4.48 | 6.27 |
| Width | m | 1.97 | 2.16 | 2.24 | 2.07 | 2.39 |
| Height | m | 1.48 | 1.49 | 1.67 | 2.04 | 2.10 |
| Wheelbase | m | 2.13 | 2.68 | 2.58 | 2.51 | 3.20 |
| Category | Variable | Definition | Source | Unit | Setting |
| Vehicle pose | Position of the vehicle reference point in the CARLA world frame | CarSim | m | 30 Hz | |
| Roll, pitch, yaw of the body | CarSim | deg | 30 Hz | ||
| Vehicle state | Planar velocity recorded in the world frame and rotated into the body frame | CarSim | m/s | 30 Hz | |
| Yaw rate about the body | CarSim | deg/s | 30 Hz | ||
| Throttle, brake, and steering commands issued by the controller | CARLA | – | Throttle and brake in , steering in | ||
| Scenario | Route progress | Distance traveled along the reference path | CARLA | m | 30 Hz |
| Method | Architecture | Conditioning | Parameters | Resolution | Frames | Inference (s/video) |
| General World Models | ||||||
| LTX-Video 2.5 | DiT | First frame + Text | 35.57 B | 1280 704 | 121 | 173.9 |
| CogVideoX 1.5 | 3D DiT | First frame + Text | 10.55 B | 1280 720 | 81 | 1082.6 |
| SkyReels-V2-I2V | 3D DiT | First frame + Text | 20.26 B | 1280 720 | 77 | 1875.6 |
| Wan 2.2-I2V | 3D DiT | First frame + Text | 11.39 B | 1280 704 | 77 | 246.2 |
| HunyuanVideo 1.5-I2V | MMDiT | First frame + Text | 17.98 B | 1280 720 | 77 | 2130.6 |
| Method | Vision Encoder / Depth Prior | Feature Dimensions | Geometric Module | Trajectory Estimation |
| DPVO | Two 4-block CNN encoders | 128 matching; 384 context | Patch graph + differentiable BA | Recurrent sparse patch matching with joint pose/depth refinement |
| DrivingGen | SIFT + UniDepthV2 | 128-D SIFT descriptor; dense metric depth | RANSAC + PnP | Relative-pose chaining; constant-velocity extrapolation with orientation perturbation |
| GEM | DROID learned CNN + Depth Anything V2 | 128 matching; 256 context; dense depth | Recurrent dense update + dense BA | Dense correspondence optimization with depth-assisted scale estimation |
| Maneuver | Response | Unit | ||
| Emergency-braking | Braking distance | 29.173 | 1.459 | m |
| Braking duration | 1.469 | 0.073 | s | |
| Peak pitch excursion | 1.248 | 0.062 | deg | |
| Peak pitch rate | 3.384 | 0.169 | deg/s | |
| Single-lane-changing | Peak lateral acceleration | 1.872 | 0.094 | |
| Peak yaw excursion | 1.639 | 0.082 | deg |
| Metric | Strategy | ||
| ADE | VehDyn empirical | 0.137 | 23.912 |
| FDE | VehDyn empirical | 0.061 | 49.116 |
| VehDyn empirical | 0.249 | 20.324 | |
| Pitch | VehDyn empirical | 0.525 | 9.854 |
| VehDyn empirical | 0.030 | 22.919 | |
| Roll | VehDyn empirical | 0.042 | 5.386 |
| Methods | All Maneuvers | Emergency-Braking | Single-Lane-Changing | ||||
| ADE | FDE | RMSE | Pitch RMSE | RMSE | Roll RMSE | Yaw RMSE | |
| DPVO | 0.514 (0.301) | 0.814 (0.346) | 0.897 (0.469) | 2.522 (2.162) | 0.697 (0.521) | 0.602 (0.468) | 3.885 (3.813) |
| DrivingGen | 2.202 (1.591) | 4.536 (2.238) | 1.572 (1.248) | 8.660 (6.726) | 3.552 (2.736) | 9.141 (5.250) | 9.095 (6.823) |
| GEM | 0.589 (0.283) | 1.071 (0.346) | 0.982 (0.288) | 2.765 (2.037) | 1.045 (0.661) | 0.624 (0.496) | 3.888 (3.863) |
| Models | Trajectory Alignment | Kinematic Consistency | Dynamic Consistency | Rank | |||||
| ADE | FDE | Pitch | FCCR | CDC Dir | CDC Mag | CDC Trend | |||
| Ground Truth | 0.471 | 0.584 | 0.897 | 2.522 | 0.890 | 0.740 | 0.780 | 0.716 | – |
| General World Models | |||||||||
| LTX-Video 2.5 | 1.884 | 2.362 | 3.217 | 2.077 | 0.623 | 0.649 | 0.675 | 0.711 | 6 |
| CogVideoX 1.5 | 3.642 | 7.462 | 5.683 | 3.018 | 0.846 | 0.699 | 0.633 | 0.734 | 9 |
| SkyReels-V2-I2V | 4.415 | 9.118 | 7.243 | 2.415 | 0.658 | 0.692 | 0.654 | 0.736 | 10 |
| VehDyn Metrics Models | Trajectory Alignment | Kinematic Consistency | Dynamic Consistency | Rank | ||||||
| ADE | FDE | Roll | Yaw | FCCR | CDC Dir | CDC Mag | CDC Trend | |||
| Ground Truth | 0.586 | 1.103 | 0.812 | 0.652 | 4.275 | 0.884 | 0.711 | 0.650 | 0.853 | – |
| General World Models | ||||||||||
| LTX-Video 2.5 | 1.365 | 1.837 | 0.939 | 0.547 | 3.751 | 0.723 | 0.526 | 0.586 | 0.538 | 1 |
| Wan 2.2-I2V | 2.754 | 5.010 | 1.909 | 0.657 | 3.746 | 0.610 | 0.510 | 0.553 | 0.642 | 3 |
| SkyReels-V2-I2V | 1.862 | 2.225 | 1.114 | 0.710 | 3.771 | 0.528 | 0.505 | 0.556 | 0.486 | 4 |
| VehDyn Metrics Models | Trajectory Alignment | Kinematic Consistency | Dynamic Consistency | Rank | ||||||
| ADE | FDE | Roll | Yaw | FCCR | CDC Dir | CDC Mag | CDC Trend | |||
| Ground Truth | 0.485 | 0.756 | 0.581 | 0.552 | 3.497 | 0.925 | 0.705 | 0.665 | 0.841 | – |
| General World Models | ||||||||||
| LTX-Video 2.5 | 1.781 | 2.711 | 0.925 | 0.568 | 2.938 | 0.698 | 0.505 | 0.550 | 0.483 | 2 |
| Wan 2.2-I2V | 2.711 | 4.975 | 1.187 | 0.658 | 2.808 | 0.605 | 0.526 | 0.540 | 0.582 | 4 |
| CogVideoX 1.5 | 2.772 | 4.996 | 1.164 | 1.000 | 2.951 | 0.681 | 0.520 | 0.486 | 0.528 | 7 |
| Models | Trajectory Alignment | Normalized Score | Rank | |
| All Maneuvers | ||||
| ADE | FDE | |||
| Ground Truth | 0.984 | 0.985 | 0.984 | – |
| General World Models | ||||
| LTX-Video 2.5 | 0.936 | 0.954 | 0.945 | 1 |
| SkyReels-V2-I2V | 0.886 | 0.904 | 0.895 | 5 |
| Models | Kinematic Consistency | Normalized Score | Rank | ||||
| Emergency-Braking | Single-Lane-Changing | ||||||
| Pitch | Roll | Yaw | |||||
| Ground Truth | 0.968 | 0.786 | 0.971 | 0.895 | 0.617 | 0.847 | – |
| General World Models | |||||||
| LTX-Video 2.5 | 0.852 | 0.834 | 0.961 | 0.903 | 0.674 | 0.845 | 2 |
| SkyReels-V2-I2V | 0.652 | 0.797 | 0.952 | 0.853 | 0.678 | 0.786 | 8 |
| Models | Dynamic Consistency | Normalized Score | Rank | |||
| All Maneuvers | All Maneuvers | |||||
| FCCR | CDC Dir | CDC Mag | CDC Trend | |||
| Ground Truth | 0.900 | 0.722 | 0.713 | 0.803 | 0.784 | – |
| General World Models | ||||||
| CogVideoX 1.5 | 0.738 | 0.582 | 0.557 | 0.565 | 0.610 | 4 |
| Wan 2.2-I2V | 0.595 | 0.571 | 0.602 | 0.643 | 0.603 | 6 |
| Models | VehDyn Benchmark | VehDyn Score | Rank | ||
| Trajectory Alignment | Kinematic Consistency | Dynamic Consistency | |||
| Ground Truth | 0.984 | 0.847 | 0.784 | 2.616 | – |
| General World Models | |||||
| LTX-Video 2.5 | 0.945 | 0.845 | 0.599 | 2.389 | 3 |
| CogVideoX 1.5 | 0.890 | 0.777 | 0.610 | 2.277 | 6 |
| SkyReels-V2-I2V | 0.895 | 0.786 | 0.580 | 2.261 | 7 |
| Other Benchmarks Models | WorldScore | DrivingGen | Rank | |||||
| Quality | Distribution | Temporal Consistency | ||||||
| 3D Consistency | Photometric Consistency | Style Consistency | Subjective Quality | FVD Score | Video Consistency | Trajectory Consistency | ||
| Ground Truth | 0.7932 | 0.1988 | 0.9471 | 0.4284 | 1.0000 | 0.9102 | 0.4469 | – |
| General World Models | ||||||||
| LTX-Video 2.5 | 0.7551 | 0.2736 | 0.9446 | 0.2999 | 0.9196 | 0.9156 | 0.4709 | 1 |
| SkyReels-V2-I2V | 0.7190 | 0.6036 | 0.8878 | 0.4085 | 0.7431 | 0.9094 | 0.5130 | 2 |
| Score | TA: does the vehicle follow the reference path? | KC: does the vehicle move like the reference vehicle? | DC: does the motion respond to the changed condition like the reference? |
| 1.0 | Same lane position and heading throughout; stops (braking) or settles in the new lane (lane-changing) at the same point as the reference, within about one vehicle length. | Deceleration or lateral shift begins and ends at the same times as the reference; nose dip, body roll, and heading change match in magnitude and timing; no jitter or implausible sway. | Between the two clips, the generated braking distance, lane-changing duration, and body attitude change in the same direction as the reference and by a visibly comparable amount. |
| 0.8–0.9 | Same path with a small lateral offset, or a stopping/settling point one to two vehicle lengths from the reference. | Correct sequence and magnitude of motion with a small timing offset, or a slightly under- or over-stated attitude response. | Correct direction of change for every response; the magnitude of the main response is somewhat under- or over-stated. |
| 0.6–0.7 | Maneuver completed on a recognizably similar path, but with a lateral drift of up to about half a lane width or a stop two to four vehicle lengths early or late. | Maneuver recognizable and correctly ordered, but the speed profile or one attitude response is clearly wrong in magnitude. | Correct direction of change for the main response, but the magnitude is clearly wrong or a secondary response changes in the wrong direction. |
| 0.4–0.5 | Maneuver completed, but the path drifts by about a lane width or the stop is several vehicle lengths early or late. | Maneuver recognizable, but the motion is visibly jerky, or two or more responses are wrong in magnitude or timing. | The response changes in the right direction for some quantities and not for others, or in the right direction but with negligible magnitude. |
| 0.2–0.3 | Maneuver only partly executed (incomplete lane-changing, braking that does not stop the vehicle), or the path diverges over most of the clip. | Motion largely inconsistent with the reference (no deceleration where the reference brakes hard; roll or heading changes that the reference does not show). | The two generated clips are nearly identical although the reference clips differ clearly, or the change is in the wrong direction for the main response. |
| 0.0–0.1 | Wrong maneuver, vehicle leaves the road, or the ego vehicle is not identifiable. | Motion physically implausible (teleportation, reversal, the vehicle body deforms) or absent. | The generated response changes in the opposite direction to the reference for the main response, or the clips are unusable for comparison. |