Video generation for autonomous driving cannot follow the web-scale route: driving data is expensive to collect, bound by privacy requirements, and cannot be scraped at will, so models must make the most of a fixed corpus. We present a systematic scaling-law study of video diffusion models trained from scratch on driving data: a family of models from 1M to 9B parameters, trained at different exposures on up to 5,500 hours of driving. Validation loss follows consistent power laws in both model size and training exposure, answering the questions that shape a training budget: whether compute is better spent on longer training or on a larger model, and whether more data is needed. Loss improves much faster with training exposure than with model size, making longer training the most effective way to improve a fixed model under limited compute. However, larger models continue to achieve lower asymptotic loss, so compute-optimal scaling still favors increasing model size when sufficient compute and data are available. Guided by these laws, we train a 9B-parameter model, to our knowledge the largest video diffusion model trained from scratch on driving data: it sets a new open-source state of the art for driving video generation, as measured on nuScenes. Our code and pretrained models are available at https://github.com/valeoai/VATIX. NATIX is separately releasing the underlying driving data in stages.
Figures & tables
Model
Pico
Nano
Micro
Tiny
Small
Medium
Base
Big
Large
XLarge
1B
9B*
Params (M)
1.6
3.8
9.2
19.7
33.8
63.3
134.9
295.5
447.3
643.1
1142.9
9096.4
TFLOPs
0.02
0.05
0.14
0.30
0.61
1.10
2.47
5.61
8.79
12.99
23.09
143.84
Table 1 : Model size and training compute. * denotes the target model size. Training TFLOPs are estimated with thop as three times the FLOPs of a forward pass, approximating one training step (one forward and two backward-pass equivalents).
Figure 1 : Model architecture. Video latents from the Wan encoder are noised and processed by N transformer blocks, each containing spatial attention, temporal attention, and an MLP, modulated by AdaLN. The timestep t and ego-trajectory waypoints are embedded with sinusoidal features, and each is passed through an MLP. They are then summed and chunked into scale and shift. The output is the velocity (v-prediction).
Figure 2 : Model scaling analysis. Validation loss consistently decreases with increasing model size, and the fitted scaling law extrapolates this trend to larger models.
Figure 3 : Training dynamics per model size. Solid curves show measured validation loss during training; dashed curves show the corresponding per-model power law fit, extrapolated across the full exposure range shown; diamonds mark the best measured loss for each model. The dynamics are consistent across scales and well captured by per-model fits, while larger models converge to lower loss.
Figure 4 : Data restriction ablation. Validation loss at fixed training exposure (color) as the pool of unique footage available to the Base model is restricted from the full 5,500-hour corpus down to 5.5 hours. Loss degrades sharply only once repetition exceeds about 2,000 epochs (5.5h restriction); below that, additional unique footage brings diminishing returns for the Base model at the exposures tested here.
Figure 5 : Compute scaling analysis. Lower envelope of validation loss versus compute (a), and iso-FLOP optimal (N,D) configurations for each budget (b).
Figure 6 : Training-scaling fit for the 9B model. Validation loss during training and the corresponding power law fit ( L(D)=0.074799+0.00381D−0.7309 ), which suggests that additional training could still be beneficial to narrow the gap to the asymptotic loss L0 .
Model
FID Inception↓
FID DINO↓
FVD I3D↓
FVD VideoMAE↓
ADE ↓
Tiny
26.87
280.19
175.99
198.40
–
Base
10.66
132.51
68.43
116.94
–
Large
8.67
117.76
48.09
99.43
–
1B
5.61
87.05
33.94
89.49
–
9B
4.91
60.81
37.16
75.86
–
1B*
4.41
63.93
32.86
49.28
3.92 -0.98
Table 2 : Scaling Model size on 2.5 s video generation quality on NATIX , conditioned on a single frame. Lower values indicate better performance; best results in each column are bold. Fréchet distances are computed with different backbones. Models marked with * use trajectory-conditioning post-training. Green subscripts indicate the ADE boost over the corresponding unconditioned model. Larger models consistently improve generation quality, consistent with lower pre-training validation loss.
Model
FID Inception↓
FID DINO↓
FVD I3D↓
FVD VideoMAE↓
DriveDreamer-2 [ 40 ]
25.0
–
105.1
–
Drive-WM [ 35 ]
15.8
–
122.7
–
GenAD [ 37 ]
15.4
–
184.0
–
Vista [ 7 ]
6.9
–
89.4
–
GEM [ 9 ]
10.5
–
158.5
–
Driving World [ 15 ]
7.4
–
90.9
–
Table 3 : Comparison on the nuScenes benchmark. Quantitative evaluation of 2.5 s video generation conditioned on a single input frame (lower is better; “–” indicates the metric is not reported). Results are evaluated on two test splits matching the protocols of Vista (5,369 sequences) and Epona (1,690 sequences). Our models achieve state-of-the-art performance, demonstrating the benefit of scaling both model size and training data, followed by domain-specific fine-tuning on nuScenes.
Figure 7 : Visual Quality on 5-second generations. Small models (<20M parameters) capture the overall scene layout but quickly lose object consistency across frames. The base model (135M) better understands the scene, e.g., it correctly predicts that the car is turning left, but struggles to maintain coherent dynamics over longer horizons. Larger models (1.1B parameters) further improve visual quality and scene structure, yet still exhibit temporal inconsistencies and break down after approximately 2.5 seconds. Only our largest 9B models accurately capture the scene dynamics while maintaining high visual fidelity and temporal coherence throughout the entire rollout.
Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primarily on static image-text pairs, video generative models capture temporal dynamics and motion priors that are naturally suited for driving. We present DriveWAM, a driving world-action model that adapts a pretrained video diffusion transformer into an autoregressive video-action policy. DriveWAM organizes video and action streams into a unified temporal token sequence and trains them under a joint flow-matching objective, preserving the pretrained video-generation architecture while adapting its large-scale video priors to action generation. To incorporate high-level scene understanding, we introduce scene-evolving driving guidance, where a frozen VLM produces chunk-specific semantic intent to guide video-action generation. To keep long-horizon rollout bounded, we further introduce selective KV memory, which maintains bounded modality-aware video and action memory pools through relevance-redundancy cache selection at inference time. Experiments on NAVSIM and the PhysicalAI-Autonomous-Vehicles benchmark show that DriveWAM achieves strong planning performance, and a data-scaling study from 4k to 100k driving clips further confirms the scaling potential of world-action modeling for end-to-end autonomous driving.
Chen Shi, Jinrui Xu, Shaoshuai Shi +3
The Chinese University of Hong Kong, Shenzhen · Voyager Research, Didi Chuxing
Video diffusion models can enable embodied agents to anticipate plausible futures from the recent past, but they are typically trained offline on curated datasets--a mismatch with the agents' learning setup at deployment: online, from a single video stream that sequentially outputs one frame at a time. We bridge this training gap and demonstrate that training autoregressive video diffusion models from such a stream, resembling the experience of embodied agents, is not only possible but can also perform comparably to standard offline training given the same number of gradient steps. We find that this robustness to video stream autocorrelation and nonstationarity can be achieved using experience replay methods that retain a subset of the video stream. To support training and evaluation in this setting, we introduce five new datasets for streaming lifelong generative video modeling: Lifelong Bouncing Balls (O), Lifelong Bouncing Balls (C), Lifelong 3D Maze, Lifelong Drive, and Lifelong PLAICraft, each consisting of one million consecutive frames from environments of increasing complexity. Together, our datasets and experiments lay the groundwork for video generative models and world models that continuously learn from single-sensor video streams rather than fixed datasets.
Jason Yoo, Yingchen He, Saeid Naderiparizi +4
University of British Columbia · KU Leuven · Vector Institute +1
Video diffusion is computationally expensive, as it requires executing a large model across many denoising steps. Even with step-distillation, inference remains expensive because every distilled step still requires a costly model evaluation. We present TRACK: TRajectory-Aware Capacity routing via top-K selection, a heterogeneous denoising strategy that switches between compatible large and small models at selected steps, reducing the average cost per denoising evaluation. The switching steps are determined using a calibration process. TRACK first rolls out a reference trajectory with the large model. Then at each step, the small model's prediction is also collected and compared against the large model's prediction to obtain a relative disagreement score. Both models receive the same latent, timestep, conditioning, and guidance inputs. Aggregating this signal over a calibration set produces a disagreement score map across diffusion steps, which determines a switching policy for an efficient inference process: quality-sensitive steps keep using the large model, while steps with low disagreement scores are routed to the small model. Inference executes only the selected model at each step, requiring no retraining, architecture or scheduler changes, or online dual-model evaluation. Across Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo, TRACK yields 1.95×, 2.04×-2.73×, 2.69×, and 2.17× speedups, respectively, with comparable aggregate quality and high diversity retention. TRACK thereby establishes automated, training-free model switching as a practical acceleration paradigm for video diffusion.