Novel view synthesis (NVS) of static and dynamic urban scenes is essential for autonomous driving simulation, yet existing methods often struggle to balance reconstruction time with quality. While state-of-the-art neural radiance fields and 3D Gaussian Splatting approaches achieve photorealism, they often rely on time-consuming per-scene optimization. Conversely, emerging feed-forward methods frequently adopt per-pixel Gaussian representations, which lead to 3D inconsistencies when aggregating multi-view predictions in complex, dynamic environments. We propose EvolSplat4D, a feed-forward framework that moves beyond existing per-pixel paradigms by unifying volume-based and pixel-based Gaussian prediction across three specialized branches. For close-range static regions, we predict consistent geometry of 3D Gaussians over multiple frames directly from a 3D feature volume, complemented by a semantically-enhanced image-based rendering module for predicting their appearance. For dynamic actors, we utilize object-centric canonical spaces and a motion-adjusted rendering module to aggregate temporal features, ensuring stable 4D reconstruction despite noisy motion priors. Far-Field scenery is handled by an efficient per-pixel Gaussian branch to ensure full-scene coverage. Experimental results on the KITTI-360, KITTI, Waymo, and PandaSet datasets show that EvolSplat4D reconstructs both static and dynamic environments with superior accuracy and consistency, outperforming both per-scene optimization and state-of-the-art feed-forward baselines.
Figures & tables
Figure 1 : Overview of EVolSplat4D . We propose EVolSplat4D, a unified feed-forward 3D Gaussian Splatting framework tailored for static & dynamic urban scenes that achieves real-time rendering speeds. Leveraging both camera and tracked 3D bounding box as inputs, EVolSplat4D completes scene reconstruction in approximately 1.3 seconds, achieving photo-realistic quality comparable to time-consuming per-scene optimization methods. EVolSplat4D also supports various downstream applications, including high-fidelity scene editing and scene decomposition.
Figure 2 : Method Overview. We reconstruct urban scenes by disentangling them as close-range volume, dynamic actors, and far-field scenery, predicting 3D Gaussians of each in a feed-forward manner. a) Given a set of images, we initialize our model with the pretrained depth model and DINO feature extractor. b) In close-range volume, we leverage the 3D context of F3D to predict the geometry attributes of 3D Gaussians and project the 3D Gaussians to the reference views to retrieve 2D context, including color window and visibility maps to decode their color. c) For dynamic actors, we model each instance using an instance-wise canonical space and perform feed-forward reconstitution through our proposed motion-adjusted IBR module. d) To model far-range regions, we employ a 2D U-Net backbone F2D with cross-view self-attention to aggregate information from nearby reference images and predict per-pixel Gaussians. e) The composition of the three parts leads to our full model for unbounded scenes.
Figure 3 : Occlusion Illustration . a) One Gaussian in 3D space may retrieve inaccurate color information from 2D reference images due to occlusions. b) The previous method ( Miao et al., 2025 ) uses depth priors to check occlusions, which may suffer from inaccurate monocular depth predictions. c) In contrast, EVolSplat4D comprises robust DINO priors to reduce the impact of invisible colors to enhance rendering quality.
Figure 4 : Motion Adjusted IBR. LiDAR points in the canonical space for an instance m are transformed using time-specific poses and projected into reference frames. A Window-based Projection strategy then samples coherent appearance features cik and visibility map vik , which are input to the dynamic Gaussian decoder Ddyn . The decoder regresses the 3D Gaussian attributes for each point, enabling the final rendering of the dynamic object’s image and binary mask.
Cat.
Method
Sensors
3D Box Tracks
Rep.
Feed-Forward
Dynamic
Real-Time
Scene Edit
Per-scene Opt.
SUDS ( Turki et al., 2023 )
LiDAR+RGB
N
Non-PA
✘
✔
✘
✘
EmerNeRF ( Yang et al., 2023a )
LiDAR+RGB
N
Non-PA
✘
✔
✘
✘
StreetGaussian ( Yan et al., 2024 )
LiDAR+RGB
Pred
Non-PA
✘
✔
✔
✔
OmniRe ( Chen et al., 2024c )
LiDAR+RGB
GT
Non-PA
✘
✔
✔
✔
DeSiRe-GS ( Peng et al., 2025 )
LiDAR+RGB
N
Non-PA
✘
✔
✔
✘
FeedForward Recon.
AnySplat ( Jiang et al., 2025a )
RGB
N
PA
✔
✘
✔
✘
Table 1 : Categorization of representative baselines. We categorize representative methods into per-scene optimization and feed-forward reconstruction methods. PA and Non-PA denote pixel-aligned and non-pixel-aligned representations, respectively.
Method
KITTI- 360 (In-Domain)
Waymo (Out-of-Domain)
FPS ↑
Mem.(GB) ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
(Render res.: 376 × 1408)
MVSNeRF ( Chen et al., 2021 )
18.44
0.638
0.317
17.86
0.595
0.433
0.025
12.03
MuRF ( Xu et al., 2023 )
22.77
0.780
0.229
23.33
0.770
0.269
0.31
26.45
EDUS ( Miao et al., 2024 )
22.13
0.745
0.178
23.18
0.745
0.164
0.14
5.75
PixelSplat ( Charatan et al., 2024 )
19.41
0.584
0.357
16.65
0.541
0.579
81.58
26.75
MVSplat ( Chen et al., 2024b )
21.22
0.695
0.268
21.33
0.665
0.308
70.46
16.14
Table 2 : Quantitative results on Static Scenes with other feed-forward baselines on both KITTI-360 and Waymo open datasets. All models are trained on the KITTI-360 dataset using a drop50% sparsity level. Metrics are averaged on five validation scenes without any finetuning. It is also worth noting that our method is more memory efficient compared with other 3DGS-based methods.
Figure 5 : Qualitative Comparison with feed-forward baselines on the KITTI-360 dataset.
Method
KITTI (In-Domain)
Waymo (In-Domain)
PandaSet (Out-of-Domain)
FPS ↑
Mem.(GB) ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
(Render res.: 320 × 480 )
DrivingRecon ( Lu et al., 2024 )
19.50
0.597
0.283
20.21
0.572
0.277
14.92
0.506
0.389
190.78
9.48
STORM ( Yang et al., 2024a )
18.96
0.658
0.256
23.76
0.764
0.182
23.13
0.766
0.207
178.57
22.31
EVolSplat
16.52
0.583
0.275
22.86
0.753
0.201
23.32
0.746
0.253
212.42
6.41
EVolSplat4D
20.76
0.738
0.162
26.32
0.822
0.121
26.27
0.823
0.132
201.06
6.71
Table 3 : Quantitative results on Dynamic Scenes with feed-forward baselines. Metrics are averaged on five validation scenes without any finetuning in drop80% sparsity level. Our EVolSplat4D generalizes better than all the baseline in a large margin on both in-domain (KITTI, Waymo) and out-of-domain (PandaSet) datasets.
Figure 6 : Qualitative Comparisons with baselines for feed-forward inference on the KITTI dataset. The rendering images are downscaled 2x to 188 × 621.
Figure 7 : Qualitative Comparisons with baselines for feed-forward inference on the Waymo Open dataset and PandaSet. Each column represents a different method, while each pair of rows corresponds to a different dataset. The rendering images resolution are 480 × 320 for Waymo and 480 × 270 for PandaSet.
Method
KITTI
Waymo
PandaSet
DriveRecon ( Lu et al., 2024 )
0.350
0.315
0.383
STORM ( Yang et al., 2024a )
0.104
0.080
0.102
EVolSplat4D (Ours)
0.062
0.063
0.080
Table 4 : Quantitative Comparison of View Extrapolation Quality . We report averaged KID ↓ (lower values indicate better performance) on diverse datasets without finetuning in Drop 80% sparsity level.
Figure 8 : Extrapolated Views Qualitative Comparison with STORM on PandaSet.
Figure 9 : Convergence Analysis on the Waymo Dataset. We plot PSNR and LPIPS on one test set at different training steps. Compared with optimization-based baselines, our method with generalizable priors converges faster ( ≈ 1000 steps) and achieves better PSNR.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Rec. Time ↓
SUDS ( Turki et al., 2023 )
25.03
0.813
0.167
45min
EmerNeRF ( Yang et al., 2023a )
27.03
0.886
0.069
21min
StreetGaussian ( Yan et al., 2024 )
27.27
0.891
0.057
7min
DeSiRe-GS ( Peng et al., 2025 )
27.35
0.889
0.094
14min
OmniRe ( Chen et al., 2024c )
28.16
0.912
0.051
9min
EVolSplat4D (Ours 0-iter )
25.39
0.841
0.119
1.3s
Table 5 : Comparison with Optimization-Based baselines on 3 sequences of the Waymo NOTR dataset under the Drop 80% sparsity level. The Rec. Time refers to the reconstruction time needed to reach convergence and achieve the reported metrics.
Method
Waymo
PandaSet
KID ↓ @ 1 m
KID ↓ @ 3 m
KID ↓ @ 1 m
KID ↓ @ 3 m
StreetGaussian ( Yan et al., 2024 )
0.055
0.203
0.062
0.198
OmniRe ( Chen et al., 2024c )
0.041
0.158
0.033
0.116
EVolSplat4D (Ours-FF)
0.072
0.230
0.095
0.244
EVolSplat4D (Ours-FT)
0.038
0.114
0.029
0.118
Table 6: Quantitative comparison of view extrapolation quality with optimization-based baselines. We report KID at 1 m and 3 m lane shifts on Waymo and PandaSet.
Figure 10 : Qualitative Results of Ablation Study. We represent rendering images via feed-forward inference on one novel scene.
PSNR ↑
SSIM ↑
LPIPS ↓
w/o volume branch
27.02
0.839
0.132
w/o pixel branch
26.29
0.817
0.194
w/o motion-adjusted IBR
26.77
0.837
0.121
w/o Occ. check
27.33
0.844
0.113
w/o Lmask
27.45
0.847
0.108
Full
27.78
0.856
0.102
Table 7: Ablation Analysis of method components.
Setting
PSNR ↑
SSIM ↑
LPIPS ↓
W=1 , Pred bbx
27.45
0.850
0.115
W=3 , Pred bbx
27.78
0.856
0.102
W=5 , Pred bbx
27.56
0.853
0.106
W=3 , Gt bbx
27.73
0.860
0.095
Table 8: Ablation study on window size ( W ) and bounding box source. Using W=3 with predicted bounding boxes (Pred bbx) yields the best results among window sizes. Notably, this performance is comparable to using ground-truth bounding boxes (Gt bbx).
Figure 11 : Qualitative Results of Scene Decompose Loss. The images are rendered only from the Gaussians within the close-range region. Without Lmask , the far-field pixel branch partially models the close-range building, causing Gaussians from the Gcr to be semi-transparent.
Noisy Setting
PSNR ↑
SSIM ↑
LPIPS ↓
Pred Box+Center noise σ=0.25m
27.28
0.836
0.120
Pred Box+Center noise σ=0.5m
26.76
0.821
0.125
Pred Box+Yaw noise σ=10∘
27.52
0.844
0.114
Pred Box+Yaw noise σ=15∘
27.31
0.841
0.117
Pred Box
27.78
0.856
0.102
Table 9 : Ablation study of bounding box perturbation on the Waymo dataset.
PSNR ↑
SSIM ↑
LPIPS ↓
i=0
27.260
0.833
0.142
i=1
27.787
0.856
0.102
i=2
27.781
0.856
0.102
i=3
27.785
0.856
0.102
Table 10: Ablation Study on recursion of Δμi . The results are tested on the Waymo dataset.
Modality
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Monocular Pred.
UniDepth ( Piccinelli et al., 2024 )
27.59
0.854
0.105
Metric3D ( Yin et al., 2023 )
27.86
0.859
0.104
LiDAR Compl.
Prior Depth Anything ( Wang et al., 2025e )
27.78
0.856
0.102
Table 11 : Depth Sensitivity Experiment. The results are averaged on five testsets from the Waymo dataset.
Figure 12 : Scene Decomposition on Waymo. Our approach enables clear decomposition of the close-range volume, dynamic actors, and the far-field scenery.
Figure 13 : Scene Editing on KITTI. Images in the first and second columns represent the results before and after editing.
SparseCNN Network Architecture
Layer
Description
In/Out Ch.
Conv 0
kernel = 3×3×3 , stride = 1
16/16
Conv 1
kernel = 3×3×3 , stride = 2
16/16
Conv 2
kernel = 3×3×3 , stride = 2
16/32
Conv 3
kernel = 3×3×3 , stride = 1
32/32
Conv 4
kernel = 3×3×3 , stride = 2
32/64
Table S1 : Architecture of SparseConvNet . Each layer consists of sparse convolution, batch normalization, and ReLU.
a) MuRF
Figure S2: Qualitative scene reconstruction comparisons with per-scene optimization baselines under the Drop80% sparsity level
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Rec. Time ↓
EmerNeRF ( Yang et al., 2023a )
26.70
0.863
0.125
43 min
StreetGaussian ( Yan et al., 2024 )
26.91
0.870
0.092
15 min
OmniRe ( Chen et al., 2024c )
27.35
0.884
0.087
21 min
EVolSplat4D (Ours-FT)
27.28
0.901
0.071
6 min 04 s
Table S2 : Comparison with optimization-based baselines on three 120-frame Waymo NOTR validation sequences under the denser Drop 25% sparsity setting and higher resolution ( 640×960 ). Rec. Time denotes the reconstruction time required to achieve the reported metrics.
Figure S3 : Efficiency scaling analysis. We report memory usage and inference time with respect to the number of input frames and image resolution.
Figure S4 : Downstream utility visualization, including novel RGB&Depth rendering, optical flow, dynamic mask, and re-rendering with modified camera parameters.
Figure S5 : Qualitative comparison of view extrapolation quality. We compare 1 m and 3 m lane-shifted rendering results with optimization-based baselines.
Figure S6 : Qualitative editing comparison with OmniRe on actor replacement and 2 m lane shift.
Figure S7 : Failure case for actor lane shift under a large translation.
Figure S8 : Qualitative results of bounding box perturbation on Waymo.
Figure S9 : More Qualitative Results on Waymo . We visualize synthesized images (odd rows) and the corresponding predicted dynamic actors (even rows) on novel scenes generated by our pretrained model through a feed-forward inference.
Figure S10 : More Qualitative Results on out-of-domain dataset PandaSet . We visualize synthesized images (odd rows) and the corresponding predicted dynamic actors (even rows) on novel scenes generated by our pretrained model through a feed-forward inference.
Figure S11 : More Qualitative Results on KITTI . We visualize synthesized images (odd rows) and the corresponding predicted dynamic actors (even rows) on novel scenes generated by our pretrained model through a feed-forward inference.
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, 100871, China · Xiaomi EV, Beijing, 100085, China