Novel view synthesis (NVS) of static and dynamic urban scenes is essential for autonomous driving simulation, yet existing methods often struggle to balance reconstruction time with quality. While state-of-the-art neural radiance fields and 3D Gaussian Splatting approaches achieve photorealism, they often rely on time-consuming per-scene optimization. Conversely, emerging feed-forward methods frequently adopt per-pixel Gaussian representations, which lead to 3D inconsistencies when aggregating multi-view predictions in complex, dynamic environments. We propose EvolSplat4D, a feed-forward framework that moves beyond existing per-pixel paradigms by unifying volume-based and pixel-based Gaussian prediction across three specialized branches. For close-range static regions, we predict consistent geometry of 3D Gaussians over multiple frames directly from a 3D feature volume, complemented by a semantically-enhanced image-based rendering module for predicting their appearance. For dynamic actors, we utilize object-centric canonical spaces and a motion-adjusted rendering module to aggregate temporal features, ensuring stable 4D reconstruction despite noisy motion priors. Far-Field scenery is handled by an efficient per-pixel Gaussian branch to ensure full-scene coverage. Experimental results on the KITTI-360, KITTI, Waymo, and PandaSet datasets show that EvolSplat4D reconstructs both static and dynamic environments with superior accuracy and consistency, outperforming both per-scene optimization and state-of-the-art feed-forward baselines.
Figures & tables
Figure 1 : Overview of EVolSplat4D . We propose EVolSplat4D, a unified feed-forward 3D Gaussian Splatting framework tailored for static & dynamic urban scenes that achieves real-time rendering speeds. Leveraging both camera and tracked 3D bounding box as inputs, EVolSplat4D completes scene reconstruction in approximately 1.3 seconds, achieving photo-realistic quality comparable to time-consuming per-scene optimization methods. EVolSplat4D also supports various downstream applications, including high-fidelity scene editing and scene decomposition.
Figure 2 : Method Overview. We reconstruct urban scenes by disentangling them as close-range volume, dynamic actors, and far-field scenery, predicting 3D Gaussians of each in a feed-forward manner. a) Given a set of images, we initialize our model with the pretrained depth model and DINO feature extractor. b) In close-range volume, we leverage the 3D context of F3D to predict the geometry attributes of 3D Gaussians and project the 3D Gaussians to the reference views to retrieve 2D context, including color window and visibility maps to decode their color. c) For dynamic actors, we model each instance using an instance-wise canonical space and perform feed-forward reconstitution through our proposed motion-adjusted IBR module. d) To model far-range regions, we employ a 2D U-Net backbone F2D with cross-view self-attention to aggregate information from nearby reference images and predict per-pixel Gaussians. e) The composition of the three parts leads to our full model for unbounded scenes.
Figure 3 : Occlusion Illustration . a) One Gaussian in 3D space may retrieve inaccurate color information from 2D reference images due to occlusions. b) The previous method ( Miao et al., 2025 ) uses depth priors to check occlusions, which may suffer from inaccurate monocular depth predictions. c) In contrast, EVolSplat4D comprises robust DINO priors to reduce the impact of invisible colors to enhance rendering quality.
Figure 4 : Motion Adjusted IBR. LiDAR points in the canonical space for an instance m are transformed using time-specific poses and projected into reference frames. A Window-based Projection strategy then samples coherent appearance features cik and visibility map vik , which are input to the dynamic Gaussian decoder Ddyn . The decoder regresses the 3D Gaussian attributes for each point, enabling the final rendering of the dynamic object’s image and binary mask.
Cat.
Method
Sensors
3D Box Tracks
Rep.
Feed-Forward
Dynamic
Real-Time
Scene Edit
Per-scene Opt.
SUDS ( Turki et al., 2023 )
LiDAR+RGB
N
Non-PA
✘
✔
✘
✘
EmerNeRF ( Yang et al., 2023a )
LiDAR+RGB
N
Non-PA
✘
✔
✘
✘
StreetGaussian ( Yan et al., 2024 )
LiDAR+RGB
Pred
Non-PA
✘
✔
✔
✔
OmniRe ( Chen et al., 2024c )
LiDAR+RGB
GT
Non-PA
✘
✔
✔
✔
DeSiRe-GS ( Peng et al., 2025 )
LiDAR+RGB
N
Non-PA
✘
✔
✔
✘
FeedForward Recon.
AnySplat ( Jiang et al., 2025a )
RGB
N
PA
✔
✘
✔
✘
Table 1 : Categorization of representative baselines. We categorize representative methods into per-scene optimization and feed-forward reconstruction methods. PA and Non-PA denote pixel-aligned and non-pixel-aligned representations, respectively.
Method
KITTI- 360 (In-Domain)
Waymo (Out-of-Domain)
FPS ↑
Mem.(GB) ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
(Render res.: 376 × 1408)
MVSNeRF ( Chen et al., 2021 )
18.44
0.638
0.317
17.86
0.595
0.433
0.025
12.03
MuRF ( Xu et al., 2023 )
22.77
0.780
0.229
23.33
0.770
0.269
0.31
26.45
EDUS ( Miao et al., 2024 )
22.13
0.745
0.178
23.18
0.745
0.164
0.14
5.75
PixelSplat ( Charatan et al., 2024 )
19.41
0.584
0.357
16.65
0.541
0.579
81.58
26.75
MVSplat ( Chen et al., 2024b )
21.22
0.695
0.268
21.33
0.665
0.308
70.46
16.14
Table 2 : Quantitative results on Static Scenes with other feed-forward baselines on both KITTI-360 and Waymo open datasets. All models are trained on the KITTI-360 dataset using a drop50% sparsity level. Metrics are averaged on five validation scenes without any finetuning. It is also worth noting that our method is more memory efficient compared with other 3DGS-based methods.
Figure 5 : Qualitative Comparison with feed-forward baselines on the KITTI-360 dataset.
Method
KITTI (In-Domain)
Waymo (In-Domain)
PandaSet (Out-of-Domain)
FPS ↑
Mem.(GB) ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
(Render res.: 320 × 480 )
DrivingRecon ( Lu et al., 2024 )
19.50
0.597
0.283
20.21
0.572
0.277
14.92
0.506
0.389
190.78
9.48
STORM ( Yang et al., 2024a )
18.96
0.658
0.256
23.76
0.764
0.182
23.13
0.766
0.207
178.57
22.31
EVolSplat
16.52
0.583
0.275
22.86
0.753
0.201
23.32
0.746
0.253
212.42
6.41
EVolSplat4D
20.76
0.738
0.162
26.32
0.822
0.121
26.27
0.823
0.132
201.06
6.71
Table 3 : Quantitative results on Dynamic Scenes with feed-forward baselines. Metrics are averaged on five validation scenes without any finetuning in drop80% sparsity level. Our EVolSplat4D generalizes better than all the baseline in a large margin on both in-domain (KITTI, Waymo) and out-of-domain (PandaSet) datasets.
Figure 6 : Qualitative Comparisons with baselines for feed-forward inference on the KITTI dataset. The rendering images are downscaled 2x to 188 × 621.
Figure 7 : Qualitative Comparisons with baselines for feed-forward inference on the Waymo Open dataset and PandaSet. Each column represents a different method, while each pair of rows corresponds to a different dataset. The rendering images resolution are 480 × 320 for Waymo and 480 × 270 for PandaSet.
Method
KITTI
Waymo
PandaSet
DriveRecon ( Lu et al., 2024 )
0.350
0.315
0.383
STORM ( Yang et al., 2024a )
0.104
0.080
0.102
EVolSplat4D (Ours)
0.062
0.063
0.080
Table 4 : Quantitative Comparison of View Extrapolation Quality . We report averaged KID ↓ (lower values indicate better performance) on diverse datasets without finetuning in Drop 80% sparsity level.
Figure 8 : Extrapolated Views Qualitative Comparison with STORM on PandaSet.
Figure 9 : Convergence Analysis on the Waymo Dataset. We plot PSNR and LPIPS on one test set at different training steps. Compared with optimization-based baselines, our method with generalizable priors converges faster ( ≈ 1000 steps) and achieves better PSNR.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Rec. Time ↓
SUDS ( Turki et al., 2023 )
25.03
0.813
0.167
45min
EmerNeRF ( Yang et al., 2023a )
27.03
0.886
0.069
21min
StreetGaussian ( Yan et al., 2024 )
27.27
0.891
0.057
7min
DeSiRe-GS ( Peng et al., 2025 )
27.35
0.889
0.094
14min
OmniRe ( Chen et al., 2024c )
28.16
0.912
0.051
9min
EVolSplat4D (Ours 0-iter )
25.39
0.841
0.119
1.3s
Table 5 : Comparison with Optimization-Based baselines on 3 sequences of the Waymo NOTR dataset under the Drop 80% sparsity level. The Rec. Time refers to the reconstruction time needed to reach convergence and achieve the reported metrics.
Method
Waymo
PandaSet
KID ↓ @ 1 m
KID ↓ @ 3 m
KID ↓ @ 1 m
KID ↓ @ 3 m
StreetGaussian ( Yan et al., 2024 )
0.055
0.203
0.062
0.198
OmniRe ( Chen et al., 2024c )
0.041
0.158
0.033
0.116
EVolSplat4D (Ours-FF)
0.072
0.230
0.095
0.244
EVolSplat4D (Ours-FT)
0.038
0.114
0.029
0.118
Table 6: Quantitative comparison of view extrapolation quality with optimization-based baselines. We report KID at 1 m and 3 m lane shifts on Waymo and PandaSet.
Figure 10 : Qualitative Results of Ablation Study. We represent rendering images via feed-forward inference on one novel scene.
PSNR ↑
SSIM ↑
LPIPS ↓
w/o volume branch
27.02
0.839
0.132
w/o pixel branch
26.29
0.817
0.194
w/o motion-adjusted IBR
26.77
0.837
0.121
w/o Occ. check
27.33
0.844
0.113
w/o Lmask
27.45
0.847
0.108
Full
27.78
0.856
0.102
Table 7: Ablation Analysis of method components.
Setting
PSNR ↑
SSIM ↑
LPIPS ↓
W=1 , Pred bbx
27.45
0.850
0.115
W=3 , Pred bbx
27.78
0.856
0.102
W=5 , Pred bbx
27.56
0.853
0.106
W=3 , Gt bbx
27.73
0.860
0.095
Table 8: Ablation study on window size ( W ) and bounding box source. Using W=3 with predicted bounding boxes (Pred bbx) yields the best results among window sizes. Notably, this performance is comparable to using ground-truth bounding boxes (Gt bbx).
Figure 11 : Qualitative Results of Scene Decompose Loss. The images are rendered only from the Gaussians within the close-range region. Without Lmask , the far-field pixel branch partially models the close-range building, causing Gaussians from the Gcr to be semi-transparent.
Noisy Setting
PSNR ↑
SSIM ↑
LPIPS ↓
Pred Box+Center noise σ=0.25m
27.28
0.836
0.120
Pred Box+Center noise σ=0.5m
26.76
0.821
0.125
Pred Box+Yaw noise σ=10∘
27.52
0.844
0.114
Pred Box+Yaw noise σ=15∘
27.31
0.841
0.117
Pred Box
27.78
0.856
0.102
Table 9 : Ablation study of bounding box perturbation on the Waymo dataset.
PSNR ↑
SSIM ↑
LPIPS ↓
i=0
27.260
0.833
0.142
i=1
27.787
0.856
0.102
i=2
27.781
0.856
0.102
i=3
27.785
0.856
0.102
Table 10: Ablation Study on recursion of Δμi . The results are tested on the Waymo dataset.
Modality
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Monocular Pred.
UniDepth ( Piccinelli et al., 2024 )
27.59
0.854
0.105
Metric3D ( Yin et al., 2023 )
27.86
0.859
0.104
LiDAR Compl.
Prior Depth Anything ( Wang et al., 2025e )
27.78
0.856
0.102
Table 11 : Depth Sensitivity Experiment. The results are averaged on five testsets from the Waymo dataset.
Figure 12 : Scene Decomposition on Waymo. Our approach enables clear decomposition of the close-range volume, dynamic actors, and the far-field scenery.
Figure 13 : Scene Editing on KITTI. Images in the first and second columns represent the results before and after editing.
SparseCNN Network Architecture
Layer
Description
In/Out Ch.
Conv 0
kernel = 3×3×3 , stride = 1
16/16
Conv 1
kernel = 3×3×3 , stride = 2
16/16
Conv 2
kernel = 3×3×3 , stride = 2
16/32
Conv 3
kernel = 3×3×3 , stride = 1
32/32
Conv 4
kernel = 3×3×3 , stride = 2
32/64
Table S1 : Architecture of SparseConvNet . Each layer consists of sparse convolution, batch normalization, and ReLU.
a) MuRF
Figure S2: Qualitative scene reconstruction comparisons with per-scene optimization baselines under the Drop80% sparsity level
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Rec. Time ↓
EmerNeRF ( Yang et al., 2023a )
26.70
0.863
0.125
43 min
StreetGaussian ( Yan et al., 2024 )
26.91
0.870
0.092
15 min
OmniRe ( Chen et al., 2024c )
27.35
0.884
0.087
21 min
EVolSplat4D (Ours-FT)
27.28
0.901
0.071
6 min 04 s
Table S2 : Comparison with optimization-based baselines on three 120-frame Waymo NOTR validation sequences under the denser Drop 25% sparsity setting and higher resolution ( 640×960 ). Rec. Time denotes the reconstruction time required to achieve the reported metrics.
Figure S3 : Efficiency scaling analysis. We report memory usage and inference time with respect to the number of input frames and image resolution.
Figure S4 : Downstream utility visualization, including novel RGB&Depth rendering, optical flow, dynamic mask, and re-rendering with modified camera parameters.
Figure S5 : Qualitative comparison of view extrapolation quality. We compare 1 m and 3 m lane-shifted rendering results with optimization-based baselines.
Figure S6 : Qualitative editing comparison with OmniRe on actor replacement and 2 m lane shift.
Figure S7 : Failure case for actor lane shift under a large translation.
Figure S8 : Qualitative results of bounding box perturbation on Waymo.
Figure S9 : More Qualitative Results on Waymo . We visualize synthesized images (odd rows) and the corresponding predicted dynamic actors (even rows) on novel scenes generated by our pretrained model through a feed-forward inference.
Figure S10 : More Qualitative Results on out-of-domain dataset PandaSet . We visualize synthesized images (odd rows) and the corresponding predicted dynamic actors (even rows) on novel scenes generated by our pretrained model through a feed-forward inference.
Figure S11 : More Qualitative Results on KITTI . We visualize synthesized images (odd rows) and the corresponding predicted dynamic actors (even rows) on novel scenes generated by our pretrained model through a feed-forward inference.
Feedforward Gaussian Splatting has recently emerged as an efficient paradigm for 4D reconstruction in autonomous driving. However, in unstructured off-road scenes, its performance degrades due to high-frequency geometry, ego-motion jitter, and increased non-rigid dynamics. These factors introduce conflicting Gaussian observations across timestamps, leading to either over-smoothed renderings or structural artifacts. To address this issue, we propose Ground4D, a spatially-grounded 4D feedforward framework for pose-free off-road reconstruction. The key idea is to resolve temporal conflicts through spatially localized conditioning. Specifically, we introduce voxel-grounded temporal Gaussian aggregation, which partitions the canonical Gaussian space into spatial voxels and performs query-conditioned temporal attention within each voxel. Intra-voxel softmax normalization ensures that temporal selectivity and spatial occupancy become mutually reinforcing rather than conflicting. We furthermore introduce surface normal cues as auxiliary geometric guidance to regularize the geometry of Gaussian primitives. Extensive experiments on ORAD-3D and RELLIS-3D demonstrate that Ground4D consistently outperforms existing feedforward methods in reconstruction quality and generalizes zero-shot to unseen off-road domains. Project page and code:https://github.com/wsnbws/Ground4D.
Shuo Wang, Jilin Mei, Fuyang Liu +6
Institute of Computing Technology, Chinese Academy of Sciences China · Xi’an Jiaotong University China · Beijing Institute of Technology China
Recent extensions of 3D Gaussian Splatting (3DGS) enable real-time novel view synthesis in dynamic scenes by learning time-conditioned Gaussian deformations. However, existing MLP-based methods typically estimate deformations independently at each timestamp, making them less robust to large or abrupt motions. To address this issue, we propose \textbf{EvoGS}, a 3DGS-based dynamic reconstruction framework that models Gaussian deformation as a temporal evolution process. EvoGS maintains persistent deformation states for each Gaussian, extrapolates future states from historical deformation states, and corrects the predictions with MLP-derived observations. The correction is adaptively weighted using a temporal residual memory and evolution statistics such as deformation velocity and trajectory deviation. To further improve reconstruction quality, EvoGS introduces deformation-aware densification. Clone and split operations are performed along corrected deformation directions, while an uncertainty-aware strategy suppresses densification for Gaussians with unstable deformation histories. Experiments show that EvoGS improves dynamic novel view synthesis quality and achieves competitive performance across benchmarks.
Wei Dong, Shahram Shirani, Jun Chen +1
The Department of Electrical & Computer Engineering, McMaster University
High-fidelity reconstruction of dynamic urban environments is a cornerstone of autonomous driving simulation and large-scale world modeling. While 3D Gaussian Splatting (3DGS) has established a new standard for real-time rendering, its reliance on expensive per-scene optimization limits scalability. Conversely, recent feedforward methods that infer Gaussian parameters offer faster speed but face fundamental bottlenecks: they are memory-prohibitive at high resolutions and struggle to fuse dense multi-view observations consistently. This paper presents L2D2-GS, a unified framework that reformulates generalizable reconstruction not as a one-shot regression, but as a robust iterative process of optimization and densification. To resolve the ambiguity of supervision in primitive generation, we propose a self-supervised densification policy that derives explicit reward signals from global reconstruction gains to guide local densification. Furthermore, we mitigate irreversible early-stage artifacts through a geometric regularization mechanism, utilizing reparameterization to constrain the optimization manifold and prevent convergence to poor local optima. Extensive experiments on the PandaSet and Waymo datasets demonstrate that our method achieves state-of-the-art reconstruction fidelity and strong zero-shot generalization, while using fewer primitives than competing baselines.
Zetian Song, Chenming Wu, Junnan Liu +6
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, 100871, China · Xiaomi EV, Beijing, 100085, China