4DGS-Fixer: Generative Sparse-View 4D Gaussian Splatting with Iterative Refinement Guided by Video Diffusion Priors
Authors: Haitao Huang, Shenghao Zhao, Boyuan Tian, Shin-Fang Chng, Songlin Yang, Sheila Lim, Huangying Zhan, Yi Xu, +2 more
Organizations: Goertek Alpha Labs Shanghai, China · Singapore Institute of Technology Singapore, Singapore · Goertek Alpha Labs Santa Clara, USA · The Hong Kong University of Science and Technology Hong Kong, Hong Kong SAR, China
This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cannot fundamentally resolve the ill-posed problem caused by insufficient observations and missing scene information. Moreover, sparse-view 4D Gaussian Splatting (4DGS) often suffers from poor geometric initialization: with only a few input views, COLMAP typically reconstructs sparse and incomplete point clouds, leaving large scene regions without sufficient Gaussian support and making them difficult to recover through subsequent optimization. To address these limitations, we propose a novel iterative refinement framework based on a video diffusion model to improve the completeness and consistency of dynamic 4D scenes. Specifically, we first estimate multi-view depth maps and fuse them into dense point clouds to provide more complete geometric initialization for a dynamic 4DGS representation. We then employ a pretrained video restoration model to refine sequences rendered along novel camera trajectories at different time steps. The restored sequences serve as pseudo-supervision to regularize and iteratively refine the 4DGS representation. Experiments on a widely used benchmark dataset demonstrate that our method substantially outperforms existing baselines, achieving nearly a 2 dB PSNR improvement over the previous best-performing method.
Figures & tables
Figure 1. Overview of 4DGS-Fixer. Dense multi-view point clouds initialize the 4DGS representation, while restored novel-trajectory renderings provide cached pseudo-supervision for iterative refinement.
Figure 2. Qualitative visualization comparisons on the Neural3DV dataset.
Method
PSNR ↑
DSSIM 1 ↓
DSSIM 2 ↓
LPIPS ↓
STGS ( Li et al., 2024 )
17.70
0.158
0.107
0.325
4DGS ( Yang et al., 2024 )
20.60
0.143
0.094
0.244
4DGaussians ( Wu et al., 2024 )
20.82
0.117
0.077
0.190
4C4D ( Zhou et al., 2026 )
22.29
0.098
0.062
0.146
4DGS-Fixer (Ours)
24.18
0.086
0.053
0.128
Table 1. Quantitative comparison on the Neural3DV dataset.
Method
PSNR ↑
DSSIM 1 ↓
DSSIM 2 ↓
LPIPS ↓
Prep. Time ↓
STGS w/ COLMAP
17.702
0.158
0.107
0.325
∼ 20 min
STGS w/ MVS
21.674
0.109
0.071
0.165
∼ 3 min
CogVideoV2V restoration (init. 4DGS)
23.563
0.105
0.071
0.107
–
Full
24.180
0.086
0.053
0.128
–
CogVideoV2V restoration (refined 4DGS)
24.334
0.096
0.065
0.099
–
Table 2. Stage-wise ablation on Neural3DV. The restored pseudo-target is evaluated directly at held-out test views.
4D Gaussian Splatting (4DGS) can render dynamic scenes photorealistically. However, with limited viewpoint coverage, some spatiotemporal regions remain sparsely observed, leading to artifacts, particularly in scenes with large motion. Existing approaches leveraging generative models rely on heuristic virtual-viewpoint selection before refining rendered views. As a result, they cannot actively explore such sparsely observed regions. To address this issue, we propose a pipeline that actively selects spatiotemporal virtual viewpoints to improve 4DGS reconstruction. Our method selects virtual viewpoints for generative enhancement based on the rendering sensitivity and motion-aware observation density of 4D Gaussians, prioritizing views that alleviate observation sparsity. In the refined images, we filter out regions that conflict with captured observations or are likely to contain generative artifacts and then fine-tune 4DGS using only the reliable regions. We evaluate our method on multi-view video benchmarks using new train/test splits designed to induce observation gaps. Results show consistent improvements over prior viewpoint selection strategies and fine-tuning methods in both qualitative and quantitative evaluations, while reducing artifacts.
Dynamic 4D Gaussian Splatting has emerged as an efficient representation for dynamic novel view synthesis through explicit scene modeling and real-time rendering. However, existing methods typically require dense multi-view videos for sufficient geometric constraints, making capture expensive and limiting sparse-camera deployment. Reducing input views lowers acquisition cost but weakens geometry supervision, often causing missing structures and floating Gaussians. Depth priors provide geometric cues, yet no single source offers both dense coverage and reliable geometry. Monocular depth provides dense structure but is scale-ambiguous and locally biased, whereas multi-view geometric depth provides incomplete anchors consistent with the reconstruction coordinate system. To exploit their complementarity, we propose D2-4DGS, a sparse-camera dynamic 4D Gaussian Splatting framework guided by dual-source depth priors. We align monocular estimates with valid multi-view geometric depths and verify their consistency to identify reliable geometric anchors. These verified anchors support consistency-aware pruning and depth supervision, while verified geometric depths and aligned mono-only estimates provide candidate geometry for densification in under-reconstructed regions. Finally, RGB-D joint optimization improves appearance fidelity and geometric consistency under sparse-view supervision. Across all nine dataset--view settings, D2-4DGS achieves the highest PSNR, improving by 1.33 dB on average over the best competing method in each setting.
Learning a 4D scene representation from a single monocular video that supports dynamic novel-view synthesis while maintaining faithful geometry over time remains challenging. Dynamic Gaussian Splatting achieves strong rendering performance through photometric optimization, yet does not explicitly enforce multi-view geometric consistency. In contrast, 3D foundation models recover coherent scene geometry and camera motion, but their point-based outputs are not designed for photorealistic rendering. We propose Ground4D, a geometry-grounded framework built on two stages. First, we perform geometry initialization via 3D foundation models, leveraging VGGT in a training-free manner to reconstruct multi-view-consistent 3D geometry and camera poses from monocular video. The recovered geometry provides a structured and reliable initialization for dynamic Gaussian representations. Second, we conduct geometry-consistency-aware refinement via dynamic Gaussian Splatting, optimizing the representation through differentiable rendering while maintaining multi-view geometric consistency across both observed and synthesized viewpoints. Furthermore, Ground4D inherently models the continuous 4D dynamics of the scene, naturally supporting rendering at arbitrary timestamps. By integrating foundation-level geometric priors into dynamic Gaussian optimization, Ground4D achieves stronger reconstruction fidelity and rendering performance, underscoring the role of geometry-grounded constraints in robust 4D scene modeling.
Qing Zhao, Weijian Deng, Pengxu Wei +1
School of Computer Science and Engineering, Sun Yat-sen University, China · Tsinghua Shenzhen International Graduate School, Tsinghua University, China