We present a 4D Gaussian Splatting framework for the Sparse-View Track of the SIGGRAPH Asia 2026 Volumetric Video Challenge, which requires dynamic scene reconstruction from only six cameras with wide baselines. To achieve robust dynamic reconstruction under such sparse views, our framework integrates three components. (1) Region-adaptive spatial priors: We use foreground masks to guide Gaussian initialization and mask voting to control densification separately for the dynamic foreground and static background. Background geometry is regularized using monocular depth aligned to metric scale. (2) Motion-consistent temporal priors: We provide supervision at intermediate times through frame interpolation and constrain projected Gaussian motion with estimated optical flow. (3) Generative assistance: We place virtual cameras in the widest angular gaps and restore their rendered images using a diffusion-based model conditioned on camera poses. The restored images are iteratively incorporated into training as pseudo-supervision. On the validation set, our framework improves full-frame PSNR from 25.60 dB for the baseline to 29.75 dB. On the official test benchmark, it achieves 30.04 dB full-frame PSNR and 27.88 dB foreground PSNR, ranking first overall in the Sparse-View Track.
Figures & tables
Figure 1. Overview of our sparse-view 4D Gaussian Splatting pipeline. Given videos from six cameras with wide baselines, our framework combines three components: (1) region-adaptive spatial priors use YOLO11 and SAM 2 masks for per-frame foreground initialization and mask voting during densification, and Depth Anything V2 predictions for background depth regularization; (2) motion-consistent temporal priors use RIFE interpolation to provide supervision at intermediate times and VideoFlow estimates to constrain projected Gaussian motion; and (3) generative assistance restores images rendered from virtual cameras in wide angular gaps using the ArtiFixer diffusion model ( de Lutio et al., 2026 ) , with the nearest training views in camera pose space as references. The restored images are iteratively incorporated into training as pseudo-supervision, yielding the final 4DGS model for novel view synthesis. Pipeline from six wide-baseline cameras to 4D Gaussian Splatting optimization. Spatial priors use foreground masks and monocular depth, temporal priors use interpolation and optical flow, and generative assistance restores virtual views before pseudo-supervised re-injection.
Full frame
Foreground
Setting
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Baseline ( θ=5e−6 )
25.60
0.9091
0.2429
27.28
0.8686
0.2276
Baseline ( θ=1e−5 )
28.08
0.9214
0.2148
26.63
0.8554
0.2539
w/o voting ( θ=5e−6 )
28.27
0.9263
0.1977
27.68
0.8741
0.2220
w/o voting ( θ=1e−5 )
28.49
0.9259
0.2017
26.92
0.8614
0.2409
w/o mask init
28.64
0.9267
0.1955
27.68
0.8743
0.2219
Table 1. Ablation study on the validation set. “Baseline” refers to the reconstruction method in ( Yang et al., 2026 ) . “Ours (recon)” includes all proposed spatial and temporal priors, while “Ours (+ gen)” additionally uses generative assistance. Parenthesized θ indicates a uniform densification threshold; otherwise, θfg and θbg are used for the foreground and background, respectively. Best and second-best results are highlighted for each metric.
Figure 2. Qualitative ablation of mask-voting densification. Qualitative comparison of dynamic foreground reconstruction without mask voting versus our full adaptive model.
Figure 3. Qualitative ablation of background depth regularization. Qualitative comparison of background reconstruction without depth regularization versus our full model.
Figure 4. Qualitative ablation of motion-consistent temporal priors. Qualitative comparison of dynamic reconstruction without temporal priors versus our full model.
Figure 5. Qualitative ablation of generative assistance. Qualitative comparison of novel-view reconstruction with and without generative assistance.
Dynamic 4D Gaussian Splatting has emerged as an efficient representation for dynamic novel view synthesis through explicit scene modeling and real-time rendering. However, existing methods typically require dense multi-view videos for sufficient geometric constraints, making capture expensive and limiting sparse-camera deployment. Reducing input views lowers acquisition cost but weakens geometry supervision, often causing missing structures and floating Gaussians. Depth priors provide geometric cues, yet no single source offers both dense coverage and reliable geometry. Monocular depth provides dense structure but is scale-ambiguous and locally biased, whereas multi-view geometric depth provides incomplete anchors consistent with the reconstruction coordinate system. To exploit their complementarity, we propose D2-4DGS, a sparse-camera dynamic 4D Gaussian Splatting framework guided by dual-source depth priors. We align monocular estimates with valid multi-view geometric depths and verify their consistency to identify reliable geometric anchors. These verified anchors support consistency-aware pruning and depth supervision, while verified geometric depths and aligned mono-only estimates provide candidate geometry for densification in under-reconstructed regions. Finally, RGB-D joint optimization improves appearance fidelity and geometric consistency under sparse-view supervision. Across all nine dataset--view settings, D2-4DGS achieves the highest PSNR, improving by 1.33 dB on average over the best competing method in each setting.
This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cannot fundamentally resolve the ill-posed problem caused by insufficient observations and missing scene information. Moreover, sparse-view 4D Gaussian Splatting (4DGS) often suffers from poor geometric initialization: with only a few input views, COLMAP typically reconstructs sparse and incomplete point clouds, leaving large scene regions without sufficient Gaussian support and making them difficult to recover through subsequent optimization. To address these limitations, we propose a novel iterative refinement framework based on a video diffusion model to improve the completeness and consistency of dynamic 4D scenes. Specifically, we first estimate multi-view depth maps and fuse them into dense point clouds to provide more complete geometric initialization for a dynamic 4DGS representation. We then employ a pretrained video restoration model to refine sequences rendered along novel camera trajectories at different time steps. The restored sequences serve as pseudo-supervision to regularize and iteratively refine the 4DGS representation. Experiments on a widely used benchmark dataset demonstrate that our method substantially outperforms existing baselines, achieving nearly a 2 dB PSNR improvement over the previous best-performing method.
Haitao Huang, Shenghao Zhao, Boyuan Tian +7
Goertek Alpha Labs Shanghai, China · Singapore Institute of Technology Singapore, Singapore · Goertek Alpha Labs Santa Clara, USA +1
Gaussian Splatting has achieved remarkable progress in multi-view surface reconstruction, yet it exhibits notable degradation when only few views are available. Although recent efforts alleviate this issue by enhancing multi-view consistency to produce plausible surfaces, they struggle to infer unseen, occluded, or weakly constrained regions beyond the input coverage. To address this limitation, we present VidSplat, a training-free generative reconstruction framework that leverages powerful video diffusion priors to iteratively synthesize novel views that compensate for missing input coverage, and thereby recover complete 3D scenes from sparse inputs. Specifically, we tackle two key challenges that enable the effective integration of generation and reconstruction. First, for 3D consistent generation, we elaborate a training-free, stage-wise denoising strategy that adaptively guides the denoising direction toward the underlying geometry using the rendered RGB and mask images. Second, to enhance the reconstruction, we develop an iterative mechanism that samples camera trajectories, explores unobserved regions, synthesizes novel views, and supplements training through confidence weighted refinement. VidSplat performs robustly to sparse input and even a single image. Extensive experiments on widely used benchmarks demonstrate our superior performance in sparse-view scene reconstruction.
Jimin Tang, Wenyuan Zhang, Junsheng Zhou +5
School of Software, Tsinghua University, China · Kuaishou Technology, China · Department of Computer Science, Wayne State University, USA