We present a 4D Gaussian Splatting framework for the Sparse-View Track of the SIGGRAPH Asia 2026 Volumetric Video Challenge, which requires dynamic scene reconstruction from only six cameras with wide baselines. To achieve robust dynamic reconstruction under such sparse views, our framework integrates three components. (1) Region-adaptive spatial priors: We use foreground masks to guide Gaussian initialization and mask voting to control densification separately for the dynamic foreground and static background. Background geometry is regularized using monocular depth aligned to metric scale. (2) Motion-consistent temporal priors: We provide supervision at intermediate times through frame interpolation and constrain projected Gaussian motion with estimated optical flow. (3) Generative assistance: We place virtual cameras in the widest angular gaps and restore their rendered images using a diffusion-based model conditioned on camera poses. The restored images are iteratively incorporated into training as pseudo-supervision. On the validation set, our framework improves full-frame PSNR from 25.60 dB for the baseline to 29.75 dB. On the official test benchmark, it achieves 30.04 dB full-frame PSNR and 27.88 dB foreground PSNR, ranking first overall in the Sparse-View Track.
Figures & tables
Figure 1. Overview of our sparse-view 4D Gaussian Splatting pipeline. Given videos from six cameras with wide baselines, our framework combines three components: (1) region-adaptive spatial priors use YOLO11 and SAM 2 masks for per-frame foreground initialization and mask voting during densification, and Depth Anything V2 predictions for background depth regularization; (2) motion-consistent temporal priors use RIFE interpolation to provide supervision at intermediate times and VideoFlow estimates to constrain projected Gaussian motion; and (3) generative assistance restores images rendered from virtual cameras in wide angular gaps using the ArtiFixer diffusion model ( de Lutio et al., 2026 ) , with the nearest training views in camera pose space as references. The restored images are iteratively incorporated into training as pseudo-supervision, yielding the final 4DGS model for novel view synthesis. Pipeline from six wide-baseline cameras to 4D Gaussian Splatting optimization. Spatial priors use foreground masks and monocular depth, temporal priors use interpolation and optical flow, and generative assistance restores virtual views before pseudo-supervised re-injection.
Full frame
Foreground
Setting
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Baseline ( θ=5e−6 )
25.60
0.9091
0.2429
27.28
0.8686
0.2276
Baseline ( θ=1e−5 )
28.08
0.9214
0.2148
26.63
0.8554
0.2539
w/o voting ( θ=5e−6 )
28.27
0.9263
0.1977
27.68
0.8741
0.2220
w/o voting ( θ=1e−5 )
28.49
0.9259
0.2017
26.92
0.8614
0.2409
w/o mask init
28.64
0.9267
0.1955
27.68
0.8743
0.2219
Table 1. Ablation study on the validation set. “Baseline” refers to the reconstruction method in ( Yang et al., 2026 ) . “Ours (recon)” includes all proposed spatial and temporal priors, while “Ours (+ gen)” additionally uses generative assistance. Parenthesized θ indicates a uniform densification threshold; otherwise, θfg and θbg are used for the foreground and background, respectively. Best and second-best results are highlighted for each metric.
Figure 2. Qualitative ablation of mask-voting densification. Qualitative comparison of dynamic foreground reconstruction without mask voting versus our full adaptive model.
Figure 3. Qualitative ablation of background depth regularization. Qualitative comparison of background reconstruction without depth regularization versus our full model.
Figure 4. Qualitative ablation of motion-consistent temporal priors. Qualitative comparison of dynamic reconstruction without temporal priors versus our full model.
Figure 5. Qualitative ablation of generative assistance. Qualitative comparison of novel-view reconstruction with and without generative assistance.