Online reconstruction of dynamic 4D scenes from long, unposed streaming videos requires both continuous processing and photorealistic rendering, which existing methods struggle to achieve simultaneously. Existing feed-forward Gaussian methods are restricted to offline processing, whereas online point-cloud approaches struggle to maintain dense geometry and high-fidelity rendering. We present DynStream, a framework for streaming 4D Gaussian reconstruction from long, unposed videos. Given a continuous video stream, DynStream reconstructs the scene within local temporal windows and incrementally aligns and fuses these local reconstructions into a globally consistent scene, enabling online 4D reconstruction without per-scene optimization. By jointly enforcing cross-window geometric consistency and modeling time-varying scene content, DynStream supports efficient reconstruction and photorealistic rendering over extended video streams. Experiments demonstrate that DynStream enables high-fidelity online dynamic reconstruction and rendering from long video streams, achieving state-of-the-art performance across diverse dynamic indoor and outdoor scenes.
Figures & tables
Figure 1: DynStream enables online, photorealistic 4D reconstruction from long, unposed video streams without per-scene optimization. It incrementally integrates incoming observations into a globally consistent representation, supports arbitrarily long streaming inputs, and generalizes across diverse dynamic and static indoor/outdoor scenes.
Figure 2: Overview of DynStream. The incoming video stream is processed sequentially in fixed-length chunks. For each chunk, the causal stream backbone predicts keyframe-relative camera poses, a local Gaussian scene, and time-varying Gaussian primitives. The local scene is aligned to the global coordinate system and incrementally fused with the previously reconstructed scene, forming a persistent global representation. Time-varying Gaussians preserve their temporal associations and are incorporated into the global scene only at the corresponding time. This design enables continuous reconstruction and rendering of long dynamic video streams while maintaining globally consistent static geometry and temporally coherent dynamic content.
Figure 3: Qualitative comparison of NVS rendering results on the Waymo Open Dataset. Each column shows one scene, and each row corresponds to GT, Ours , DGGT, and NeoVerse, respectively.
Real-world Dynamic
Synthetic Dynamic
Method
Waymo
SpatialVID
KITTI
PointOdyssey
TartanAir
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
AnySplat
24.109
0.7822
0.1998
22.096
0.7581
0.1528
16.003
0.4768
0.6493
20.554
0.6911
0.3157
22.302
0.6664
0.3025
DGGT
24.396
0.8125
0.1462
21.232
0.7350
0.1569
17.010
0.6497
0.2911
17.148
0.5704
0.4264
20.537
0.6488
0.2988
MVSplat
18.474
0.5291
0.3776
15.965
0.4217
0.3917
14.845
0.3697
0.3322
17.973
0.6093
0.4369
16.991
0.4731
0.4868
NeoVerse
19.732
0.6067
0.3596
17.400
0.5009
0.2982
13.469
0.4817
0.4432
21.868
0.6913
0.2834
18.932
0.5762
0.3259
Table 1: Rendering quality comparison on dynamic-scene datasets. For each scene, metrics are first averaged over 14 uniformly sampled frames, followed by an equal-weight average across scenes. Best results are shown in bold , and second-best results are underlined .
Figure 4: Qualitative comparison of novel-view synthesis (NVS) results across diverse scene types. Each column shows a scene, while each row corresponds to the ground truth or a baseline method.
Method
Sim(3) ATE (m) ↓
SE(3) ATE (m) ↓
Raw ATE (m) ↓
Raw RPE-t (m) ↓
Scale-correctedRPE-t (m) ↓
RPE-r (deg) ↓
STream3R
227.226
236.056
410.016
1.122
11.788
6.451
StreamVGGT
205.286
236.072
410.059
1.088
18.370
10.867
VGGT-SLAM
219.499
416.326
582.679
342.319
39.076
9.496
Ours
45.627
142.064
283.220
0.614
0.255
0.274
Table 2: Camera tracking performance on the KITTI Odometry benchmark. We evaluate on all 11 sequences comprising 23,201 frames and report sequence-level macro-averaged trajectory errors. Lower is better for all metrics. Best results are shown in bold , and second-best results are underlined .
Waymo
SpatialVID
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
DynStream
26.777
0.8675
0.0972
23.994
0.8213
0.1034
DynStream w/o dynamic
23.011
0.7662
0.1814
21.161
0.6529
0.1981
Table 3: Ablation of dynamic-aware scene modeling on Waymo and SpatialVID. Removing the dynamic component consistently degrades rendering quality. Best results are shown in bold .
KITTI
VKITTI2
Method
Sim(3) ATE ↓
SE(3) ATE ↓
Sim(3) ATE ↓
SE(3) ATE ↓
DynStream
45.63
142.06
10.33
25.05
DynStream w/o scale head
49.66
223.80
17.76
85.00
DynStream w/o relative pose
46.58
155.29
19.84
37.99
Table 4: Ablation of relative-pose estimation and metric-scale prediction on KITTI and VKITTI2. We report Sim(3)- and SE(3)-aligned absolute trajectory error (ATE).
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Scene Type
Dynamic
Real / Synthetic
# Frames
# Scenes
SpatialVID Wang et al. (2026a)
Indoor / Outdoor
✓
Real
7,089 h
2.7M clips
Waymo Sun et al. (2020)
Outdoor
✓
Real
390K
2,030
PointOdyssey Zheng et al. (2023)
Indoor / Outdoor
✓
Synthetic
∼ 200K
159
DL3DV-10K Ling et al. (2024)
Indoor / Outdoor
✗
Real
51.2M
10,510
HM3D Ramakrishnan et al. (2021)
Indoor
✗
Real
–
1,000
Replica Straub et al. (2019)
Indoor
✗
Real
–
18
Appendix
Table I: Overview of the datasets used for training.
Dataset
Method
Scenes / Frames
IoU ↑
Coverage ↑
Precision ↑
Dice ↑
AP ↑
Sec./Frame ↓
Waymo
DynStream
100 / 1400
72.00
97.74
72.97
83.22
96.58
0.0243
Waymo
DGGT
100 / 1400
34.88
39.13
76.68
46.69
63.37
0.0329
Waymo
NeoVerse
100 / 1400
23.02
35.40
48.15
33.09
30.51
0.0386
SpatialVID
DynStream
100 / 1400
58.82
95.79
59.91
72.69
91.86
0.0251
SpatialVID
NeoVerse
100 / 1400
28.25
44.85
47.72
39.16
37.77
0.0462
SpatialVID
MoVieS
100 / 1400
12.40
51.48
18.68
19.70
36.04
0.0980
Appendix
Table II: Consistency with SAM3-generated foreground reference masks.
Dataset
Method
Scenes
IoU ↑
Coverage ↑
Precision ↑
Dice ↑
Waymo
DynStream
100
83.31
88.90
92.74
90.61
Waymo
DGGT
100
35.45
40.67
74.15
47.52
SpatialVID
DynStream
100
74.93
82.46
87.28
84.32
Appendix
Table III: Dynamic segmentation using a common probability threshold of 0.5 .
End-to-end runtime under different chunk numbers ↓
Method
2
3
5
100
DGGT
300.4
308.8
405.2
Failed
MVSplat
410.2
505.2
568.6
Failed
NeoVerse
1482.3
320.3
1313.2
Failed
Ours
753.4
444.6
639.3
315.5
Appendix
Table IV: Average end-to-end runtime across all available online chunk experiments. Runtime is measured as milliseconds per input frame. Lower is better.
Streaming 4D reconstruction has been demonstrated only indoors, on dense camera rigs surrounding subjects that move at human pace. Outdoor 4D reconstruction exists but relies either on cameras mounted on the moving vehicle itself, or on limited-coverage arrays observing quasi-static subjects offline. The case that actually matters for spectators is a fast-moving subject, watched from a sparse ring of allocentric cameras, streaming. No method targets this, and no benchmark exists to evaluate one. To this end, we introduce FastFlowGS, a streaming 4D Gaussian Splatting method for reconstructing fast-moving subjects from a small set of fixed external cameras, and Monaco4D, a photorealistic Unreal Engine 5 benchmark for high-speed outdoor reconstruction. FastFlowGS fuses sparse matches, semi-dense tracks, and dense optical flow by lifting each signal to 3D with geometric uncertainty and combining them through a Kalman-style temporal update. Monaco4D provides Formula 1 sequences under varied illumination from trackside, onboard, and drone viewpoints with dense ground truth. On CMU-Panoptic, FastFlowGS exceeds the strongest baseline by 12.6% VMAF at 35% greater efficiency. On Monaco4D, where existing streaming methods degrade severely, it improves dynamic-region PSNR by up to 18.6% with 28.3% lower per-frame optimization time. Dataset and additional details can be found at https://humansensinglab.github.io/monaco4d/.
Saswat Subhajyoti Mallick, Riu Cherdchusakulchai, Marc Ruiz Olle +4
Dynamic scene reconstruction from monocular video remains a fundamental challenge in computer vision. Existing feed-forward methods predict 3D Gaussians pixel-wise for each frame, suffering from duplicated Gaussians and view-dependent biases that hinder effective learning of scene motion. We present C4G, a feed-forward 4D reconstruction framework built upon a compact set of timestamp-conditioned learnable Gaussian query tokens. Each token aggregates corresponding features across the full temporal context and decodes a 3D Gaussian whose position is modulated by the target timestamp, enabling globally coherent motion modeling without per-scene optimization. To capture fine-grained details, we further introduce a video diffusion model-based rendering enhancement module. Since our framework effectively aggregates features into Gaussians, we extend this capability to feature lifting, producing a 4D feature field that supports point tracking and dynamic scene understanding. C4G achieves strong novel-view synthesis performance using significantly fewer Gaussians and without requiring camera poses, while exhibiting stronger motion modeling and robustness to large temporal gaps.
Current 4D generation paradigms are often bottlenecked by a sequential decoupling design: video is generated first, followed by 3D reconstruction, leading to high interaction latency. This limits applications in interactive real-time scenarios. To this end, we propose \textbf{Streaming4D}, a tightly coupled synchronous pipeline that integrates block-wise autoregressive video generation with incremental 3D reconstruction. Unlike traditional frame-by-frame emission and delayed geometry recovery, Streaming4D generates temporal video blocks and immediately triggers reconstruction for each completed block, enabling parallel execution between synthesis and geometric updates. This approach allows the world representation to evolve online with the video stream, reducing feedback latency while preserving geometric fidelity. We instantiate \textbf{Streaming4D} using a Self-Forcing-style autoregressive generator and an incremental reconstruction backend. Experiments show consistent runtime improvements across resolutions on a single RTX 4090 (1.24× speedup), while maintaining high-quality 4D geometry and multi-view consistency.
Xiaoyan Liu, Jiaxin Liu, Kangrui Li +1
The Chinese University of Hong Kong · The Hong Kong Polytechnic University · The University of New South Wales +1