Reconstructing the Dynamic World: A Representation-Centric View of 4D Scene Reconstruction
Authors: Ziren Gong, Guo Chen, Yongjia Li, Yihua Shao, Fabio Tosi, Stefano Mattoccia, Matteo Poggi, Hao Tang, +8 more
Organizations: Department of Computer Science and Engineering, University of Bologna, Italy · Wangxuan Institute of Computer Technology, Peking University, China · Department of Computing, The Hong Kong Polytechnic University, Hong Kong · School of Computer Science, Peking University, China · Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), China · School of Electronics, Electrical Engineering and Computer Science, Queen’s University Belfast, United Kingdom · Department of Information Engineering and Computer Science, University of Trento, Italy · University of the Chinese Academy of Sciences, China · Faculty of IT, Monash University, Australia · Huawei · Google DeepMind and the University of California, Merced, United States
4D scene reconstruction aims to recover the evolving geometry, appearance, and motion of dynamic environments from visual observations. Despite substantial progress in neural scene representations, reconstructing dynamic scenes remains challenging due to non-rigid motion, occlusions, temporal inconsistencies, and the trade-offs between reconstruction fidelity and computational efficiency. Recent advances in Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have introduced diverse approaches to representing and reconstructing dynamic scenes, yet their relationships, underlying design choices, and evaluation protocols remain fragmented. In this paper, we present a unified perspective on 4D scene reconstruction, organizing existing methods around their scene representations, temporal modeling strategies, reconstruction pipelines, and optimization objectives. Through this framework, we examine how different design choices affect geometric fidelity, appearance consistency, motion representation, and computational efficiency. We further consolidate commonly used datasets and evaluation metrics, identify limitations in current experimental practices, and discuss open challenges in reconstructing complex, dynamic real-world environments. By connecting methodological developments with their underlying assumptions and evaluation evidence, this work provides a structured foundation for understanding existing approaches and identifying future research directions. An evolving collection of relevant papers and resources is available at https://github.com/ZiyangYan/Awesome-4D-Scene-Reconstruction.
Figures & tables
Fig. 1: Trends in 4D reconstruction. The increasing adoption of NeRF- and GS-based methods for dynamic scene modeling has led to a rapid growth in related publications in recent years.
Fig. 2: Comparison between NeRF and 3DGS . NeRF (left) evaluates an MLP along each ray, whereas 3DGS (right) renders by rasterizing and blending Gaussians.
Survey
4D Scene Types
Evaluation Coverage
Datasets
Methods
Taxonomy
Fan et al. [ 25 ]
Human and animal motion
–
–
90
–
Zhu et al. [ 194 ]
General
NVS, Efficiency
10
52
✓
He et al. [ 41 ]
Autonomous driving
–
–
36
–
Cao et al. [ 9 ]
General
–
–
111
–
Zhao et al. [ 186 ]
Object, human, and animal motion
–
21
62
–
Ours
General
NVS, Geometry, Efficiency
22
102
✓
TABLE I: Comparison of existing surveys on 4D scene reconstruction.
Fig. 3: General pipeline of NeRF-based 4D scene reconstruction methods. The pipeline illustrates representative strategies, including deformation-based, 4D primitive-based, and 4D feature volume-based frameworks. Temporal prior-based methods are not included due to their diversity.
Method
Venue
Inputs
Scenario
Target Domain
4D-style
Scene Encoding
Flow
Normal
Segment.
Extra Prior
D-NeRF [ 110 ]
CVPR2021
RGB
Indoor
Entity-Centric
Deformation Fields
MLP
NR-NeRF [ 138 ]
CVPR2021
RGB
In-the-Wild
Scene-Centric
Deformation Fields
MLP
STaR [ 179 ]
CVPR2021
RGB
Indoor
Scene-Centric
Deformation Fields
MLP
Nerfies [ 105 ]
ICCV2021
RGB
Indoor
Entity-Centric
Deformation Fields
MLP
HyperNeRF [ 106 ]
TOG2021
RGB
Indoor
Entity-Centric
Deformation Fields
MLP
NDR [ 7 ]
NIPS2022
RGBD
Indoor
Entity-Centric
Deformation Fields
MLP
TABLE II: Overview of NeRF-based 4D dynamic scene reconstruction methods . Methods are categorized into four types. For each method, we summarize its scene representation, key components, and additional priors.
Fig. 4: General pipeline of 3DGS-style 4D scene reconstruction methods. The pipeline presents the representative 4D strategies in explicit 4D primitive-based, deformation-field-based, and frame-wise-training frameworks.
Method
Venue
Input
Scenario
Target Domain
4D-style
Text
Flow
Normal
Segment.
Extra Prior
4DGS [ 172 ]
ICLR2024
RGB
Indoor
Multi.-Centric
Explicit 4D Primitive
CD-3DGS [ 53 ]
ECCV2024
RGB
Indoor
Entity-Centric
Explicit 4D Primitive
✓
RAFT
SpacetimeGS [ 74 ]
CVPR2024
RGB
Indoor
Entity-Centric
Explicit 4D Primitive
4DRotorGS [ 22 ]
ACM SIG. 2024
RGB
Indoor
Entity-Centric
Explicit 4D Primitive
✓
FreeTimeGS [ 148 ]
CVPR2025
RGB
Indoor
Multi.-Centric
Explicit 4D Primitive
ROMA
DeSiRe-GS [ 108 ]
CVPR2025
RGBD
Auto. Driving
Scene-Centric
Explicit 4D Primitive
✓
TABLE III: Overview of 3DGS-based 4D dynamic scene reconstruction methods. Methods are categorized into three types. For each method, we summarize its key components and additional priors.
Dataset
Scene Type
Sensor Setup
Resolution
Frame Rate
Scene/Seq
Temporal Scale
Synthetic Datasets (Ground Truth Geometry/Motion)
D-NeRF [ 110 ]
Indoor
1 Cam
800 × 800
–
8
50–200 frames
ParticleNeRF [ 1 ]
Indoor
40 Cams
–
–
6
–
SS3DM [ 45 ]
Autonomous Driving
6 Cams + 5 LiDAR
–
10 FPS
28
13K frames
Real-world: Monocular & Sparse View
DAVIS [ 109 ]
Outdoor
1 Cam
–
–
150
10k frames total
TABLE IV: Taxonomy of dynamic scene datasets based on benchmark properties. Datasets are grouped by their primary research focus and capture characteristics.
Fig. 5: Qualitative reconstruction point map and depth of NeRF-style methods on the NuScenes [ 6 ] dataset. Image from [ 178 ] .
Methods
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
DyNeRF
29.6
0.961
0.083
StreamRF
28.3
-
-
HexPlane
29.5
-
0.097
K-Planes
31.6
0.964
-
TIDNeRF
29.9
-
0.096
HyperReel
31.1
0.927
0.096
TABLE V: Neu3D [ 72 ] NeRF-style 4D reconstruction results. PSNR ( ↑ ), SSIM ( ↑ ), and LPIPS ( ↓ ) are used as metrics.
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
4DGS
32.01
-
0.055
4DRotorGS
31.62
0.940
0.140
FreeTimeGS
33.19
-
0.036
SpacetimeGS
32.05
-
0.044
CD-3DGS
30.46
0.955
0.150
SaRO-GS
32.15
-
0.044
TABLE VI: Neu3D [ 72 ] 3DGS-style 4D reconstruction results. PSNR ( ↑ ), SSIM ( ↑ ), and LPIPS ( ↓ ) are used as the evaluation metrics.
Methods
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
D-NeRF
30.50
0.95
0.070
TiNeuVox
32.67
0.97
0.041
HexPlane
31.04
0.97
0.040
K-Planes
31.61
0.97
0.049
TIDNeRF
32.73
0.97
0.033
Ced-NeRF
34.21
0.99
0.037
TABLE VII: D-NeRF [ 110 ] NeRF-style 4D reconstruction results. PSNR ( ↑ ), SSIM ( ↑ ), and LPIPS ( ↓ ) are used as metrics.
Methods
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
Deformable-3DGS
24.10
0.85
0.18
SC-GS
24.10
0.89
0.14
4D-GS
24.18
0.88
0.14
SP-GS
23.33
0.84
0.21
MotionGS
24.54
0.87
0.17
DN-4DGS
24.36
0.87
0.17
TABLE VIII: NeRF-DS [ 163 ] 3DGS-style 4D reconstruction results. PSNR ( ↑ ), SSIM ( ↑ ), and LPIPS ( ↓ ) are used as metrics.
Fig. 6: Qualitative novel view synthesis results of 4D reconstruction methods on the Waymo [ 133 ] dataset. Image from [ 15 ] .
Fig. 7: Qualitative novel view synthesis results of NeRF-style methods on the NVIDIA Dynamic Scene [ 176 ] dataset. Image from [ 34 ] .
Fig. 8: Qualitative Novel View Synthesis results of 3DGS-style framework on the Neu3D [ 72 ] dataset. Image sourced from [ 161 ] .
methods
CD ( ↓ )
F-Score ( ↑ )
RMSE ( ↓ )
NeRF-style
D-NeRF
0.33
0.85
7.11
TiNeuVox-B
0.39
0.86
7.21
K-Planes
0.30
0.89
6.80
LiDAR4D*
0.24
0.89
6.78
STGC-NeRF*
0.22
0.91
6.54
TABLE IX: NuScenes [ 6 ] 3D geometric reconstruction results. * denotes methods with LiDAR supervision; † uses protocols from [ 15 ] .
Methods
4D-style
FPS
training time (h)
Params (Mb)
NeRF-style
D-NeRF
Deformation fields
<1
22.3
3
DyNeRF
4D Primitive
<1
1344
7
NeRFPlayer
4D feature volumes
<1
6
-
HyperReel
4D feature volumes
6.1
2.2
360
MixVoxel
4D feature volumes
4.3
1.3
500
TABLE X: Performance analysis of 4D reconstruction methods . GPU memory, frame per second (FPS), and training time are evaluated.
School of Computer Science and Engineering, Sun Yat-sen University, China · Tsinghua Shenzhen International Graduate School, Tsinghua University, China