While Dynamic Gaussian Splatting enables high-fidelity 4D reconstruction, its deployment is severely hindered by a fundamental dilemma: unconstrained densification leads to excessive memory consumption incompatible with edge devices, whereas heuristic pruning fails to achieve optimal rendering quality under preset Gaussian budgets. In this work, we propose Constrained Dynamic Gaussian Splatting (CDGS), a novel framework that formulates dynamic scene reconstruction as a budget-constrained optimization problem to enforce a strict, user-defined Gaussian budget during training. Our key insight is to introduce a differentiable budget controller as the core optimization driver. Guided by a multi-modal unified importance score, this controller fuses geometric, motion, and perceptual cues for precise capacity regulation. To maximize the utility of this fixed budget, we further introduce an adaptive static-dynamic allocation strategy that separates the Gaussian representation into static and dynamic branches and distributes the shared global capacity between them according to motion complexity. Furthermore, we implement a three-phase training strategy to seamlessly integrate these constraints, ensuring precise adherence to the target count. After training, a dual-mode hybrid compression scheme further reduces storage overhead. CDGS therefore not only strictly adheres to the specified Gaussian-count budget (error<2%) but also achieves favorable rate-distortion performance. Extensive experiments demonstrate that CDGS delivers optimal rendering quality under varying capacity limits and favorable rate-distortion performance, achieving over 3x model compression compared with the state-of-the-art method Ex4DGS.
Figures & tables
Fig. 1: Left : Our CDGS leverages differentiable budget control for precise Gaussian number regulation, achieving adaptive static-dynamic allocation across varying target numbers and optimal rendering quality. Middle : Visual comparison with state-of-the-art methods, highlighting advantages in visual quality, model size, and Gaussian count controllability (second row: actual/target counts). Right : Superior rate-distortion performance and precise Gaussian number control of our approach, outperforming the compared prior works (e.g., 4DGS [ 21 ] , STGS [ 22 ] , and Ex4DGS [ 23 ] ).
Fig. 2: Overview of the proposed CDGS framework. ( Top ) Three-Phase Pipeline: The training progresses from a Warm-up phase to establish foundational priors, through a Budget Enforcement phase where constraints are actively applied, to a final Fine-tuning phase that maximizes quality under the fixed count Ntarget . ( Bottom Left ) Adaptive Static-Dynamic Allocation: This module analyzes the distribution of motion magnitudes to identify a distribution-adaptive separation threshold τmotion . It decomposes the scene into a Static set S and a Dynamic set D(t) . ( Bottom Right ) Differentiable Budget Control: This module regulates capacity by computing a Unified Importance Score Mi . It fuses Perceptual cues Fperceptual and Geometric/Motion cues Fgeom/motion . The resulting score guides the differentiable budget loss Lbudget to precisely prune or densify Gaussians towards Ntarget .
Fig. 3: Illustration of the Differentiable Budget Controller. The controller aggregates geometric, motion, and perceptual cues to compute a Unified Importance Score Mi . This score passes through a hard-sigmoid gate to estimate the effective Gaussian count Np , which is regulated toward the current global sub-target N(r) via a quadratic budget loss Lbudget . Guided by this progressively updated constraint, the closed-loop policy performs densification on high-importance Gaussians and pruning on low-importance ones until the final target Ntarget is reached.
Fig. 4: Illustration of our dual-mode hybrid compression strategy, which separates outliers from both static and dynamic Gaussian data, then applies distinct compression approaches.
Method
PSNR ↑ (dB)
SSIM ↑
Size ↓ (MB)
Render ↑ (FPS)
User-Defined Gaussian-Count Control
K-Planes [ 52 ]
29.91
0.920
300
0.15
✕
ReRF [ 17 ]
29.71
0.918
231
2.0
✕
TeTriRF [ 86 ]
30.65
0.931
227
2.7
✕
StreamRF [ 83 ]
30.61
0.930
2280
8.3
✕
3DGStream [ 70 ]
31.54
0.942
2430
215
✕
4DGC [ 71 ]
31.58
0.943
150
168
✕
TABLE I: Quantitative comparison on the N3DV [ 54 ] dataset. The PSNR, SSIM, and rendering speed are averaged over all 300 frames for each scene. The reported model size is the total for the entire sequence. Ours-l and Ours-s use target Gaussian numbers of 300,000 and 100,000, respectively. Best and second-best results are highlighted.
Method
Target
PSNR
Static
Dynamic
Overall
Ratio
Ex4DGS [ 23 ]
-
28.79
292.2k
47.7k
339.9k
-
Ours
100k
28.53
72.4k
27.6k
99.9k
0.1%
200k
28.68
156.8k
41.0k
197.8k
1.1%
300k
28.81
244.4k
51.4k
295.8k
1.4%
400k
28.95
316.4k
81.2k
397.6k
0.6%
TABLE II: Validation of our precise control capability over the total number of Gaussians. We report the average PSNR on the coffee_martini sequence, the number of static/dynamic/overall Gaussians, and the error ratio of Gaussian number.
Dataset
Method
PSNR ↑ (dB)
SSIM ↑
Size ↓ (MB)
Render ↑ (FPS)
MeetRoom Dataset [ 83 ]
ReRF [ 17 ]
26.43
0.911
189
2.9
TeTriRF [ 86 ]
27.37
0.917
183
3.8
StreamRF [ 83 ]
26.71
0.913
2469
10
3DGStream [ 70 ]
28.03
0.921
2430
288
4DGC [ 71 ]
28.08
0.922
126
213
4DGCPro [ 72 ]
28.02
0.921
123
222
TABLE III: Quantitative comparison on the MeetRoom dataset [ 83 ] and Technicolor dataset [ 82 ] .
Fig. 5: Rate-distortion curves on different datasets, illustrating the superiority of our method over ReRF [ 17 ] , TeTriRF [ 86 ] , 4DGC [ 71 ] , 4DGCPro [ 72 ] , RD4DGS [ 66 ] , and Ex4DGS [ 23 ] .
Dataset
ReRF [ 17 ]
TeTriRF [ 86 ]
4DGCPro [ 72 ]
RD4DGS [ 66 ]
Ours
N3DV [ 54 ]
-1.99
-1.12
0.08
1.33
1.90
MeetRoom [ 83 ]
-1.84
-0.86
-0.02
-
1.72
TABLE IV: The BD-PSNR(dB) results of our CDGS, ReRF [ 17 ] , TeTriRF [ 86 ] , 4DGCPro [ 72 ] and RD4DGS [ 66 ] when compared with 4DGC [ 71 ] on different datasets.
Time
ReRF [ 17 ]
TeTriRF [ 86 ]
4DGC [ 71 ]
Ex4DGS [ 23 ]
Ours
Encode(s)
246
219
810
-
16
Decode(s)
18.3
16.8
28.2
-
0.5
Train(h)
>100
5.2
4.2
1.2
1.0
Render(ms)
497
372
5.6
7.8
5.4
TABLE V: Complexity comparison of our method with dynamic scene reconstruction and compression methods.
Fig. 6: Qualitative comparison of our CDGS against STGS [ 22 ] and Ex4DGS [ 23 ] on the N3DV [ 54 ] , MeetRoom [ 83 ] and Technicolor [ 82 ] datasets, demonstrating the performance of CDGS under different target Gaussian numbers.
Fig. 7: Visualization of densification strategies at Iter. 7k. Unlike the gradient-based baseline (c) which generates excessive redundant primitives in the background, our importance-driven method (d) effectively suppresses irrelevant growth, concentrating Gaussians strictly on geometric and motion boundaries.
PSNR(dB) ↑
Size(MB) ↓
Ratio ↓
w/o Importance score
31.91
31.2
1.6%
w/o Fgeom/motion
31.97
31.5
1.3%
w/o Fperceptual
31.99
31.3
1.4%
λgm=1.5
32.08
31.4
1.3%
λgm=2.5
32.11
31.5
1.4%
w/o Budget loss
32.15
31.7
4.8%
TABLE VI: Ablation study of differentiable budget control and adaptive static-dynamic allocation.
N3DV [ 54 ]
MeetRoom [ 83 ]
PSNR(dB)
Size(MB)
PSNR(dB)
Size(MB)
w/o ∇iμ
32.05
31.8
29.08
10.6
w/o αimax
32.07
31.5
29.07
10.3
w/o di−1
32.09
31.5
29.12
10.7
w/o λmax(Σi)
32.08
31.6
29.10
10.4
w/o Moi
32.02
31.3
29.02
10.5
TABLE VII: Component-wise ablation of the Unified Importance Score on N3DV [ 54 ] and MeetRoom [ 83 ] . PSNR and model size are averaged over all scenes.
Fig. 8: Evolution of Gaussian counts under varying target budgets. The subplots correspond to target limits of 1×105 ( Left ), 2×105 ( Middle ), and 3×105 ( Right ). The curves illustrate the growth trajectories of Static , Dynamic , and Total Gaussians. Under our constrained optimization framework, the total number steadily increases and converges precisely to the predefined Target Number .
Fig. 9: Visualization of our adaptive static-dynamic allocation strategy ( bottom ) against the pre-defined ratio approach ( top ).
Fig. 10: Normalized motion distributions and automatically selected static-dynamic thresholds for N3DV ( coffee_martini ), MeetRoom ( discussion ), and Technicolor ( birthday ), from left to right. The horizontal axis denotes the normalized motion statistic, the vertical axis denotes the Gaussian count in logarithmic scale, and red dashed lines indicate the selected thresholds.
Fig. 11: Comparison of storage distribution before ( left ) and after ( right ) applying our dual-mode hybrid compression on the flame_salmon sequence. The visualization highlights the significant reduction in model size and the explicit isolation of background and outlier components in the compressed representation.
PSNR(dB) ↑
Size(MB) ↓
w/o Initialization
30.90
31.8
w/o Fine-tuning
31.86
31.5
w/o Compression
32.21
98.3
w/o Static Compression
32.19
65.5
w/o Dynamic Compression
32.16
64.3
w/o Regularization Loss
31.39
31.2
TABLE VIII: Evaluation results of our three-phase training strategy and dual-mode hybrid compression.
Platform
Decoding (ms)
Rendering (ms)
RTX 3090
497
5.4
iPad M2
662
15
iPhone A15
836
32
TABLE IX: Cross-device decoding and rendering efficiency of CDGS on a desktop GPU and mobile platforms. The compressed representation is decoded once before rendering.
Recent advances in 4D Gaussian Splatting (4DGS) enable high-fidelity, real-time spatiotemporal rendering, but expose a fundamental trade-off between motion expressiveness and storage efficiency. While anchor-based designs achieve compactness through anchor-level parameter sharing, their rigid uniform parametrization enforces fixed Neural Gaussian counts and feature budgets per anchor. Consequently, insufficient fidelity is addressed by excessive anchor density, rather than lightweight, targeted increases in Neural Gaussian count or feature capacity, resulting in memory waste. To overcome this rigidity, we introduce an adaptive-capacity anchor-based framework that dynamically allocates the representational capacity based on local spatiotemporal demands. Adaptive Anchor Cardinality varies the number of Neural Gaussians per anchor, concentrating primitives in regions of high geometric or motion complexity while suppressing redundancy. In parallel, Adaptive Anchor Feature Masking modulates anchor-level feature channels, assigning rich features to complex regions and lightweight representations to simpler ones. Experiments on MPEG, Panoptic Sports, and N3DV datasets demonstrate substantial storage reduction without degrading visual quality. Notably, on challenging MPEG sequences with complex motion, our method achieves up to 1.5x higher compression than state-of-the-art anchor-based methods while preserving comparable quality.
Seunghyeon Song, Joo Chan Lee, Chanung Park +4
Sungkyunkwan University Suwon, Gyeonggi-do Republic of Korea · Electronics and Telecommunications Research Institute Daejeon, Republic of Korea · Yonsei University Seoul, Republic of Korea
3D Gaussian Splatting (3DGS) has emerged as a powerful explicit representation enabling real-time, high-fidelity 3D reconstruction and novel view synthesis. However, its practical use is hindered by the massive memory and computational demands required to store and render millions of Gaussians. These challenges become even more severe in 4D dynamic scenes. To address these issues, the field of Efficient Gaussian Splatting has rapidly evolved, proposing methods that reduce redundancy while preserving reconstruction quality. This survey provides the first unified overview of efficient 3D and 4D Gaussian Splatting techniques. For both 3D and 4D settings, we systematically categorize existing methods into two major directions, Parameter Compression and Restructuring Compression, and comprehensively summarize the core ideas and methodological trends within each category. We further cover widely used datasets, evaluation metrics, and representative benchmark comparisons. Finally, we discuss current limitations and outline promising research directions toward scalable, compact, and real-time Gaussian Splatting for both static and dynamic 3D scene representation.
Seokhyun Youn, Soohyun Lee, Geonho Kim +3
Kyung Hee University, South Korea · Kyung Hee University, Seoul, South Korea
Recent progress in 4D Gaussian Splatting (4DGS) has achieved impressive dynamic scene reconstruction results. While these methods demonstrate remarkable performance, the specific factors behind their gains remain underexplored, making a systematic understanding of the underlying principles challenging. In this paper, we perform a comprehensive analysis of these hidden factors to provide a clearer perspective on the 4DGS framework. We first establish a controlled baseline, FreeTimeGS_ours, by formalizing and reproducing the heuristics of the state-of-the-art FreeTimeGS. Using this framework, we examine 4DGS along its fundamental axes and identify practical secrets, including the emergent temporal partitioning driven by Gaussian durations and the decoupling between photometric fidelity and motion behavior. Based on these insights, we propose FreeTimeGS++, a principled method that employs gated marginalization, UFM-guided initialization, and color correction to improve stability and reproducibility. Our approach yields reproducible results with reduced run-to-run variance.