Mobile-4DGS: Unified Static-Dynamic Real-time Mobile Gaussian Splatting
Authors: Xiaobiao Du, Beixi Hao, Zhen Fang, Tianqing Zhu, Richard Hartley, Xin Yu
Organizations: University of Technology Sydney, Sydney 2009, Australia · Yale University, New Haven, 06520, USA · City University of Macau, Macau, 999078, China · Australian National University, Canberra, 2600, Australia · Adelaide University, Adelaide, 5005, Australia
Recent advances in 3D Gaussian Splatting (3DGS) have achieved remarkable performance in novel view synthesis, yet deploying both static and dynamic Gaussian representations on resource-constrained mobile devices remains challenging due to heavy storage, redundant primitives, and costly per-frame computation. We present Mobile-4DGS, a unified lightweight framework for high-fidelity real-time static and dynamic Gaussian rendering on mobile platforms. For compact appearance modeling, we introduce a Monte Carlo Specular Energy Aggregator that compresses high-order radiance residuals into the first-order Spherical Harmonics (SH), together with an Attribute-Conditioned SH Enhancement module whose predicted offsets are pre-baked before inference. We further propose a Multi-View Alpha-Based Densification and Pruning strategy to suppress redundant primitives while maintaining multi-view consistency. For dynamic scenes, we develop a compact explicit 4D representation by constructing second-order Gaussian motion, learnable temporal support, and a binary static-dynamic partition, enabling continuous-time modeling without runtime deformation networks. Based on this partition, a Depth-Order Certificate selectively reuses previously committed depth orders to reduce re-projection, sorting, merging, and index-buffer updates during playback. Extensive experiments on static and dynamic scenes demonstrate that Mobile-4DGS substantially reduces storage and rendering overhead while maintaining competitive visual quality, enabling real-time 3D and 4D Gaussian Splatting on mobile devices. \textcolor{magenta}{Code has been released: https://xiaobiaodu.github.io/mobile-4dgs-project/}.
Figures & tables
Fig. 1: (a)(b) Mobile-4DGS achieves rendering quality comparable to both 3DGS [ 1 ] and Mobile-GS [ 2 ] , while reducing the number of Gaussian primitives and achieving significantly higher FPS on a mobile device with a Snapdragon 8 Gen 3 GPU. (c) The proposed Mobile-4DGS utilizes WebGL to enable seamless cross-platform rendering and support both static and dynamic novel view rendering.
Fig. 2: Gaussian parameter distribution and Spherical Harmonic fidelity analysis. (a) Per Gaussian memory footprint across 3DGS variants. Mobile-4DGS achieves significant compression (61% and 26% reductions) by optimizing Spherical Harmonics (SH) coefficients and decoupling SH into the base and view-independent components. (b) Qualitative comparison demonstrates that Mobile-4DGS with only first-order SH can render high-fidelity high-frequency details comparable to 3DGS.
Fig. 3: Overview of the Mobile-4DGS framework. Our method optimizes third-order SH for the initial 3k iterations, then transitions via Monte Carlo Specular Energy Aggregator for high-frequency representation. With the first-order direction moments inherited from the original third-order SH, we leverage a neural network to aggregate these latents into first-order SH c′ for rendering. During inference, the model requires only a one-time decoding step to obtain the first-order SH coefficients, significantly reducing storage while introducing no per-frame decoding overhead.
Fig. 4: Illustration of our proposed Multi-view Alpha-based Densification and Pruning strategy. (a) Traditional 3DGS leverages the single-view Gaussian position gradient determinist. (b) We propose to employ a multi-view loss-driven mechanism, coupled with alpha-based Gaussian evaluation to densify more Gaussians in the bad-reconstructed regions and moderately prune Gaussians for well-reconstructed areas.
Fig. 5: Overview of the temporal modeling in Mobile-4DGS. Each Gaussian is equipped with an explicit second-order motion state (vi,ai,ti′,ρi) and a learned binary static–dynamic gate gi . Given a query time t∈[0,1] , its center is updated from the canonical state by Eq. ( 19 ), so that the displacement of a primitive with gi=0 is identically zero rather than merely small. A Gaussian temporal window, parameterized by the learnable duration σi , modulates its effective opacity through Eq. ( 22 ), enabling smooth appearance and disappearance while handling occlusions and topology changes. The entire formulation is differentiable with respect to the motion and temporal parameters, and the committed partition exposes an exactly time-invariant subset that can be cached and compressed separately.
frame
worker
main thread
payload
committed, Eq. ( 33 )
O(N+B) : projection, sorts, merge
O(N) index upload
4N bytes
reused, Eq. ( 41 )
O(1) certificate
—
O(1)
TABLE I: Cost of one playback frame. The B=216 counting-sort scratch space is reused across frames, while the static stream is re-projected and re-sorted only when the camera depth mapping changes.
Category
Indoor
Outdoor
Method & Metrics
PSNR ↑
#G ×106↓
Storage (MB) ↓
FPS ↑
Train (min) ↓
PSNR ↑
#G ×106↓
Storage (MB) ↓
FPS ↑
Train (min) ↓
3DGS [ 1 ]
30.41
1.45
478
-
27
24.61
3.14
1361
-
36
3DGS*
30.04
1.45
46
11
27
24.39
3.14
85
5
36
Speedy-Splat [ 27 ]
30.11
0.37
64
21
14
24.41
0.56
88
14
13
C3DGS [ 28 ]
30.01
0.75
21
18
31
24.38
0.91
34
13
45
LocoGS-S [ 73 ]
30.08
0.82
6.1
13
46
24.31
1.41
10
19
61
TABLE II: Comprehensive evaluation on the mobile device with Snapdragon 8 Gen 3 GPU on the Mip-NeRF 360 dataset [ 79 ] . To facilitate a comprehensive analysis of Flux-GS, scenes are categorized into indoor and outdoor subsets. #G denotes the number of Gaussian primitives. 3DGS* represents the quantized version through Huffman encoding for mobile rendering. Mobile-GS* denotes the version of Mobile-GS [ 2 ] without MLP in the inference stage. We compare with this version without MLP for fairness. The best and second-best results are highlighted.
Dataset
Tanks&Temples
Deep Blending
Method & Metrics
PSNR ↑
SSIM ↑
LPIPS ↓
Storage (MB) ↓
FPS ↑
Train (min) ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Storage (MB) ↓
FPS ↑
Train (min) ↓
3DGS [ 1 ]
23.14
0.841
0.183
358.7
-
34
29.41
0.903
0.243
697.3
-
24
Speedy-Splat [ 27 ]
23.08
0.821
0.241
62.4
16
34
29.11
0.864
0.309
71.2
24
16
C3DGS [ 28 ]
23.32
0.831
0.202
21.8
15
49
29.73
0.900
0.258
24.7
16
35
LocoGS-S [ 73 ]
23.23
0.837
0.204
6.8
21
65
29.76
0.903
0.251
7.8
14
47
Mobile-GS [ 2 ]
23.09
0.831
0.208
2.5
124
82
29.93
0.906
0.243
4.6
135
75
TABLE III: Quantitative evaluation of state-of-the-art light-weight Gaussian-based methods on the real-world datasets. We report performance on the Tank&Temples [ 80 ] and Deep Blending [ 81 ] datasets.
Fig. 6: Qualitative and efficiency comparison with previous state-of-the-art methods. We compare rendering quality, Gaussian number, and storage costs across 3DGS, Mobile-GS, and our Mobile-4DGS. Zoomed-in regions highlight details and structural consistency for clearer differentiation. Our method achieves comparable or superior visual fidelity while using significantly fewer Gaussians and substantially lower storage for real-time rendering on mobiles. The number of Gaussians and total model size are reported below each result, demonstrating the improved efficiency–quality trade-off of our approach.
Method
SSIM ↑
SH ↓
FPS ↑
#Gaussian ↓
Storage (MB) ↓
Train (min) ↓
Real-Time4DGS [ 83 ]
0.949
3
3
2,770,350
1701
374
OMG4 [ 78 ]
0.942
3
12
283,297
5.8
417
Mobile-4DGS ( Ours )
0.945
1
47
285,707
3.8
62
TABLE IV: Quantitative performance on N3DV [ 82 ] dataset. We report FPS results on an iPhone 14 smartphone with the A15 Bionic chip. The best , second best , and third best results are highlighted, respectively.
Fig. 7: Qualitative and efficiency comparison with previous state-of-the-art methods for dynamic scenes on the N3DV dataset [ 82 ] . We compare rendering quality, Gaussian count, and storage overhead across Real-Time4DGS, OMG4, and Mobile-4DGS.
Method
PSNR ↑
Storage (MB) ↓
Peak VRAM ↓
FPS ↑
#Points ×106↓
Train (min) ↓
Mobile-4DGS
27.16
3.9
416
132
0.36
12
w/o MC-SEA
26.72
3.9
416
132
0.36
12
w/o Δc
26.86
3.9
416
132
0.36
10
w/o Multi-view densify
27.21
18
1725
21
1.37
21
w/o Multi-view prune
27.19
6.2
641
104
0.59
15
TABLE V: Ablation study of our Mobile-4DGS on the Mip-NeRF 360 dataset. MC-SEA denotes the proposed Monte Carlo Specular Energy Aggregator. Δc represents the SH offset predicted by our Attribute-Conditioned SH Enhancement module. We also ablate the proposed Multi-view Alpha-based Densification and Pruning.
Fig. 8: Decomposition of the spherical harmonic components. The visual results are rendered using Mobile-4DGS under different spherical harmonic (SH) configurations. The comparison illustrates the contribution of different SH components to view-dependent appearance modeling and highlights their effects on color fidelity, fine-grained details, and overall rendering quality.
Policy
reuse ↑
ms/s ↓
MB/s ↓
Δ ms/s ↓
drift ( r )
Monolithic
0.0%
124.7
23.2
+53%
0.01
Two-stream (separation)
0.0%
81.3
23.2
—
0.01
DOC certified ( r=0 )
0.6%
104.1
23.1
+28%
0.01
DOC r=0.25
65.5%
25.8
8.0
−68%
0.26
DOC r=0.5
81.0%
13.9
4.4
−83%
0.44
DOC r=1
90.2%
7.2
2.3
−91%
1.28
TABLE VI: Ablation of the depth-order certificate (200k Gaussians, 8.1% animated, 6 s clip, 60 Hz playback). ms/s is worker time spent on projection, sorting and merging. MB/s is index-buffer upload. drift is the worst depth violation of the presented order in mean Gaussian radii.
Fig. 9: Impact of Monte Carlo Sampling Points K on Reconstruction Quality. We evaluate PSNR across varying sampling densities on the Mip-NeRF 360 dataset. Performance follows an upward trend as K increases, with a notable saturation point appearing at K=2048 , indicating an optimal reconstruction fidelity.
Fig. 10: Impact of the number of cameras in multi-view alpha-based densification. We analyze the scaling impact of varying camera counts in our multi-view densification strategy on the Mip-NeRF 360 dataset. While PSNR steadily improves and stabilizes with additional views, the total number of Gaussian primitives decreases significantly, demonstrating the effectiveness of our densification in avoiding aggressive densification while maintaining high-fidelity reconstruction.
Xiaobiao Du is currently pursuing a Ph.D degree at the University of Technology Sydney, Australia. His research interests include static 3D reconstruction, 4D reconstruction, 3D object generation, and video generation. He is particularly interested in improving few-shot 3D reconstruction with generative prior.
Table 17
Beixi Hao is currently the Chief Scientist at TrustAI, leading technical strategy and product development for trustworthy AI systems. Prior to this, she was a Founding Algorithm Engineer at NPlace Inc., where she developed real-time 3D spatial reconstruction pipelines for iOS applications. She holds an M.S. in Computer Science from Yale University (2023) and a B.S. in Computer Science from Indiana State University (2020). Her research spans 3D computer vision, reinforcement learning, distributed AI infrastructure, and quantitative finance, with publications on attention-based architectures for question answering and medical dialogue diagnosis.
Table 18
Zhen Fang is the Australian DECRA Fellow, and he received the PhD degree in the Faculty of Engineering and Information Technology, University of Technology Sydney, Ultimo, Australia. He is a member of the Decision Systems and e-Service Intelligence (DeSI) Research Laboratory, Australian Artificial Intelligence Institute, University of Technology Sydney (UTS). Currently, he is the lecturer at UTS. His research interests include transfer learning and out-of-distribution learning. He has published over 60 papers in top-tier conferences and journals, e.g., IEEE TPAMI, IEEE TNNLS, IEEE TCYB, ICML, NeurIPS, and ICLR. He received the NeurIPS 2022 Outstanding Paper Award, the 2023 Australasian AI Emerging Researcher Award and NeurIPS 2025 Top Area Chair.
Table 19
Tianqing Zhu received the B.Eng. and M.Eng. degrees from Wuhan University, Wuhan, China in 2000 and 2004 respectively, and the Ph.D. degree in computer science from Deakin University, Australia, in 2014. She is currently a Professor and the Dean with the Faculty of Data Science at the City University of Macau. Before that, she was an Associate Professor with the School of Computer Science, University of Technology Sydney, and a Lecturer with the School of Information Technology, Deakin University, from 2014 to 2018. Her research interests include privacy-preserving and AI security. She has published more than 400 papers in refereed international journals and refereed international conferences proceedings, including many articles in IEEE Transactions and journals.
Table 20
Richard Hartley (Fellow, IEEE) is a member of the computer vision group with the Research School of Engineering, ANU, where he has been since January, 2001. He is also a member of the computer vision research group in NICTA. He worked with the GE Research and Development Center from 1985 to 2001, working first in VLSI design, and later in computer vision. He became involved with Image Understanding and Scene Reconstruction working with GE’s Simulation and Control Systems Division. He is an author (with A. Zisserman) of the book Multiple View Geometry in Computer Vision.
Table 21
Xin Yu received the BS degree in electronic engineering from the University of Electronic Science and Technology of China, Chengdu, China, in 2009, the PhD degree from the Department of Electronic Engineering, Tsinghua University, Beijing, China, in 2015, and the PhD degree from the College of Engineering and Computer Science, Australian National University, Canberra, Australia, in 2019. He is currently a senior lecturer with the University of Queensland. His research interests include computer vision and image processing.
Recent advances in 3D Gaussian Splatting have demonstrated unprecedented success in novel view synthesis. However, the substantial inference and storage overhead driven by high-order Spherical Harmonics (SH) are primary bottlenecks for mobile platforms. In this paper, we present Flux-GS, a real-time Gaussian Splatting method designed to achieve high-fidelity rendering with significantly reduced overhead for resource-constrained mobile platforms. We first propose a Monte Carlo Specular Energy Aggregator, sampling third-order radiance residuals and aggregating specular energy into a compact latent space. In this way, our method effectively preserves visually salient lighting features in lower-order bands without expensive distillation or pre-training. To mitigate the high-frequency details lost during compression, we introduce an Attribute-Conditioned SH Enhancement module. This module predicts Gaussian-aware offsets based on intrinsic Gaussian attributes, which enhance the first-order SH representation prior to inference, without extra inference costs. Furthermore, the original single-view gradient-based densification is prone to producing excessive Gaussians and overfitting to a certain view. We address these limitations by proposing a Multi-view Alpha-based Densification and Pruning strategy. By leveraging multi-view guidance, we ensure multi-view structure consistency and the precise removal of redundant primitives. Extensive experiments demonstrate that Flux-GS achieves substantial parameter reduction while maintaining competitive visual quality, offering a robust and scalable solution for real-time mobile rendering. Code: \textcolor{magenta}{https://xiaobiaodu.github.io/flux-gs-project/}.
Xiaobiao Du, YuAn Wang, Hao Li +3
University of Technology Sydney, Australia · Baidu Inc., China · Australian Institute for Machine Learning, Adelaide University, Australia
Dynamic 3D Gaussian splatting faces a fundamental tension between motion consistency and visual fidelity. Deformation-based approaches preserve temporal correspondence but suffer from motion over-factorization, oversmoothing high-frequency dynamics. In contrast, 4D-primitive methods capture fine visual details yet incur temporal overparameterization, breaking object identity and leading to severe storage overhead. To resolve this, we introduce Multi4D, a framework for high-fidelity dynamic Gaussian Splatting based on multi-level competitive allocation. Instead of a monolithic representation, we distribute modeling capacity across three structured levels: static structure, persistent dynamic geometry, and transient appearance primitives. Through shared rasterization and residual-driven optimization, these levels dynamically compete to explain photometric error, enabling adaptive specialization without pre-assigned decomposition. This allocation preserves long-term motion consistency while capturing fine dynamic detail, achieving state-of-the-art rendering quality and real-time performance with significantly fewer dynamic primitives. Furthermore, because our representation explicitly tracks compact persistent Gaussians over time, semantic features can be embedded afterward, enabling Multi4D to achieve state-of-the-art 4D segmentation accuracy with an order-of-magnitude speedup. Project page: https://batfacewayne.github.io/Multi4D.io/
Dynamic 4D Gaussian Splatting has emerged as an efficient representation for dynamic novel view synthesis through explicit scene modeling and real-time rendering. However, existing methods typically require dense multi-view videos for sufficient geometric constraints, making capture expensive and limiting sparse-camera deployment. Reducing input views lowers acquisition cost but weakens geometry supervision, often causing missing structures and floating Gaussians. Depth priors provide geometric cues, yet no single source offers both dense coverage and reliable geometry. Monocular depth provides dense structure but is scale-ambiguous and locally biased, whereas multi-view geometric depth provides incomplete anchors consistent with the reconstruction coordinate system. To exploit their complementarity, we propose D2-4DGS, a sparse-camera dynamic 4D Gaussian Splatting framework guided by dual-source depth priors. We align monocular estimates with valid multi-view geometric depths and verify their consistency to identify reliable geometric anchors. These verified anchors support consistency-aware pruning and depth supervision, while verified geometric depths and aligned mono-only estimates provide candidate geometry for densification in under-reconstructed regions. Finally, RGB-D joint optimization improves appearance fidelity and geometric consistency under sparse-view supervision. Across all nine dataset--view settings, D2-4DGS achieves the highest PSNR, improving by 1.33 dB on average over the best competing method in each setting.