Gaussian Splatting-based Volumetric Video Compression with Sparse 4D Anchors
Authors: Ge Gao, Siyue Teng, Chanqgi Wang, Fan Zhang, Nantheera Anantrasirichai, Jui Chiu Chiang, Wen-Hsiao Peng, David Bull
Organizations: Visual Information Lab, University of Bristol, UK · National Chung Cheng University, Taiwan · National Yang Ming Chiao Tung University, Taiwan
Immersive video communication requires photorealistic, render-efficient, and compact dynamic scene representations. 3D Gaussian Splatting (3DGS) offers a promising representation, but dynamic 3DGS remains difficult to compress due to dense primitives and spatiotemporal redundancy. Anchor-based formulations improve compactness with sparse scaffolds that share geometry and appearance across primitives. However, existing designs often rely on deforming a single canonical scaffold and condition each primitive on its associated anchor in isolation, limiting their ability to handle non-local dynamics and disocclusion while under-exploiting inter-anchor correlations, particularly in motion- or texture-dense regions. To address these limitations, we propose SAGA, a volumetric video codec built upon Sparse Anchor-assisted GAussian splatting representations. SAGA represents dynamic 3D scenes using hierarchically organized sparse 4D anchors, where coordinate-based INR decoders generate fine anchors and Gaussian primitives from inter-anchor interpolations, enabling compact parameter sharing across spatiotemporal structures. For long-range dependencies among unstructured anchors, we further introduce fixed-size memory slots with orthogonality-informed updates for accurate entropy-context modeling. Experiments show that SAGA achieves strong rate-distortion performance against GIFStream, with PSNR BD-rate reductions of 80.39% and 83.94% on Neu3D and MPEG MIV, respectively.
Figures & tables
Figure 1 : ( Left ) Visual comparison on the examples from Neu3D and MPEG MIV datasets. ( Right ) RD performance on Neu3D, Panoptic Sports, and MPEG MIV in terms of PSNR.
Figure 2 : Overview of SAGA. In the representation stage, coarse anchors provide shared spatiotemporal context for fine anchors through localized kernel interpolation, while fine anchors generate Gaussian primitives via INR decoders and residual hash refinement. In the compression stage, anchors, hash grids, and neural decoder parameters are quantized and entropy coded with a hierarchical context model, where a recurrent memory bank captures long-range dependencies across sequential coding groups for improved compression efficiency.
Figure 3
Neu3D Li et al. (2022)
Panoptic Sports Joo et al. (2017)
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Train ↓
FPS ↑
Size (MB) ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Train ↓
FPS ↑
Size (MB) ↓
3DGStream Sun et al. (2024) (CVPR’24)
32.65
0.947
0.201
0.806
218
67
21.11
0.720
0.448
0.213
389
97
4DGS Wu et al. (2024a) (CVPR’24)
31.57
0.993
0.0572
120
228
202
28.68
0.911
0.157
–
–
974
E-D3DGS Bae et al. (2024) (ECCV’24)
31.42
0.945
0.037
–
–
137
25.61
0.896
0.172
–
50
297.9
STG Li et al. (2024d) (CVPR’24)
32.05
0.946
0.050
0.77
140
200
25.09
0.900
0.181
0.41
265
181
GIFStream Li et al. (2025) (CVPR’25)
31.75
0.938
0.051
0.991
95
50
28.03
0.9160
0.0868
0.731
141
53
Table 1 : Novel View Synthesis. We report PSNR, SSIM, LPIPS Zhang et al. (2018) , training time in hours, decoding FPS, and model size on Neu3D and Panoptic Sports. Missing entries are denoted by “–” when results or implementations are unavailable. For a fair comparison, we report the raw model size for GIFStream and the proposed SAGA, rather than the size after quantization and entropy coding.
Neu3D Li et al. (2022)
Panoptic Sports Joo et al. (2017)
MPEG MIV Boyce et al. (2021)
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FPS ↑
PSNR ↑
SSIM ↑
LPIPS ↓
FPS ↑
PSNR ↑
SSIM ↑
LPIPS ↓
FPS ↑
TMIV-24.0 Boyce et al. (2021)
-15.27%
-19.81%
-13.39%
35
-1.88%
-18.95%
-20.77%
45
+1.12%
-4.23%
-3.48%
42
4DGC Hu et al. (2025a) (CVPR’25)
-62.39%
-41.59%
-61.35%
110
–
–
–
–
-32.38%
-36.13%
-42.61%
115
GIFStream Li et al. (2025) (CVPR’25)
-80.39%
-88.21%
-85.01%
95
-75.88%
-77.31%
-76.95%
141
-83.94%
-86.20%
-86.74%
100
SAGA (Ours)
0.00%
0.00%
0.00%
89
0.00%
0.00%
0.00%
117
0.00%
0.00%
0.00%
91
Table 2 : 4D Volumetric Video Compression. Comparison on Neu3D, Panoptic Sports, and MPEG MIV, evaluated in terms of PSNR, SSIM, LPIPS, and decoding FPS. Each BD-rate value is computed using the corresponding benchmark codec as the anchor. We only report baselines for which BD-rate can be reliably computed (i.e., rate-distortion ranges overlap sufficiently for BD-rate calculations).
Figure 5 : RD comparison on Neu3D, Panoptic Sports and MPEG MIV in SSIM and LPIPS.
Component
Before EC MB (%)
After EC MB (%)
Runtime
Anchor
8.84 (61.2)
0.98 (66.2)
N/A
Hash
2.16 (15.0)
0.33 (22.3)
2.7ms
MLPs
3.44 (23.8)
0.17 (11.5)
2.6ms
Entropy
N/A
N/A
4.6ms
Total
14.44
1.48
9.9ms
Table 3 : Ablation, memory-slot analysis, and bitstream/runtime breakdown on Neu3D. BD-rate is computed in terms of PSNR using SAGA as the anchor.
Recent advances in 4D Gaussian Splatting (4DGS) enable high-fidelity, real-time spatiotemporal rendering, but expose a fundamental trade-off between motion expressiveness and storage efficiency. While anchor-based designs achieve compactness through anchor-level parameter sharing, their rigid uniform parametrization enforces fixed Neural Gaussian counts and feature budgets per anchor. Consequently, insufficient fidelity is addressed by excessive anchor density, rather than lightweight, targeted increases in Neural Gaussian count or feature capacity, resulting in memory waste. To overcome this rigidity, we introduce an adaptive-capacity anchor-based framework that dynamically allocates the representational capacity based on local spatiotemporal demands. Adaptive Anchor Cardinality varies the number of Neural Gaussians per anchor, concentrating primitives in regions of high geometric or motion complexity while suppressing redundancy. In parallel, Adaptive Anchor Feature Masking modulates anchor-level feature channels, assigning rich features to complex regions and lightweight representations to simpler ones. Experiments on MPEG, Panoptic Sports, and N3DV datasets demonstrate substantial storage reduction without degrading visual quality. Notably, on challenging MPEG sequences with complex motion, our method achieves up to 1.5x higher compression than state-of-the-art anchor-based methods while preserving comparable quality.
Seunghyeon Song, Joo Chan Lee, Chanung Park +4
Sungkyunkwan University Suwon, Gyeonggi-do Republic of Korea · Electronics and Telecommunications Research Institute Daejeon, Republic of Korea · Yonsei University Seoul, Republic of Korea
Open-vocabulary 3D scene understanding is commonly achieved by embedding 2D vision-language features such as CLIP into a 3D Gaussian Splatting scene, turning it into a text-queryable semantic field. However, attaching a high-dimensional feature to each of millions of Gaussians inflates a single scene to gigabytes, which makes storage and deployment the real bottleneck of these fields. Existing compact methods each learn and ship a per-scene codec, an autoencoder, a quantized codebook, or a distilled feature field, entangling field construction with field storage and never compressing the per-Gaussian assignment that holds the bulk of the cost. We argue that construction and storage should be decoupled, and that storage is a rate-distortion problem over the per-Gaussian binding to a small anchor table, a structure no prior open-vocabulary method compresses. We present CoSAG, which constructs the field without any per-scene training through a closed-form transmittance-weighted lift, spatially grounded semantic anchors, and multi-view denoising, and stores it with a spatially predictive entropy coder that ships no decoder. Because the anchors are spatially grounded, the binding is predictable and therefore highly compressible. The transmittance-weighted lift and multi-view denoising yield a clean, view-consistent assignment, so the entropy coder spends almost no rate on correcting noise and instead codes only the residual against its spatial prediction. CoSAG reaches sub-megabyte storage while matching or exceeding the state of the art across the 2D-rendered, 3D-selection, and dense-LSeg protocols, reducing field size by 37 to 76x relative to LangSplatV2 at higher accuracy.
Yuang Jia, Jinlong Wang, Junhong Lin +2
SECE, Peking University · University of Electronic Science and Technology of China
Volumetric video enables immersive free viewpoint rendering of dynamic real world scenes, yet existing methods struggle with long sequences and complex motions, often leading to temporal instability and visual artifacts. To address these challenges, we propose \ourname, a Gaussian splatting based framework for volumetric video reconstruction. Our key insight is that explicitly tracking long term complex motion with individual Gaussian primitives is inherently unstable. Instead, we organize Gaussians around time conditioned anchors that localize their spatial and temporal support, thereby reducing long range motion complexity. We further introduce a temporal windowing strategy to activate only anchors relevant to the queried time, which improves scalability and temporal coherence. In addition, to ensure spatial and temporal stability, we design a compact set of multi level anchor features that encode global features, local spatial features, and local temporal features, jointly constraining Gaussian generation. Extensive experiments demonstrate that \ourname \ consistently outperforms prior methods on long sequence volumetric videos with complex motions. Project page: https://github.com/WuJH2001/ATGS.
Jiahao Wu, Jie Liang, Die Hu +6
Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology, Shenzhen Graduate School, Peking University, Pengcheng Laboratory, China · Pengcheng Laboratory, China · Shenzhen Graduate School, Peking University, Pengcheng Laboratory, China +1