Efficient Gaussian Splatting Sequence Compression with Standard Video Codecs
Authors: Qi Yang, Shuting Xia, Le Yang, Geert Van Der Auwera, Zhu Li
Organizations: School of Science and Engineering, University of Missouri - Kansas City · Cooperative Medianet Innovation Center, Shanghai Jiaotong University · Electrical and Computer Engineering, University of Canterbury · Qualcomm
This paper presents a novel effective Gaussian Splatting (GS) sequence Compression method that utilizes the Video codec (GSCV). Existing video-based GS sequence compression relies on the Parallel Linear Assignment Sorting (PLAS) and tracked primitive information to convert GS into smooth 2D videos. However, tracked information is not available for most practical applications, and without it, using the vanilla PLAS can generate images exhibiting weak inter-frame correlation, due to its stochastic nature. GSCV incorporates a simple yet efficient Inter-PLAS method to produce close images between the I- and P-frames of GS, enhancing the inter-frame performance of video codec greatly. GSCV also realizes a new pipeline based on the state-of-the-art video codecs with high bit-depth GS images, achieving higher compressibility while simultaneously providing a higher quality upper bound. Experimental results show that the proposed GSCV exhibits obviously improved performance over MPEG video and point cloud-based anchors in GS sequence compression. The code is available at https://github.com/Qi-Yangsjtu/GSCV.
Figures & tables
Figure 1 . Examples of current inter-frame GS coding based on the video codecs.
Figure 2 . First row: PLAS results of “bartender” frames 1 and 2. Second row: frame 1/2 color DC and scale PLAS image.
Figure 3 . Framework of Inter-PLAS.
Figure 4 . PLAS loss and block size variation curve of “bartender” frame 1 (I-frame) and 2 (P-frame).
Figure 5 . Anchor-based PLAS refinement. First row: PLAS results for the 1st frame of “bartender”; second row: PLAS results for Pini′ ; third row: the proposed anchor-based PLAS refinement for Pini′ .
Figure 6 . Diagram of the proposed GSCV.
Figure 7 . Performance comparison on MPEG dataset. “T” and “ST” mean tracked and semi-tracked.
Figure 8 . Ablation study on (a)-(b): PLAS initialization and (c): effectiveness of Inter-PLAS.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
QP
Rate
Coordinate
Color DC
Color SH (degree 1/2/3)
opacity
scale
rotation
R05
lossless
0
0/0/0
0
0
0
R04
7
7/12/17
7
7
2
R03
17
17/22/27
17
7
2
R02
37
22/27/32
17
12
7
R01
47
32/37/42
22
12
17
Appendix
Table 1 . QP of five bitrates of GSCV
Figure 9 . PSNR performance comparison on MPEG tracked and semi-tracked dataset.
Sequence
Tracking mode
Proposed (HEVC)
Proposed (VVC)
Ref. GSCodec
Ref. GPCC
Ref. GSCodec
Ref. GPCC
Bartender
Tracked
−52.18%
−14.81%
−51.64%
−13.59%
Semi-tracked
−47.33%
−10.18%
−45.86%
−6.70%
Breakfast
Tracked
−50.01%
−13.81%
−50.29%
−13.43%
Semi-tracked
−46.98%
−11.06%
−45.29%
−7.60%
Cinema
Tracked
−51.64%
−12.75%
−51.81%
−12.60%
Appendix
Table 2 . BD-Rate results of the proposed method under HEVC and VVC configurations. Negative values indicate bitrate savings relative to the corresponding reference method.
Primitive Number
Sequence
Bartender
Breakfast
Cinema
Frame Index
Tracked
Semi-tracked
Tracked
Semi-tracked
Tracked
Semi-tracked
1
570255
531263
426968
2
570104
530997
426752
3
569842
530998
426497
4
569733
530857
426209
Appendix
Table 3 . Number of primitives for MPEG tracked and semi-tracked sequences.
Figure 10 . Bitrate allocation of GSCV.
Figure 11 . Coding time of GSCV.
Figure 12 . Additional RD curves of GSCV on tracked “bartender”.
Figure 13 . Influence of data shuffle on GSCodec Studio on tracked dataset.
Figure 14 . Ablation study of anchor-based PLAS refinement.
Figure 15 . Influence of feature channels on Inter-PLAS
Figure 16 . RD curves of single-frame results.
Figure 20
Figure 19 . Illustration of residual map and statistic histogram.
Figure 20 . RD curves of different components.
Component
QP
Color DC
0
22
27
37
47
Color SH
degree1
0
7
17
22
32
degree2
0
12
22
27
37
degree3
0
17
27
32
42
Opacity
0
17
22
27
32
Scaling
0
7
12
17
22
Appendix
Table 4 . QP of different components
Figure 21 . RD curves of different components with smooth QP variation.
Codec
GSCV-HM18.0
GSCV-VTM23.11
GSCV-libx265
GSCV-libx264
GSCodec Studio
GPCC v1
Rate Point
enc
dec
enc
dec
enc
dec
enc
dec
enc
dec
enc
dec
5
1183.5
3.9
11288.9
4.2
4.6
1.2
0.8
0.6
8.9
1.4
304.1
86.0
4
1012.5
2.9
11234.2
3.3
4.1
0.9
0.7
0.6
8.5
1.4
298.0
88.4
3
540.3
2.1
6385.6
2.3
2.8
0.7
0.6
0.5
7.3
1.3
291.7
86.7
2
353.4
1.6
3482.8
1.7
2.3
0.6
0.5
0.5
6.4
1.3
291.9
86.0
1
301.5
1.3
2391.8
1.2
1.9
0.5
0.5
0.5
5.2
1.1
293.2
85.4
Appendix
Table 5 . Encoding and decoding time of different codecs on one GoP (8 frames)
Figure 22 . RD curves of using FFMPEG codecs
Figure 23 . RD curves of A-3DGS methods.
Figure 24 . RD curves on “dance” and “basketball”.
Figure 25 . RD curves of using the Sandwich network.
Feed-forward 3D Gaussian Splatting (3DGS) enables scalable scene reconstruction without per-scene optimization, yet produces dense Gaussians that are costly to store and transmit. Existing feed-forward Gaussian compression methods formulate decoding as deterministic representation recovery, which becomes inadequate at low bitrates when high-frequency textures and view-dependent appearance are discarded. Although generative models offer a promising alternative, using them as standalone post-processing decouples generation from the transmitted scene structure, thereby compromising cross-view consistency. To address these limitations, we propose GenSplatCodec, a unified feed-forward Gaussian codec that reformulates low-bitrate Gaussian compression as geometry-guided generative decoding. We present a detail-aware feed-forward Gaussian coding scheme within a dual-stream formulation, where the resulting compact Gaussian structural stream is complemented by a lightweight reference appearance stream. We further introduce a geometry-guided one-step generative decoding approach that jointly exploits decoded structural and appearance cues through hierarchical geometry control to reconstruct high-fidelity and view-consistent novel views. Finally, we develop a three-stage optimization strategy that stabilizes the learning of the unified codec and adapts the generative decoder to codec-derived structural and appearance cues. Extensive experiments across multiple datasets demonstrate that GenSplatCodec consistently achieves superior rate-distortion (RD) performance over existing methods.
Qiang Hu, Zhenlong Wu, Lei Huang +3
Cooperative Medianet Innovation Center, Shanghai Jiao Tong University, Shanghai, 200240, China
While feed-forward 3D Gaussian splatting reconstructs renderable Gaussian primitives from sparse context views without per-scene optimization, existing pipelines do not provide a compact scene representation for storage or transmission. A natural solution is to apply existing 3DGS compression methods to the generated Gaussian primitives. However, this approach operates on the final irregular 3D representation and is decoupled from the internal feature-to-Gaussian generation process, which limits compression efficiency. To address this, we introduce CodecSplat, an ultra-compact latent coding framework for feed-forward 3D Gaussian splatting. CodecSplat first encodes an intermediate 2D Gaussian-generation feature into an entropy-coded scene bitstream. At the decoder, the latent feature is reconstructed and used to predict depth and Gaussian parameters, which are then mapped to 3D Gaussian primitives. Note that, by integrating compression into the feed-forward Gaussian generation pipeline, CodecSplat avoids inefficient compression over irregular 3D Gaussian primitives and allows the codec to exploit the structured intermediate feature representation. We instantiate CodecSplat on a feed-forward Gaussian splatting backbone with depth-guided multi-view feature refinement and a hierarchical learned feature codec. On DL3DV and RealEstate10K datasets, CodecSplat achieves 23.56-26.36 dB and 24.76-27.05 dB PSNR with only 20.00-107.77 KiB and 3.37-12.51 KiB per scene, respectively. This is roughly one order of magnitude smaller than compressing feed-forward generated Gaussian primitives, while preserving controllable rate-distortion behavior.
Pengpeng Yu, Runqing Jiang, Qi Zhang +3
Sun Yat-sen University, China · Pengcheng Laboratory, China · Peking University, China
Dynamic 3D Gaussian Splatting (3DGS) holds great promise as a 3D video streaming technology since it can represent complex 3D scenes with high fidelity. In this approach, every frame in a 3D video represents the environment as a collection of Gaussians with position and other attributes such as scale, rotation, opacity, and color. Frames capture fine details, permit views from any arbitrary perspective, but are an order of magnitude, or more, larger than 2D video frames. A line of recent work has explored how to compress dynamic 3DGS frames, but these approaches are often slow, in part because their compression techniques are not amenable to efficient acceleration. GS-NFS accelerates dynamic 3DGS compression and decompression on a GPU, to the point where it can encode and decode at full frame rate. It achieves this by developing novel GPU-based parallelizations of existing algorithms for encoding both positions and attributes of Gaussians. As a result, it is 1-2 orders of magnitude faster than the state-of-the-art in encoding and decoding a frame, while offering competitive compression performance and rendering quality.