Feed-forward 3D Gaussian Splatting (3DGS) enables efficient novel-view synthesis from sparse multi-view images, yet its representations remain costly to store and transmit. Existing approaches compress either the input images, incurring heavy receiver-side reconstruction, or the reconstructed Gaussian primitives, which are difficult to compress due to their heterogeneous and irregular attributes. We instead compress compact intermediate features, providing a better balance between compression efficiency and receiver-side complexity. Based on this paradigm, we propose FeCoSplat, a feedback-guided compression framework for feed-forward 3DGS. FeCoSplat first compresses multi-view features to obtain an intermediate 3DGS, whose rendered views are used as feedback to guide a second-stage compression for further refinement. The resulting bitstreams are decoded into a compact implicit state, from which the final Gaussian primitives are reconstructed with a lightweight predictor. Experiments demonstrate that FeCoSplat achieves favorable rate--distortion performance, particularly at low bitrates, while requiring only 3.45M parameters for receiver-side Gaussian reconstruction. Code will be released soon.
Figures & tables
Figure 1: Different compression pipelines for feed-forward 3DGS. Image compression incurs high receiver-side computation and error propagation, while Gaussian compression is challenged by heterogeneous attributes and irregular primitives. Feature compression operates on compact, structured intermediate feature representations, balancing efficiency and receiver-side complexity.
Figure 2: Overview of FeCoSplat . Top: FeCoSplat performs feature compression in two stages, encoding multi-view features in Stage 1 ( k=1 ) and feedback-guided features in Stage 2 ( k=2 ). A unified feature codec integrates the decoded information into an implicit Gaussian state across stages, followed by geometry-aware 3D interaction for Gaussian reconstruction. Bottom Left: Encoding process, where reconstruction feedback is introduced to capture complementary information, and S1 serves as decoded context to reduce redundancy in the second-stage coding. Bottom Right: Decoding process, which requires only lightweight modules and avoids explicitly reconstructing the intermediate 3DGS. The encoding and decoding of camera parameters are omitted for clarity.
Methods
Params (M) ↓
Latency (s) ↓
Peak GPU Mem (GiB) ↓
Sender
Receiver
Sender
Receiver
Image Compression
ELIC → DepthSplat
26.45
62.40
0.077
0.181
1.505
ELIC → ReSplat
26.45
100.69
0.074
0.305
2.119
Gaussian Compression
DepthSplat → TinySplat
38.34
N/A *
1.285
0.375
1.365
Table 1: Comparison of sender/receiver parameters, coding latency, and peak GPU memory on RealEstate10K. FeCoSplat achieves a lightweight receiver with only 3.45M parameters and 0.117s latency, while also maintaining low sender-side complexity and the lowest peak GPU memory among the compared learning-based approaches.
Figure 3: Performance comparison.
Figure 4: Qualitative comparison on RE10K and ACID. The rate per scene (KiB), PSNR (dB), and LPIPS are reported for each result. FeCoSplat achieves favorable reconstruction quality even at extremely low bitrates compared with competing methods. More results are in Appendix D .
Figure 5: Ablation study on component design, Gaussian resolution and feedback iterations.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Detailed architectures of the building blocks and feature extractors used in FeCoSplat . The figure illustrates the designs of the 2D and 3D blocks, together with the multi-view feature extractor Emv and the feedback-guided feature extractor Efb .
Figure A2: Performance comparison with two base-size methods.
Methods
Sender Params (M) ↓
Receiver Params (M) ↓
ELIC → DepthSplat (Base)
26.45
144.81
ELIC → ReSplat (Base)
26.45
246.20
FeCoSplat
37.28
3.45
Appendix
Table A1: Comparison of model parameters at the sender and receiver.
Figure A3: Visualization of the two-stage reconstruction process of FeCoSplat .
While feed-forward 3D Gaussian splatting reconstructs renderable Gaussian primitives from sparse context views without per-scene optimization, existing pipelines do not provide a compact scene representation for storage or transmission. A natural solution is to apply existing 3DGS compression methods to the generated Gaussian primitives. However, this approach operates on the final irregular 3D representation and is decoupled from the internal feature-to-Gaussian generation process, which limits compression efficiency. To address this, we introduce CodecSplat, an ultra-compact latent coding framework for feed-forward 3D Gaussian splatting. CodecSplat first encodes an intermediate 2D Gaussian-generation feature into an entropy-coded scene bitstream. At the decoder, the latent feature is reconstructed and used to predict depth and Gaussian parameters, which are then mapped to 3D Gaussian primitives. Note that, by integrating compression into the feed-forward Gaussian generation pipeline, CodecSplat avoids inefficient compression over irregular 3D Gaussian primitives and allows the codec to exploit the structured intermediate feature representation. We instantiate CodecSplat on a feed-forward Gaussian splatting backbone with depth-guided multi-view feature refinement and a hierarchical learned feature codec. On DL3DV and RealEstate10K datasets, CodecSplat achieves 23.56-26.36 dB and 24.76-27.05 dB PSNR with only 20.00-107.77 KiB and 3.37-12.51 KiB per scene, respectively. This is roughly one order of magnitude smaller than compressing feed-forward generated Gaussian primitives, while preserving controllable rate-distortion behavior.
Pengpeng Yu, Runqing Jiang, Qi Zhang +3
Sun Yat-sen University, China · Pengcheng Laboratory, China · Peking University, China
Feed-forward 3D Gaussian Splatting (3DGS) enables scalable scene reconstruction without per-scene optimization, yet produces dense Gaussians that are costly to store and transmit. Existing feed-forward Gaussian compression methods formulate decoding as deterministic representation recovery, which becomes inadequate at low bitrates when high-frequency textures and view-dependent appearance are discarded. Although generative models offer a promising alternative, using them as standalone post-processing decouples generation from the transmitted scene structure, thereby compromising cross-view consistency. To address these limitations, we propose GenSplatCodec, a unified feed-forward Gaussian codec that reformulates low-bitrate Gaussian compression as geometry-guided generative decoding. We present a detail-aware feed-forward Gaussian coding scheme within a dual-stream formulation, where the resulting compact Gaussian structural stream is complemented by a lightweight reference appearance stream. We further introduce a geometry-guided one-step generative decoding approach that jointly exploits decoded structural and appearance cues through hierarchical geometry control to reconstruct high-fidelity and view-consistent novel views. Finally, we develop a three-stage optimization strategy that stabilizes the learning of the unified codec and adapts the generative decoder to codec-derived structural and appearance cues. Extensive experiments across multiple datasets demonstrate that GenSplatCodec consistently achieves superior rate-distortion (RD) performance over existing methods.
Qiang Hu, Zhenlong Wu, Lei Huang +3
Cooperative Medianet Innovation Center, Shanghai Jiao Tong University, Shanghai, 200240, China
3D Gaussian Splatting (3DGS) enables high-quality novel-view synthesis but requires substantial storage. Existing compression methods often rely on spatial context modeling over irregular 3D representations, increasing the complexity of training and coding. Meanwhile, floating-point context inference can introduce numerical inconsistencies across platforms, causing entropy-decoding failures. To address these practical challenges, we propose COSA-GS, which constructs context without spatial aggregation through anchor-wise causal factorization. Specifically, we use geometry context derived from each anchor's coordinates to model a compact learnable anchor latent. The anchor latent is then fused with the geometry context to form an anchor context for attribute coding. The resulting context model features a simple architecture composed solely of linear transformations and activations. We train COSA-GS using rate--distortion optimization with adaptive Gaussian pruning. Further, we develop quantization-aware training and integer inference for the context model to achieve bit-exact consistency of entropy-decoded symbols across platforms. Experiments demonstrate that COSA-GS achieves state-of-the-art compression performance while retaining fast and consistent cross-platform decoding, providing a simple yet effective framework for practical 3DGS compression. Code is available at https://github.com/pengpeng-yu/COSA-GS.
Pengpeng Yu, Yueru Chen, Fei Song +4
Sun Yat-sen University, China · Pengcheng Laboratory, China · Academy of Broadcasting Science, National Radio and Television Administration, China +1