Learned image codecs (LICs) achieve high reconstruction quality, but their decoding speed is often insufficient for immersive virtual reality (VR). Gaussian splatting (GS) codecs render much faster, yet still lag in reconstruction quality and typically decide primitive allocation without considering the coding cost of each primitive. We introduce OIC-GS, an omnidirectional GS codec with a new hierarchical HEALPix primitive grid representation. Gaussian primitives are anchored at predefined spherical locations, eliminating explicit coordinate coding. Finer levels refine their coarser ancestors, naturally supporting coarse-to-fine reconstruction and layered transmission. The predefined grid also enables efficient viewport decoding by selecting only view-relevant primitives. We further introduce a lightweight entropy model for quantized primitives and optimize the codec under a spherical rate-distortion objective. Primitives with insufficient rate-distortion benefit are automatically removed when their quantized opacity becomes zero, allowing OIC-GS to adapt both primitive density and level of detail without a fixed primitive budget. A single bitstream supports full-sphere, viewport-dependent, and progressive decoding. The first viewport reaches final quality after decoding only 52% of the bitstream, and is then rendered at 1,270 FPS. On a 100-image omnidirectional benchmark, OIC-GS outperforms all evaluated GS codecs, reducing WS-PSNR BD-rate by 49.6% over GaussianImage++ and 68.6% over SGI, which uses a learned entropy model.
Figures & tables
Figure 1: Overview of OIC-GS. Primitives are anchored on hierarchical HEALPix levels and carry a quantized feature and opacity, coded level by level. The levels are rendered from coarse to fine; a primitive whose quantized opacity is zero becomes inactive, and its cell falls back to the coarser levels. The rendering state (F(ℓ),r(ℓ)) is reconstructed identically at the encoder and the decoder.
Figure 2: The rendering-conditioned entropy model (RCC). Even cells (pass 1) are coded with the coarse and parent contexts; the decoded even cells are then added to the context of the odd cells (pass 2). The ContextMLP hψ predicts a 3-component Gaussian mixture for the arithmetic coder, and the decoder reconstructs the identical context by construction. Cell indices are illustrative.
Figure 3: Rate-distortion performance on the 100-image main benchmark. OIC-GS outperforms every GS codec at every tested bitrate on all four metrics.
Figure 4: Rate-distortion performance on the 100 SUN360 images. OIC-GS again outperforms every GS codec at every tested bitrate on all four metrics.
Figure 5: Reconstructed viewports. Every baseline uses at least as many bits as OIC-GS. The labels give the viewport PSNR and, for each 2× zoomed-in patch, the PSNR of the textured 1282 window. App. G shows five scenes in all six directions.
Main benchmark
SUN360
Method
WS-PSNR
V-PSNR
V-SSIM
V-LPIPS
WS-PSNR
V-PSNR
V-SSIM
V-LPIPS
Gaussian splatting image codecs
SGI
+77.6
+54.0
+73.1
+89.9
+84.2
+69.5
+76.1
+95.2
GaussianImage
+42.3
+22.9
+49.8
+68.0
+69.0
+56.4
+67.5
+80.2
GaussianImage++
+5.5
−9.5
+31.3
+66.5
+26.8
+16.9
+68.1
+103.6
LIG
+165.9
+95.4
+254.2
+369.7
+214.0
+159.3
+308.7
+396.2
Table 1: BD-rate ( Bjontegaard, 2001 ) against JPEG on the main benchmark and on the 100 SUN360 images; a negative value means fewer bits at equal quality.
Figure 6: Schematic comparison of first-viewport delivery over a 10 Mbps link for full-ERP codecs, tiled JPEG2000 and OIC-GS. Size: bytes transmitted before the first viewport is shown; Transmit: transfer time of these bytes.
Method
Category
Entropy decoding (ms)
Reconstruction (FPS)
JPEG ( Wallace, 1992 )
Traditional
97
JPEG2000 ( Skodras et al., 2001 )
Traditional
3.8
ELIC ( He et al., 2022 )
LIC, checkerboard context
466
22.1
MLIC++ ( Jiang et al., 2025 )
LIC, multi-reference context
621
17.4
mbt2018 ( Minnen et al., 2018 )
LIC, autoregressive
30,480
67.1
cheng2020-attn ( Cheng et al., 2020 )
LIC, autoregressive
30,490
27.4
Table 2: Decoding time and rendering speed on one 2048×1024 ERP image (one idle RTX 4090). Entropy decoding includes the context networks of the LIC rows and the rendering-conditioned contexts of OIC-GS; reconstruction is the synthesis network of the LIC rows and the rendering of the GS rows. JPEG and JPEG2000 report the full CPU decoding speed, and – marks codecs without a learned entropy model. The viewport row decodes only the streams of the first viewport. Details are given in App. D .
Table 3: Ablation studies of the main components and of the design parameters of each module on the 100 images of the main benchmark: rate difference against the full model at equal quality at λR=3×10−3 (positive is worse). The protocol, the variant specifications and the SUN360 results are given in App. E.2 .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Constant
Value
Evaluation grid of Eq. ( 4 )
target grid, nside=512 , M=3,145,728 directions
Candidate set C(ℓ)(v)
K=9 cells (the cell containing v and its 8 neighbors);
all ∣Ω(0)∣=192 primitives at ℓ=0
Denominator floor of Eq. ( 4 )
ϵ=10−8
Fill opacity in Eq. ( 3 )
αfill=0.1
Selection threshold in Eq. ( 2 )
τ=0.2
Appendix
Table 4: Rendering and selection constants.
Item
Setting
Optimizer
Adam ( Kingma and Ba, 2014 )
Learning rate, features
10−2 at the finest level, decreasing linearly with the level to 10−3 at ℓ=0
Learning rate, opacities
10−2
Learning rate, entropy model
3×10−3
Learning rate, ColorMLP
3×10−3
Learning rate, kernel widths
10−3
Appendix
Table 5: Training hyperparameters.
λR
bpp
WS-PSNR (dB)
V-PSNR (dB)
V-SSIM
V-LPIPS ↓
9.0×10−3
0.1448
23.83
27.61
0.7099
0.3750
3.0×10−3
0.2885
25.36
29.79
0.7888
0.2522
2.0×10−3
0.3843
26.05
30.72
0.8202
0.2083
1.4×10−3
0.4931
26.67
31.58
0.8469
0.1733
0.9×10−3
0.6576
27.45
32.63
0.8764
0.1372
4.7×10−4
0.9494
28.50
34.06
0.9103
0.0981
Appendix
Table 6: Operating points of OIC-GS: mean values over the 100 panoramas, measured end-to-end against the original ERP images (Sec. 4.1 ). These are the points of OIC-GS plotted in Fig. 3 . Bitrates are the written bitstream sizes in bits (bytes × 8) divided by the number of ERP pixels. The decoding measurements of App. D use the 500 trained models of the first five points.
Baseline
OIC-GS BD-rate (%) vs. baseline ↓
WS-PSNR
V-PSNR
V-SSIM
V-LPIPS
SGI
−68.6
−67.1
−57.7
−66.3
GaussianImage++
−49.6
−45.6
−47.4
−65.1
GaussianImage
−58.2
−56.1
−50.4
−66.7
LIG
−84.2
−82.1
−83.2
−88.6
Appendix
Table 7: Direct pairwise BD-rates of OIC-GS against the GS baselines, each computed on the common quality range of that pair.
Figure 7: Bit allocation over the sphere on the benchmark panoramas at matched bitrates. Bits are assigned to 64 latitude bands, divided by the solid angle of each band, and normalized by the global mean of each method; for OIC-GS, the feature bits are shown. (a) The ERP-based codecs follow the sampling density of the projection, while OIC-GS is distributed substantially more evenly; GaussianImage and LIG are omitted from (a) for readability. (b) The polar bands (the 11 outermost bands on each side, ∣lat∣>59∘ ) cover 14.2% of the sphere and receive 13.1% of the feature bits of OIC-GS, close to their area share, compared with 27.9 - 34.4% for the ERP-based codecs, close to the 34.4% ERP pixel share.
Fixed viewport
Stage (per frame)
ms
share
Level-0 splatting ( 192 primitives)
0.046
6.3%
Levels 1-6 splatting ( K=9 )
0.063
8.5%
Coarse-level downsampling
0.140
19.0%
Blending (Eq. ( 5 ))
0.022
3.0%
ColorMLP
0.443
60.2%
Appendix
Table 8: Per-frame time breakdown of the repeated rendering of a fixed viewport (one idle RTX 4090). The viewport rendering evaluates the 417,508 target-grid directions inside the padded 90∘×60∘ frustum and maps them to a 1024×768 frame that stays on the GPU; the host copy is not included. Stage times are GPU kernel times recorded with the PyTorch profiler over 200 frames, as medians over 15 models (scenes 1, 7 and 46 ×5 bitrates; two viewports, yaw 0∘ and 90∘ , per model).
Format
File size
Δ size
Saved (yaw 0∘ / 90∘ )
32 equal streams per pass (default)
116,089 B ( 0.4429 bpp)
–
–
Tile-aligned ( 4,096 cells/tile)
117,352 B ( 0.4477 bpp)
+1.1%
50.0% / 39.4%
Appendix
Table 9: Two storage formats of the same symbols on the representative scene ( 90∘×60∘ viewports). The tile-aligned format increases the file size by 1.1% ( 1.2% benchmark mean) and allows 39 - 50% of the file to be skipped for one viewport; the default format has no tile boundaries to skip.
λR
First view (bpp)
Deferred
T1 (KB)
Unchanged
T2 (FPS)
0.9×10−3
0.338
48%
82
3/3
314 / 891
Appendix
Table 10: Progressive streaming: the models of scenes 1, 7 and 46 at the λR=0.9×10−3 point of Table 6 played end-to-end in fresh processes ( 3 sessions; first viewport at yaw 0∘ , followed by a 180 -frame pan to 90∘ ). “ T1 (KB)” is the data size of phase T1 , and “ T2 (FPS)” is the rendering speed of phase T2 . “Deferred” is the fraction of the file not yet received when the first viewport is displayed in each session. “Unchanged” counts the sessions whose first-viewport pixels are bit-identical after completion; T2 gives the re-splatting/cached speeds of Table 11 .
Strategy
ms/frame
FPS
GT-PSNR
Notes
Re-splatting, visible set with margin
3.22
310
27.994
4 updates of the visible set
Cached sphere, lookup only
1.08
923
27.994
bit-identical to re-splatting
Appendix
Table 11: Head-rotation speed on the representative scene ( 180 frames from yaw 0∘ to 90∘ , 1024×768 at 90∘×60∘ , idle RTX 4090). FPS is computed from the unrounded mean frame time. Both strategies are measured in a single run that decodes once and then pans; over the sessions of Table 10 , they run at 314 / 891 FPS on average. GT-PSNR, the mean viewport PSNR along the trajectory against the ground-truth projection, is identical for all strategies, so the acceleration does not affect quality.
Figure 8: Progressive decoding by level, from the written bitstream of scene 1 at the λR=0.9×10−3 operating point ( 195.5 KB; Lanczos-3 resampling to ERP). Each panel shows the full sphere reconstructed only from the levels ℓ≤L′ , labeled with the bytes of this prefix (share of the file) and the end-to-end WS-PSNR. Because of the coarse-to-fine order of Sec. 3.2 , every prefix gives a complete sphere with less detail: a recognizable preview requires the first 10% of the file (levels ℓ≤3 ), and the final prefix is bit-exact with full decoding.
Reference round-trip
Native sphere PSNR (dB)
End-to-end ERP WS-PSNR (dB)
Scene
fidelity
L6
L7
Δ
L6
L7
Δ
87
30.921
32.952
45.466
+12.514
28.702
30.737
+2.035
88
32.016
33.349
43.963
+10.614
29.225
31.660
+2.435
84
32.659
32.515
41.484
+8.969
28.959
31.959
+3.000
86
33.541
32.290
40.666
+8.375
28.892
32.574
+3.682
73
34.341
35.009
44.354
+9.345
31.496
33.848
+2.352
Appendix
Table 12: Paired L6/L7 resolution analysis. “Reference round-trip fidelity” is the WS-PSNR of the original ERP image after the reference ERP → HEALPix → ERP round trip; it is used to stratify the image selection and as the reference score in Fig. 9 , and it is not a mathematical upper bound. Both the native and the end-to-end scores improve on all 10/10 scenes. These fits use only the distortion loss, continuous features and all primitives active, and are therefore not codec operating points or results at matched bitrates.
Figure 9: Resolution analysis. Left: the finest level at nside=512 improves the end-to-end ERP WS-PSNR for every selected image ( +2.752 dB on average; range +2.035 - +3.682 dB). Right: images with lower round-trip fidelity of the reference ERP → HEALPix → ERP path, which indicates stronger resampling loss and more high-frequency content, obtain a larger native L7 gain (Pearson correlation −0.739 ). The native gain is also correlated with the global and nadir high-frequency energy ( +0.642 and +0.721 ), and is +10.480 dB on average at the nadir against +8.022 dB at the equator.
Metric
BD-rate (%)
Δ at matched bitrate
Negative on
WS-PSNR
−7.7
+0.24 dB
20/20
V-PSNR
−8.2
+0.35 dB
20/20
native sphere PSNR
−11.7
+0.46 dB
20/20
Appendix
Table 13: The L7 codec against the L6 configuration on the 20-image subset of App. F.2 , at the three operating points λR∈{2.0,1.4,0.9}×10−3 , computed with the estimator of Table 1 on the common quality range of each pair. A negative value means fewer bits at equal quality; “Negative on” counts the images with a defined BD-rate. Native sphere PSNR is measured on the HEALPix representation and excludes the effect of resampling.
Level 7
Selection ratio, L6 config → L7 config
Active primitives
λR
median (range)
Level 4
Level 5
Level 6
L6 → L7
2.0×10−3
0.14% ( 0.02 - 0.60 )
71.9→76.1%
34.0→42.3%
8.7→6.3%
186 K →191 K
1.4×10−3
0.25% ( 0.04 - 1.02 )
73.6→77.8%
37.2→46.6%
11.8→5.2%
217 K →197 K
0.9×10−3
0.56% ( 0.11 - 1.87 )
75.4→78.6%
45.8→49.9%
18.2→8.4%
285 K →241 K
Appendix
Table 14: Selection ratios after training, once the inactive primitives are removed, on the 20-image subset of App. F.2 at the three operating points λR∈{2.0,1.4,0.9}×10−3 . Level 7 contains 3,145,728 cells; the median column gives the per-image median selection ratio (range in parentheses), and the remaining columns compare the same level between the two configurations at the same λR .
Figure 10: Locations of active primitives of the additional level on scene 87: active level-7 cells are shown in red over the ground truth. They concentrate on foliage, rocky ground and the shoreline, and are absent from the sky and open water; decreasing λR (bottom) selects more of them without changing where they are located.
Figure 11: Additional qualitative results (part 1 of 3): scene 55 in all six viewing directions, under the protocol described in the text. Across the five scenes, OIC-GS achieves the highest viewport PSNR among the GS codecs and COIN in 25 of the 30 panels, and the highest patch PSNR in 27 . On this scene, the only exception is the zenith, where GaussianImage++ and GaussianImage exceed OIC-GS by 1.7 and 1.0 dB at 0.784 and 0.613 bpp, compared with 0.578 bpp for OIC-GS; the other four exceptions are shown in Fig. 12 . OIC-GS matches or exceeds JPEG2000 in 11 of the 12 equatorial panels of scenes 7, 32 and 55, while JPEG2000 remains ahead at the nadir of every scene.
Figure 12: Additional qualitative results (part 2 of 3): scenes 74 (top) and 46 (bottom) in all six viewing directions, under the protocol of Fig. 11 . Every baseline uses more bits than OIC-GS on both scenes ( 0.554 against 0.600 - 0.894 bpp on scene 74, and 0.436 against 0.561 - 0.792 bpp on scene 46). Four of the five panels in Figs. 11 – 13 where a GS codec exceeds OIC-GS in viewport PSNR are shown here, all for GaussianImage++ at a 30 - 82% higher bitrate and by 0.1 - 0.6 dB: the yaw- 180∘ and nadir views of scene 74 and the zenith and nadir of scene 46. The nadir patch of scene 74 is the only panel where no textured window favors OIC-GS over JPEG; at the zenith of scene 55, the best window ties.
Figure 13: Additional qualitative results (part 3 of 3): scenes 7 (top) and 32 (bottom) in all six viewing directions, under the protocol of Fig. 11 . Every baseline uses more bits than OIC-GS on both scenes ( 0.527 against 0.597 - 0.783 bpp on scene 7, and 0.581 against 0.600 - 0.754 bpp on scene 32), and OIC-GS achieves the highest viewport PSNR among the GS codecs and COIN in all twelve panels.