Organizations: Pengcheng Laboratory, Shenzhen, China · Sun Yat-sen University, Shenzhen, China · Peking University, Shenzhen, China · Xidian University, Xi’an, China
Transform-based methods provide an effective framework for point cloud attribute compression by representing attributes as transform coefficients. Introducing learned spatial context into this framework requires mapping spatial representations to the transform domain, but this known basis change is often left for the network to learn implicitly. We propose Transform-Aligned Learned Features (TALF) by applying the attribute transform to learned spatial representations, explicitly aligning them with the coding targets. Our analysis shows that the resulting features exactly represent the first-order prediction term of a smooth nonlinear model, with a bounded Taylor remainder. We integrate TALF into a transform-based attribute codec with explicit coefficient prediction and conditional residual entropy modeling under a unified coefficient-domain rate--distortion objective, while retaining explicit quantization-step control. Extensive experiments across three benchmark datasets and multiple transform bases demonstrate that TALF improves rate--distortion performance over conventional and learned baselines.
Figures & tables
Dataset
3DAC
Unicorn
TALF
Ford
+1.81
−5.63†
−20.93
KITTI
+9.89
−4.68†
−14.01
Table 1: Comparisons of coding performance using average BD-Rate (%) against the baseline methods and runtime (s/frame). G-PCCv33 is the BD-Rate anchor, except for Unicorn † , which uses the reported Unicorn-G-PCC (RAHT21) anchor.
Predictor
Nonlinearity
Parameters
Innovation MSE
Residual energy reduction
Shared linear predictor
None
513
139.2518
30.37
Shared one-hidden-layer MLP
GELU
263,169
135.2843
32.35
Table 2: Analysis of transform alignment on Ford. (a) Residual energy reduction is relative to zero-innovation prediction. (b) Prediction MSE is evaluated on valid AC coefficients before quantization. (c) BD-Rate is relative to G-PCCv33. All reductions, increases, and BD-Rate values are in %.
Prediction variant
BD-Rate vs. G-PCCv33
BD-Rate vs. No μ
BD-Rate vs. Integer μnq
No μ
-7.68
0.00
–
Integer μnq
-15.50
-8.46
0.00
Full μ
-20.93
-14.02
-6.00
Table 3: Ablation of explicit coefficient correction on Ford. BD-Rate (%) is reported relative to G-PCCv33, No μ , and Integer μnq .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Reflectance
Color
Maximum references per node
32
32
Attention blocks
5
5
Attention width
128
128
Output feature dimension
512
512
Input feature dimension
5
9
Shared MLP dimensions
512→256
512→256
Appendix
Table 4: Network and training configuration used for the reported models.
QS
Pr(μnq=0)
E[∣μnq∣]
Pr(qμ=qn)
4
50.92
1.460
19.21
8
34.93
0.678
14.93
16
19.76
0.291
11.14
32
8.00
0.101
6.08
64
2.28
0.027
2.48
Appendix
Table 5: Mechanistic analysis of the explicit coefficient correction on Ford. The integer-only and full corrections are compared with the no-correction path using the codec’s actual quantizer. Probabilities, zero-symbol ratios, and relative changes are reported in %.
G-PCCv33
TALF
Point
bpp
PSNR-Refl
QS
bpp
PSNR-Refl
R8
4.4898
50.17
4.00
2.9921
46.28
R7
3.5089
46.24
5.03
2.6260
43.84
R6
2.5609
41.00
6.34
2.3246
42.03
R5
1.6631
35.45
8.00
2.0329
40.16
R4
0.8915
30.23
10.06
1.8060
38.83
Appendix
Table 6: Dataset-average rate–distortion results for Ford.
G-PCCv33
TALF
Point
bpp
PSNR-Refl
QS
bpp
PSNR-Refl
R8
3.6923
50.43
4.00
2.3908
46.44
R7
2.7439
46.44
5.03
2.0212
43.91
R6
1.8133
41.06
6.34
1.7255
42.07
R5
0.8855
35.56
8.00
1.4498
40.30
R4
0.2518
31.28
10.06
1.2438
38.97
Appendix
Table 7: Dataset-average rate–distortion results for KITTI.
G-PCCv33
TALF
Point
bpp
Y
Cb
Cr
YCbCr
QS
Neural bpp
RL bpp
Y
Cb
Cr
YCbCr
R8
7.6590
49.98
49.98
49.73
49.95
4.00
4.2864
4.6183
46.20
46.95
47.27
46.41
R7
4.8095
45.92
46.39
46.62
46.06
5.03
3.7922
4.0720
43.86
46.25
46.74
44.39
R6
2.5896
40.63
42.42
43.99
41.14
6.34
2.9026
3.1165
42.16
44.46
45.32
42.69
R5
1.3095
35.49
39.82
41.97
36.33
8.00
2.3043
2.4613
40.51
43.09
44.35
41.11
R4
0.5993
30.82
38.06
40.03
31.85
10.06
1.9140
2.0176
39.28
42.03
43.65
39.92
Appendix
Table 8: Dataset-average rate–distortion results for ScanNet.
Scalable compression is essential for bandwidth-adaptive transmission, yet most learned codecs are optimized for a fixed rate-distortion point, making rate adaptation costly due to re-encoding or maintaining multiple bitstreams. In this work, we propose TAFA-GSGC, a scalable learned point cloud geometry codec that enables multi-quality decoding from a single bitstream and a single trained model. TAFA-GSGC combines layered residual refinement with channel-group entropy coding, and introduces a Target-Aligned Feature Aggregation module to reduce cross-layer redundancy in enhancement residuals. Our framework supports up to 9 decodable quality levels with monotonic quality improvement as more subbitstreams are received, while maintaining strong compression efficiency. Compared with the PCGCv2 baseline, TAFA-GSGC demonstrates improved RD performance, achieving average BD-rate reductions of 4.99% and 5.92% in terms of D1-PSNR and D2-PSNR, respectively.
Xiumei Li, Alexander Kopte, André Kaup
Chair of Multimedia Communications and Signal Processing Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU) Erlangen, Germany
Given encoded 3D point cloud geometry available at the decoder, we study the problem of lossy attribute compression in a multi-resolution B-spline projection framework. A target continuous 3D attribute function is first projected onto a sequence of nested subspaces Fl0(p)⊆⋯⊆FL(p), where Fl(p) is a family of functions spanned by a B-spline basis function of order p at a chosen scale and its integer shifts. The projected low-pass coefficients Fl∗ are computed by variable-complexity unrolling of a rate-distortion (RD) optimization algorithm into a feed-forward network, where the rate term is the sparsity-promoting ℓ1-norm. Thus, the projection operation is end-to-end differentiable. For a chosen coarse-to-fine predictor, the coefficients are then adjusted to account for the prediction from a lower-resolution to a higher-resolution, which is also optimized in a data-driven manner.
Tam Thuc Do, Philip A. Chou, Gene Cheung
department of EECS, York University, 4700 Keele Street, Toronto, M3J 1P3, Canada
Plenoptic point clouds (PPC) are novel data structures that represent the light from different viewing directions in order to provide a higher degree of realism to regular point clouds. This is achieved by associating each point to multiple colors instead of a single one. Here, we present a method to efficiently compress the attributes of a PPC, consisting of a Karhunen-Loève transform over the color attributes followed by multiple attribute coders with intra prediction capability. This compression scheme can be incorporated within the MPEG's geometry-based PCC (G-PCC) standard, using any of G-PCC's existing solutions for attribute coding. Compression performance assessment using PPCs of different spatial resolutions reveals competitive results in comparison to existing methods, such as RAHT-based or video-based PCC solutions. We believe our coder to be the new state of the art.
Davi R. Freitas, Gustavo L. Sandri, Ricardo L. de Queiroz
Inria - Rennes Bretagne-Atlantique Rennes, France · Instituto Federal de Bras´ılia Bras´ılia, Brazil · Universidade de Bras´ılia Bras´ılia, Brazil