Organizations: Pengcheng Laboratory, Shenzhen, China · Sun Yat-sen University, Shenzhen, China · Peking University, Shenzhen, China · Xidian University, Xi’an, China
Transform-based methods provide an effective framework for point cloud attribute compression by representing attributes as transform coefficients. Introducing learned spatial context into this framework requires mapping spatial representations to the transform domain, but this known basis change is often left for the network to learn implicitly. We propose Transform-Aligned Learned Features (TALF) by applying the attribute transform to learned spatial representations, explicitly aligning them with the coding targets. Our analysis shows that the resulting features exactly represent the first-order prediction term of a smooth nonlinear model, with a bounded Taylor remainder. We integrate TALF into a transform-based attribute codec with explicit coefficient prediction and conditional residual entropy modeling under a unified coefficient-domain rate--distortion objective, while retaining explicit quantization-step control. Extensive experiments across three benchmark datasets and multiple transform bases demonstrate that TALF improves rate--distortion performance over conventional and learned baselines.
Figures & tables
Dataset
3DAC
Unicorn
TALF
Ford
+1.81
−5.63†
−20.93
KITTI
+9.89
−4.68†
−14.01
Table 1: Comparisons of coding performance using average BD-Rate (%) against the baseline methods and runtime (s/frame). G-PCCv33 is the BD-Rate anchor, except for Unicorn † , which uses the reported Unicorn-G-PCC (RAHT21) anchor.
Predictor
Nonlinearity
Parameters
Innovation MSE
Residual energy reduction
Shared linear predictor
None
513
139.2518
30.37
Shared one-hidden-layer MLP
GELU
263,169
135.2843
32.35
Table 2: Analysis of transform alignment on Ford. (a) Residual energy reduction is relative to zero-innovation prediction. (b) Prediction MSE is evaluated on valid AC coefficients before quantization. (c) BD-Rate is relative to G-PCCv33. All reductions, increases, and BD-Rate values are in %.
Prediction variant
BD-Rate vs. G-PCCv33
BD-Rate vs. No μ
BD-Rate vs. Integer μnq
No μ
-7.68
0.00
–
Integer μnq
-15.50
-8.46
0.00
Full μ
-20.93
-14.02
-6.00
Table 3: Ablation of explicit coefficient correction on Ford. BD-Rate (%) is reported relative to G-PCCv33, No μ , and Integer μnq .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Reflectance
Color
Maximum references per node
32
32
Attention blocks
5
5
Attention width
128
128
Output feature dimension
512
512
Input feature dimension
5
9
Shared MLP dimensions
512→256
512→256
Appendix
Table 4: Network and training configuration used for the reported models.
QS
Pr(μnq=0)
E[∣μnq∣]
Pr(qμ=qn)
4
50.92
1.460
19.21
8
34.93
0.678
14.93
16
19.76
0.291
11.14
32
8.00
0.101
6.08
64
2.28
0.027
2.48
Appendix
Table 5: Mechanistic analysis of the explicit coefficient correction on Ford. The integer-only and full corrections are compared with the no-correction path using the codec’s actual quantizer. Probabilities, zero-symbol ratios, and relative changes are reported in %.
G-PCCv33
TALF
Point
bpp
PSNR-Refl
QS
bpp
PSNR-Refl
R8
4.4898
50.17
4.00
2.9921
46.28
R7
3.5089
46.24
5.03
2.6260
43.84
R6
2.5609
41.00
6.34
2.3246
42.03
R5
1.6631
35.45
8.00
2.0329
40.16
R4
0.8915
30.23
10.06
1.8060
38.83
Appendix
Table 6: Dataset-average rate–distortion results for Ford.
G-PCCv33
TALF
Point
bpp
PSNR-Refl
QS
bpp
PSNR-Refl
R8
3.6923
50.43
4.00
2.3908
46.44
R7
2.7439
46.44
5.03
2.0212
43.91
R6
1.8133
41.06
6.34
1.7255
42.07
R5
0.8855
35.56
8.00
1.4498
40.30
R4
0.2518
31.28
10.06
1.2438
38.97
Appendix
Table 7: Dataset-average rate–distortion results for KITTI.
G-PCCv33
TALF
Point
bpp
Y
Cb
Cr
YCbCr
QS
Neural bpp
RL bpp
Y
Cb
Cr
YCbCr
R8
7.6590
49.98
49.98
49.73
49.95
4.00
4.2864
4.6183
46.20
46.95
47.27
46.41
R7
4.8095
45.92
46.39
46.62
46.06
5.03
3.7922
4.0720
43.86
46.25
46.74
44.39
R6
2.5896
40.63
42.42
43.99
41.14
6.34
2.9026
3.1165
42.16
44.46
45.32
42.69
R5
1.3095
35.49
39.82
41.97
36.33
8.00
2.3043
2.4613
40.51
43.09
44.35
41.11
R4
0.5993
30.82
38.06
40.03
31.85
10.06
1.9140
2.0176
39.28
42.03
43.65
39.92
Appendix
Table 8: Dataset-average rate–distortion results for ScanNet.