We propose Squeeze3D, a novel framework that leverages implicit prior knowledge learnt by existing pre-trained encoders and decoders to compress 3D data at extremely high compression ratios. Our approach bridges the latent spaces between a pre-trained encoder and a pretrained decoder model through trainable mapping networks. Any 3D asset represented as a mesh, point cloud, or radiance field is first encoded by the pre-trained encoder and then transformed (i.e. compressed) into a highly compact latent code by a mapping network. This latent code can effectively be used as an extremely compressed representation of the mesh, point cloud, or radiance field. A mapping network transforms the compressed latent code into the latent space of a powerful generative model; the decoder of this generative model then recreates the original 3D asset (i.e. decompression). Squeeze3D is trained entirely on generated synthetic data and does not require any 3D datasets. The Squeeze3D architecture can be flexibly used with existing pre-trained 3D encoders and existing generative models. It can flexibly support different formats, including meshes, point clouds, and radiance fields. Our experiments demonstrate that Squeeze3D achieves compression ratios of up to 2187× for textured meshes, 58.5× for point clouds, and more than 650× for radiance fields while maintaining visual quality comparable to many existing methods. Squeeze3D only incurs a small compression and decompression latency since it does not involve training object-specific networks to compress an object.
Figures & tables
Figure 2 : Overview of our Method. Squeeze3D bridges arbitrary latent spaces between 3D encoders and generators through trainable mapping networks. During compression, a 3D geometry is encoded and then transformed into a compact representation via the forward mapping network. During decompression, the reverse mapping network converts this representation into the generator’s latent space, which is then used to reconstruct the original geometry.
Symb.
Description
Symb.
Description
G
A 3D geometry in some format
E
3D encoder model: E(G)↦zE∈RdE
G
3D decoder model: G(zG,c)↦G′
zE
Latent code from the encoder: zE∈RdE
zG
Latent code for the generator: zG∈RdG
zcomp
Compressed representation zcomp∈RdC
c
Conditioning information ( e.g. text prompt, image) for G
FθE
Forward mapping network: FθE(zE)↦zcomp
FθD
Reverse mapping network: FθD(zcomp)↦zG
dC
Dimensionality of compressed representation
dG
Dimensionality of generator latent space
dE
Dimensionality of encoder latent space
Table 1 : Notation. The notation we use to describe our method.
Figure 3 : Training Squeeze3D . We show an overview of (a) our process of creating synthetic data to train the mapping networks and (b) our process of training the mapping networks.
Figure 4 : Qualitative mesh compression results. We compare Squeeze3D to state-of-the-art methods. Our approach maintains visually important geometric details. Additional qualitative comparisons are provided in Appendix C.4 .
Figure 5 : Qualitative radiance field compression results. We show qualitative results comparing Squeeze3D to state-of-the-art methods. Our approach achieves a significantly higher compression ratio while maintaining visually important geometric details.
Figure 6 : Compression Ratio and Reconstruction Quality Trade-off. We show the trade-off between compression ratio and reconstruction quality (LPIPS) for Squeeze3D and Draco. Each Squeeze3D operating point varies dC and uses a separately trained forward/reverse mapping pair. Squeeze3D maintains high quality even at extreme compression ratios, whereas Draco’s quality degrades significantly as its compression ratio increases.
Figure 7 : Qualitative point-cloud compression. Reconstructions from Squeeze3D and established codecs; quantitative results are reported in Tables 3 and 10 .
Method
CR ( × ) ↑
PSNR ↑
MS-SSIM ↑
LPIPS ↓
Shap-E Native Baseline
1.61
25.80
0.9600
0.0410
Squeeze3D (Shap-E)
1607.50
25.92
0.9615
0.0398
Table 5 : Shap-E Comparison. Comparison between Squeeze3D (MeshAnything → Shap-E) and the native Shap-E encoder-decoder.
Method
CR ( × ) ↑
PCQM ↑
PointSSIM ↑
Chamfer ↓
LION Native Baseline
3.60
2.1000
0.5100
0.0210
Squeeze3D (LION)
58.50
1.8437
0.4484
0.0285
Table 6 : LION Comparison. Comparison between Squeeze3D (PointNet++ → LION) and the native LION encoder-decoder.
Decoder
Total size
Squeeze3D
Shap-E
905 MB
InstantMesh
1.51 GB
OpenLRM
1.04 GB
LION
270 MB
NeRF-MAE
306 MB
Table 8 : Storage size of each decoder used at decompression.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : Network Architectures.
Encoder/Generator Model
Latent Space Size
MeshAnything ( Chen et al., 2025c )
257×1024
InstantMesh TriPlane Features ( Xu et al., 2024 )
3×80×64×64
LION ( Zeng et al., 2022 )
8320
PointNet++ ( Qi et al., 2017 )
1024
NeRF-MAE ( Irshad et al., 2024 )
256×20×20×20
Appendix
Table 10 : Latent Space Sizes. Sizes of latent spaces for encoder and generator models we use.
Encoder → decoder
dC
Latent size (KB)
MeshAnything → OpenLRM
1024
4.00
MeshAnything → Shap-E
1024
4.00
MeshAnything → InstantMesh
770
3.01
PointNet++ → LION
512
2.00
PointNet++ → LION
1024
4.00
PointNet++ → LION
2048
8.00
Appendix
Table 11: Latent Sizes. The sizes of the compressed representations for our experiments.
Decoder
Mapping training (h)
Paired-data creation (h)
InstantMesh
12
20
LION
5
20
NeRF-MAE
15
3
Appendix
Table 12 : One-time Training and Data-creation Costs. These costs are incurred once for the listed encoder-generator setting; subsequent objects use feed-forward inference.
Hyperparameter
LRM
Shap-E
InstantMesh
NeRF-MAE
Training Precision
FP-32
FP-32
FP-32
FP-32
dC
1024
1024
770
24000
Dropout
0.35
0.35
0.35
0.2
Epochs
700
700
700
200
Batch Size
16
8
16
4
Optimizer
Muon
Adam
Muon
Muon
Appendix
Table 13 : Training Hyperparameters. We show the training hyperparameters for the mapping networks we train.
Mesh method
Bits/vertex ↓
Point-cloud method
Bits/point ↓
Draco ∗ ( Galligan et al., 2018 )
141.06
Draco ‡ ( Galligan et al., 2018 )
29.92
Draco † ( Galligan et al., 2018 )
145.61
Draco § ( Galligan et al., 2018 )
20.88
Draco ‡ ( Galligan et al., 2018 )
147.13
G-PCC ( Schwarz et al., 2018 )
12.60
Draco § ( Galligan et al., 2018 )
157.75
V-PCC ( Liu et al., 2019 )
9.32
Corto ( Lab, 2025 )
21.24
SparsePCGC ( Wang et al., 2023 )
13.89
Neural Subd. ( Liu et al., 2020a )
86.46
Squeeze3D (LION, dC=512 )
8.00
Appendix
Table 14: Standardized Rate Measures. Bits per input vertex for mesh compression and bits per input point for point-cloud compression.
Figure 9 : Interpolation. The compressed representations we obtain can also be interpolated. In these examples, we obtain the compressed representation for the leftmost and rightmost meshes and linearly interpolate between them.
Figure 10 : Point-cloud rate–distortion. PointSSIM (x-axis) versus bits/input point (log y-axis) for Squeeze3D and four codecs. Each Squeeze3D point uses separately trained mappings.
Figure 11 : Multi-view visualization of compressed and reconstructed meshes. The consistent appearance across different viewing angles demonstrates that Squeeze3D learns correct transformations between latent spaces and produces coherent 3D reconstructions. This confirms that our compressed representation encodes complete 3D information rather than view-dependent features.
Figure 12 : Squeeze3D preserves geometry details. We show some meshes compressed with Squeeze3D as wireframes. Notice that Squeeze3D preserves many fine-grained geometric details.
Figure 21
Size ( zcomp )
Compression (ms)
Decompression (ms)
PCQM ↑
PointSSIM ↑
512
3.85
12.74
1.4001
0.3640
1024
3.85
12.74
1.4047
0.3640
2048
4.22
13.39
1.3311
0.4249
4096
4.85
14.04
1.8437
0.4318
8192
5.30
14.19
1.5665
0.4473
Appendix
Table 18 : Ablations. Ablation study on compressed representation size. We analyze how varying the dimensionality of the compressed latent space affects compression time, decompression time, and reconstruction quality (measured by PCQM and PointSSIM).
Figure 15 : Failure Cases. We show examples where Squeeze3D fails to accurately reconstruct the input. (Left) Input mesh with highly intricate details or text. (Right) The reconstruction is smoothed, losing the fine-grained text or surface texture.
Case
PSNR ( ↑ )
MS-SSIM ( ↑ )
LPIPS ( ↓ )
Chamfer ( ↓ )
EMD ( ↓ )
XYZRGB Dragon
21.13
0.9315
0.0611
0.05154
0.12827
Lucy
19.47
0.8979
0.1043
0.19425
0.35322
Appendix
Table 20: Quantitative failure-case analysis. Image metrics are means over 36 geometry-only views; normalized Chamfer distance and approximate 3D EMD use 2,048 surface samples. Both meshes are independently normalized to a unit box without rigid ICP alignment.
σ / bbox diagonal
PSNR ( ↑ )
MS-SSIM ( ↑ )
LPIPS ( ↓ )
Chamfer ( ↓ )
EMD ( ↓ )
0 (clean)
27.50
0.9796
0.0274
0.3930
0.0084
0.005
26.49
0.9767
0.0329
0.4517
0.0174
0.010
24.19
0.9663
0.0533
0.6257
0.0204
0.025
20.12
0.9280
0.1369
1.1186
0.0305
0.050
19.07
0.9125
0.1916
1.9091
0.0954
0.100
18.54
0.9046
0.2140
2.0378
0.1107
Appendix
Table 22 : Input-noise sensitivity.
σ / input dynamic range
PSNR ( ↑ )
MS-SSIM ( ↑ )
LPIPS ( ↓ )
0 (clean)
26.62
0.9533
0.0743
0.005
26.55
0.9528
0.0752
0.010
26.43
0.9520
0.0768
0.025
26.10
0.9498
0.0805
0.050
25.62
0.9460
0.0869
0.100
24.90
0.9395
0.0980
Appendix
Table 23 : Radiance-field input-noise sensitivity. Gaussian noise scale is expressed relative to the input dynamic range.
Decoder
Shared storage
Original/object
Latent/object
Break-even objects
Shap-E
2.89 GB
6.43 MB
4.00 KB
449
InstantMesh
3.43 GB
6.43 MB
3.01 KB
534
OpenLRM
2.79 GB
6.43 MB
4.00 KB
435
LION ( dC=512 )
312.25 MB
117 KB
2.00 KB
2,716
NeRF-MAE
478.92 MB
58.07 MB
93.75 KB
9
Appendix
Table 24: Decompression-side amortization thresholds. Exact break-even calculations including both the frozen decoder and reverse mapping network.