High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a Volumetric Anchor Lattice coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by 8.7%, coverage by 5.96 absolute points, and Betti error by 9.2% over the strongest baseline, while using 70.0% fewer tokens than the next-most compact baseline and over 98% fewer tokens than sparse or hierarchical tokenizers, effectively reducing training memory by 40.4% and inference time by 58.5%. Qualitative results further show improved preservation of thin structures, repeated components, and long-range connectivity.
Figures & tables
Figure 1: High-resolution 3D generation through topology-preserving slice latents. Given a single input image, 0.64314 0.68627 0.84706S0.34902 0.7451 0.88235I __color_backend_reset: __color_backend_reset: 0.41569 0.76863 0.6902L0.47843 0.78824 0.50196S0.5451 0.81176 0.3098A __color_backend_reset: __color_backend_reset: __color_backend_reset: generates high-resolution 3D shapes with coherent structure and detailed local geometry, including thin structures, holes, repeated components, and long-range connectivity. By modeling shapes through compact overlapping slice latents and cross-axis volumetric coordination, SILSA maintains cross-sectional continuity and produces coherent global structure.
Figure 2: 0.64314 0.68627 0.84706S0.34902 0.7451 0.88235I __color_backend_reset: __color_backend_reset: 0.41569 0.76863 0.6902L0.47843 0.78824 0.50196S0.5451 0.81176 0.3098A __color_backend_reset: __color_backend_reset: __color_backend_reset: Overview. SILSA represents each shape with overlapping slice latents along the x , y , and z axes. A topology-aware SliceVAE encodes surfaces into fixed multi-axis slice latents and decodes them through sparse volumetric upsampling into a high-resolution mesh. An image-conditioned rectified-flow transformer generates these latents, using a Volumetric Anchor Lattice as shared 3D memory for cross-axis coordination. Slice-wise persistent-homology and Betti-transition losses supervise connected components, holes, and topological consistency during VAE training.
Model
CLIP ↑
FD ↓
KD ↓
PSNR ↑
LPIPS ↓
COV(%) ↑
MMD(‰) ↓
Shap-E
80.16
34.64
0.87
16.84
0.21
61.41
19.19
LN3Diff
82.79
26.98
0.76
18.73
0.19
55.21
19.84
Direct3D
74.12
24.97
0.33
22.36
0.17
58.72
18.46
3DTopia-XL
76.46
24.21
0.29
22.06
0.18
58.93
17.62
InstantMesh
84.41
20.13
0.29
25.72
0.11
66.84
16.72
Gau.Any.
80.91
22.46
0.44
23.84
0.15
60.01
15.47
Table 1: Image-to-3D generation. We compare SILSA with representative image-conditioned and native 3D generation methods. Best and second best highlighted.
Table 2: VAE reconstruction quality. We compare SILSA against representative 3D reconstruction tokenizers using geometric, volumetric, and topology-aware metrics. Betti-Err measures the average mismatch in connected components and holes. Best and second best results highlighted.
Table 3: Efficiency comparison. Token counts for variable-length methods are reported as mean ± std over the test set. Memory and training time are measured with batch size 4 on a single A100. Inference time is reported end-to-end per shape. Best and second best highlighted.
Figure 3: Image-to-3D generation in the wild.
Figure 4: VAE reconstruction quality. We compare reconstructed meshes from SILSA and representative 3D models. Normal maps are shown in the top-right inset, and surface-error maps are shown in the bottom-right inset. Surface error is visualized as from low (blue) to high (red).
Variant
CD ↓
IoU ↑
Betti ↓
Slice-wise topology loss
No Ltopo
0.78
81.76
4.43
Only LPH
0.72
84.41
2.91
Only Ltrans
0.69
87.16
2.87
Full Ltopo
0.59
93.01
1.58
Cross-axis communication
Table 4: Key ablations. We ablate topology supervision and VAL. Best and second best highlighted.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Topology signals for slice-wise supervision. For each canonical slicing direction, cross-sections form a sequence over depth. Blue intervals denote connected components (β0) and red intervals denote holes (β1) that persist across ranges of slices. Our loss uses these signals in two ways: per-slice persistence matching supervises the topology within each cross-section, while Betti-transition matching supervises where components and holes appear, disappear, merge, or split across neighboring slices.
Model
CD ↓
F-Score@0.01 ↑
F-Score@0.005 ↑
IoU ↑
Betti-Err ↓
Trellis (SLAT)
0.07
99.71
91.18
96.84
0.01
SparseFlex
0.05
100.00
93.07
98.42
0.01
SILSA
0.04
100.00
93.21
98.71
0.00
Appendix
Table 5: VAE reconstruction on open-surface shapes. Best and second best highlighted.
Figure 6: Additional image-to-3D results.
Variant
#Tokens
CD ↓
F@0.01 ↑
IoU ↑
Betti-Err ↓
Number of slices N
N=64
192
0.82
94.36
88.24
2.86
N=96
288
0.67
95.81
91.27
2.04
N=128 (default)
384
0.59
96.79
93.01
1.58
N=256
768
0.57
97.02
93.28
1.53
N=512
1536
0.55
97.24
93.51
1.63
Appendix
Table 6: Ablation on slice representation. We ablate the number of slices per axis N and the sliding-window width w . Default settings are N=128 , w=8 , and three canonical axes, yielding 3N=384 slice tokens. Reported on the VAE reconstruction task.
Variant
CD ↓
F@0.01 ↑
F@0.005 ↑
IoU ↑
Betti-Err ↓
Loss components
No Ltopo
0.78
93.84
78.16
81.76
4.43
Only LPH
0.72
94.51
79.72
84.41
2.91
Only Ltrans
0.69
95.03
80.46
87.16
2.87
Full Ltopo (default)
0.59
96.79
84.03
93.01
1.58
Topology loss weight λtopo
Appendix
Table 7: Ablation of the slice-wise topology-preserving loss. All variants are evaluated on VAE reconstruction. Default settings are shaded.
Variant
Axes
N
#Tokens
CD ↓
IoU ↑
Betti-Err ↓
Single axis
1 axis ( z only)
z
384
384
1.34
76.43
5.92
1 axis ( z only)
z
128
128
1.87
71.29
7.83
1 axis ( x only)
x
384
384
1.41
75.18
6.18
1 axis ( x only)
x
128
128
1.94
70.42
8.07
Two axes
Appendix
Table 8: Ablation on the number of canonical axes. Default setting is shaded.
Variant
CD ↓
F@0.01 ↑
IoU ↑
Betti-Err ↓
Overwrite (no gating)
0.66
95.61
90.86
1.91
Additive write (no gating)
0.63
96.12
92.04
1.74
Gated write (default)
0.59
96.79
93.01
1.58
Appendix
Table 9: Ablation on the VAL update mechanism. Default setting is shaded.
Scaling image-to-3D generation to ultra-high resolutions requires controlling rapidly growing computational costs without sacrificing fine geometric detail. We present \textbf{Filigree3D}, a sparse latent flow-matching framework that generates 3D geometry from a single image at voxel resolutions up to 20483, with straightforward extensibility to 40963. To make training tractable, we introduce Structure-Aware Sparse Scaling, which combines spatial bounding with alternating local-global attention to constrain token growth while preserving both fine-scale details and long-range structural context. To enhance detail reconstruction, we curate training samples based on their high-resolution geometric gains and inject multi-scale image features into a sparse 3D DiT, effectively coupling structural semantics with fine-grained visual cues. Furthermore, a visibility-aware voxel regularization strategy improves robustness against sparse perturbations and facilitates the completion of unobserved geometry. Under our default configuration, Filigree3D maintains peak GPU memory consumption within practical limits for contemporary hardware, enabling the generation of highly intricate 3D geometry in approximately one minute. Extensive experiments demonstrate that our method yields substantial improvements in overall geometric fidelity and fine-detail preservation compared to existing baselines, validating practical, detail-preserving 3D generation at unprecedented resolutions.
Autoregressive multimodal large language models (MLLMs) enable 3D generation but struggle to scale to high-resolution shapes due to inadequate 3D tokenizations. Compact set-based representations discard deterministic spatial ordering, leading to ambiguous sequence prediction, while uniform or octree-based voxel grids preserve ordering at the cost of severe redundancy and excessively long sequences. This structural trade-off limits stable and efficient autoregressive 3D generation. We present SuperVoxelGPT, a representation-first framework that resolves this tension through adaptive and deterministically ordered supervoxel tokenization. Given a prompt, we first predict a coarse geometric saliency distribution and construct a shape-adaptive supervoxel partition using saliency-guided centroidal Voronoi tessellation, allocating fine-grained cells to complex regions and larger cells to smooth regions. Conditioned on this prompt and ordered supervoxel layout, we introduce a SuperVoxelVAE and fine-tune a pretrained MLLM to autoregressively generate supervoxel tokens. Experiments using Trellis-500K data show that SuperVoxelGPT reduces token sequence length to 12.8% of uniform voxel tokenization while achieving state-of-the-art generation quality and an average 10x speedup over prior methods.
Yuan Li, Congyi Zhang, Xifeng Gao +1
University of Texas at Dallas, USA · LightSpeed, USA
Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize latent representations either over spatial regions or as fixed-size sets of global tokens, both suffering sharp reconstruction degradation when compressed to extremely low token budgets. In this paper, we present ZipTok3D, a 3D tokenizer designed for high-fidelity reconstruction from extremely short token sequences. Its key idea is to organize object geometry into progressively informative global-token prefixes and unfold these compact representations through iterative decoding. Specifically, nested dropout randomly truncates the latent sequence after encoding during training and requires each retained prefix to reconstruct the complete object, thereby prioritizing essential geometric information in the leading tokens. The decoder then repeatedly applies a parameter-shared Transformer block to recover fine-grained geometry from each prefix without a separate generative sampling stage. With the same token dimension, ZipTok3D achieves reconstruction quality comparable to the 32-token COD-VAE baseline using only one token on ShapeNet and four on TRELLIS, yielding 32× and 8× shorter token sequences, respectively.
Mingda Lin, Weijie Wang, Zeyu Zhang +7
1Zhejiang University · 2Monash University · University of Adelaide