High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a Volumetric Anchor Lattice coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by 8.7%, coverage by 5.96 absolute points, and Betti error by 9.2% over the strongest baseline, while using 70.0% fewer tokens than the next-most compact baseline and over 98% fewer tokens than sparse or hierarchical tokenizers, effectively reducing training memory by 40.4% and inference time by 58.5%. Qualitative results further show improved preservation of thin structures, repeated components, and long-range connectivity.
Figures & tables
Figure 1: High-resolution 3D generation through topology-preserving slice latents. Given a single input image, 0.64314 0.68627 0.84706S0.34902 0.7451 0.88235I __color_backend_reset: __color_backend_reset: 0.41569 0.76863 0.6902L0.47843 0.78824 0.50196S0.5451 0.81176 0.3098A __color_backend_reset: __color_backend_reset: __color_backend_reset: generates high-resolution 3D shapes with coherent structure and detailed local geometry, including thin structures, holes, repeated components, and long-range connectivity. By modeling shapes through compact overlapping slice latents and cross-axis volumetric coordination, SILSA maintains cross-sectional continuity and produces coherent global structure.
Figure 2: 0.64314 0.68627 0.84706S0.34902 0.7451 0.88235I __color_backend_reset: __color_backend_reset: 0.41569 0.76863 0.6902L0.47843 0.78824 0.50196S0.5451 0.81176 0.3098A __color_backend_reset: __color_backend_reset: __color_backend_reset: Overview. SILSA represents each shape with overlapping slice latents along the x , y , and z axes. A topology-aware SliceVAE encodes surfaces into fixed multi-axis slice latents and decodes them through sparse volumetric upsampling into a high-resolution mesh. An image-conditioned rectified-flow transformer generates these latents, using a Volumetric Anchor Lattice as shared 3D memory for cross-axis coordination. Slice-wise persistent-homology and Betti-transition losses supervise connected components, holes, and topological consistency during VAE training.
Model
CLIP ↑
FD ↓
KD ↓
PSNR ↑
LPIPS ↓
COV(%) ↑
MMD(‰) ↓
Shap-E
80.16
34.64
0.87
16.84
0.21
61.41
19.19
LN3Diff
82.79
26.98
0.76
18.73
0.19
55.21
19.84
Direct3D
74.12
24.97
0.33
22.36
0.17
58.72
18.46
3DTopia-XL
76.46
24.21
0.29
22.06
0.18
58.93
17.62
InstantMesh
84.41
20.13
0.29
25.72
0.11
66.84
16.72
Gau.Any.
80.91
22.46
0.44
23.84
0.15
60.01
15.47
Table 1: Image-to-3D generation. We compare SILSA with representative image-conditioned and native 3D generation methods. Best and second best highlighted.
Table 2: VAE reconstruction quality. We compare SILSA against representative 3D reconstruction tokenizers using geometric, volumetric, and topology-aware metrics. Betti-Err measures the average mismatch in connected components and holes. Best and second best results highlighted.
Table 3: Efficiency comparison. Token counts for variable-length methods are reported as mean ± std over the test set. Memory and training time are measured with batch size 4 on a single A100. Inference time is reported end-to-end per shape. Best and second best highlighted.
Figure 3: Image-to-3D generation in the wild.
Figure 4: VAE reconstruction quality. We compare reconstructed meshes from SILSA and representative 3D models. Normal maps are shown in the top-right inset, and surface-error maps are shown in the bottom-right inset. Surface error is visualized as from low (blue) to high (red).
Variant
CD ↓
IoU ↑
Betti ↓
Slice-wise topology loss
No Ltopo
0.78
81.76
4.43
Only LPH
0.72
84.41
2.91
Only Ltrans
0.69
87.16
2.87
Full Ltopo
0.59
93.01
1.58
Cross-axis communication
Table 4: Key ablations. We ablate topology supervision and VAL. Best and second best highlighted.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Topology signals for slice-wise supervision. For each canonical slicing direction, cross-sections form a sequence over depth. Blue intervals denote connected components (β0) and red intervals denote holes (β1) that persist across ranges of slices. Our loss uses these signals in two ways: per-slice persistence matching supervises the topology within each cross-section, while Betti-transition matching supervises where components and holes appear, disappear, merge, or split across neighboring slices.
Model
CD ↓
F-Score@0.01 ↑
F-Score@0.005 ↑
IoU ↑
Betti-Err ↓
Trellis (SLAT)
0.07
99.71
91.18
96.84
0.01
SparseFlex
0.05
100.00
93.07
98.42
0.01
SILSA
0.04
100.00
93.21
98.71
0.00
Appendix
Table 5: VAE reconstruction on open-surface shapes. Best and second best highlighted.
Figure 6: Additional image-to-3D results.
Variant
#Tokens
CD ↓
F@0.01 ↑
IoU ↑
Betti-Err ↓
Number of slices N
N=64
192
0.82
94.36
88.24
2.86
N=96
288
0.67
95.81
91.27
2.04
N=128 (default)
384
0.59
96.79
93.01
1.58
N=256
768
0.57
97.02
93.28
1.53
N=512
1536
0.55
97.24
93.51
1.63
Appendix
Table 6: Ablation on slice representation. We ablate the number of slices per axis N and the sliding-window width w . Default settings are N=128 , w=8 , and three canonical axes, yielding 3N=384 slice tokens. Reported on the VAE reconstruction task.
Variant
CD ↓
F@0.01 ↑
F@0.005 ↑
IoU ↑
Betti-Err ↓
Loss components
No Ltopo
0.78
93.84
78.16
81.76
4.43
Only LPH
0.72
94.51
79.72
84.41
2.91
Only Ltrans
0.69
95.03
80.46
87.16
2.87
Full Ltopo (default)
0.59
96.79
84.03
93.01
1.58
Topology loss weight λtopo
Appendix
Table 7: Ablation of the slice-wise topology-preserving loss. All variants are evaluated on VAE reconstruction. Default settings are shaded.
Variant
Axes
N
#Tokens
CD ↓
IoU ↑
Betti-Err ↓
Single axis
1 axis ( z only)
z
384
384
1.34
76.43
5.92
1 axis ( z only)
z
128
128
1.87
71.29
7.83
1 axis ( x only)
x
384
384
1.41
75.18
6.18
1 axis ( x only)
x
128
128
1.94
70.42
8.07
Two axes
Appendix
Table 8: Ablation on the number of canonical axes. Default setting is shaded.
Variant
CD ↓
F@0.01 ↑
IoU ↑
Betti-Err ↓
Overwrite (no gating)
0.66
95.61
90.86
1.91
Additive write (no gating)
0.63
96.12
92.04
1.74
Gated write (default)
0.59
96.79
93.01
1.58
Appendix
Table 9: Ablation on the VAL update mechanism. Default setting is shaded.