Prior artisan mesh generation works largely predict face tokens autoregressively, which makes inference slow. Recent methods instead flow match continuous latents built by Variational AutoEncoders (VAEs), but reconstruction quality drops significantly when geometry and topology are jointly encoded, and further when the latent space is compressed. We present MeshCarve, a flow matching method that generates entirely in compact latent spaces, generating vertex positions and edge connections separately and sidestepping the difficulty of a joint compact latent. To shorten the token sequence, we propose a hierarchical sparse transformer backbone, instantiated as VertexVAE and EdgeVAE. Instead of encoding fields over the surface voxels, both VAEs anchor on discrete vertices in their latent spaces, which drastically reduces the token sequence length, and our spatial-aware compression shortens it further without costing reconstruction. VertexVAE directly encodes vertex occupancy. For connectivity, we propose vertex-link encoding, which turns arbitrary connectivity between vertices into fixed-length continuous per-vertex embeddings and recovers complex artistic topology faithfully. MeshCarve combines these VAEs with an anchor generator and flow matches on the shortened token sequences. It shows advantages over state-of-the-art autoregressive and flow matching methods on Objaverse and generalizes to Toys4K. To the best of our knowledge, it is among the first artisan mesh generation methods whose every generative stage runs in a spatially compressed latent, with a token sequence only a fraction of the most compressed previous autoregressive and flow matching works.
Figures & tables
Autoregressive
Diffusion / Flow Matching
Method
Tokens / face
Method
Token
Target
Tokens / face
MeshAnything ( Chen et al., 2025b )
9
LATO ( Zhao et al., 2026 )
coarse surface voxel
joint
6.5
MeshGPT ( Siddiqui et al., 2024 )
6
LATO.2 ( Long et al., 2026 )
coarse surface voxel / vertex
G+T
0.88 / 0.32
EdgeRunner ( Tang et al., 2025 )
4.2
MeshFlow ( Li et al., 2026b )
vertex
joint
0.52
MeshAnything V2 ( Chen et al., 2025c )
4.1
Nexus ( Wang et al., 2026b )
vertex
G+T
0.79 / 0.52
DeepMesh ( Zhao et al., 2025 )
2.5
MeshCraft ( He et al., 2025 )
face
joint
1.0
Table 1: Tokens per triangle face, lower is a shorter sequence for the generative transformer. Target: G = geometry, T = topology, joint = one latent for both. Two-stage methods list both stages, and the two stages of MeshCarve share the same anchor and hence one value.
Figure 2: Overview of MeshCarve. Top: VertexVAE and EdgeVAE share one sparse transformer backbone, reaching the compressed 643 latent through the spatial down projection of Eq. 1 and decoding the vertex voxels and their connections through the mirrored up projection. Bottom: end-to-end generation from a dense point cloud through the anchor generator and two flows.
Objaverse
Toys4K
Method
CD ℓ2↓
CD ℓ1↓
HD ↓
NC ↑
CD ℓ2↓
CD ℓ1↓
HD ↓
NC ↑
w/ GT vertices
LATO.2 T-Flow ( Long et al., 2026 )
.00595
.00824
.0709
.8775
.00483
.00689
.0490
.9245
MeshCarve edge flow (ours)
.00462
.00644
.0495
.9218
.00374
.00538
.0461
.9488
End-to-end generation
MeshAnything ( Chen et al., 2025b )
.02838
.03655
.1661
.7537
.03525
.04798
.1765
.7754
Table 2: Shape-conditioned artisan mesh generation on Objaverse and Toys4K. The top block gives the edge stage the reference vertices. Best in bold within each block.
Figure 3: Qualitative comparison on Objaverse and Toys4K. References are in clay, baselines in blue and MeshCarve in amber; all outputs are shown as generated with wireframes. Zoom in for better view.
Table 5
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
vertices / ref.
faces / ref.
faces / vertex
Reference
1.00
1.00
1.48
MeshAnything ( Chen et al., 2025b )
0.24
0.33
1.74
MeshAnythingV2 ( Chen et al., 2025c )
0.74
1.01
1.89
TreeMeshGPT ( Lionar et al., 2025 )
0.58
0.78
1.88
BPT ( Weng et al., 2025 )
0.43
0.61
1.96
DeepMesh ( Zhao et al., 2025 )
1.03
1.47
1.90
Appendix
Table 5: Vertex and face budgets relative to the reference on Toys4K, means per mesh. Bold: closest to the reference among the flow-based methods. Underlined: among the autoregressive methods.
Objaverse
Toys4K
Method
B
M
NM
Val.
W1
B
M
NM
Val.
W1
Reference
5.6
94.4
0.0
5.52
0.00
23.0
76.8
0.2
4.84
0.00
MeshAnything ( Chen et al., 2025b )
11.3
86.3
2.4
5.04
0.87
10.6
85.7
3.7
5.33
1.20
MeshAnythingV2 ( Chen et al., 2025c )
10.5
83.0
6.5
5.40
0.89
12.3
81.8
5.9
5.71
1.37
TreeMeshGPT ( Lionar et al., 2025 )
9.2
89.7
1.1
5.42
0.61
4.4
95.1
0.6
5.74
1.17
BPT ( Weng et al., 2025 )
9.6
84.2
6.2
5.72
0.87
5.3
92.1
2.6
5.86
1.28
Appendix
Table 6: Structural statistics of the generated meshes, means per mesh. B / M / NM: boundary, manifold and non-manifold edge shares in %. Val.: mean vertex valence. W1 : distance of the valence histogram to the reference’s (lower is closer).
Figure 4: Latent sequence length of each VAE against mesh size on 5,000 random training meshes, averaged within 1K-wide bins of the vertex-voxel count at 5123 . Every length is determined by the ground-truth mesh alone, so the figure isolates compression from generation. Surface voxels are those hit by a dense sample of the surface.
n
Objaverse
Toys4K
Method
Obj.
Toys
CD ℓ2↓
CD ℓ1↓
HD ↓
NC ↑
CD ℓ2↓
CD ℓ1↓
HD ↓
NC ↑
MeshAnything, cap 800
within the cap
15
131
.02838
.03655
.1661
.7537
.03525
.04798
.1765
.7754
MeshCarve, same meshes
15
131
.00615
.00872
.0693
.8950
.00563
.00809
.0649
.9296
full set
185
500
.03088
.04105
.1854
.7083
.03624
.05020
.1766
.7539
MeshAnything V2, cap 1,600
Appendix
Table 7: Autoregressive baselines inside and outside their published face caps, with MeshCarve on the same meshes inside each cap and on the full sets. n is the number of meshes scored per set.
Figure 5: Additional results of MeshCarve on Objaverse and Toys4K. In each pair the reference artisan mesh is on the left in ivory and the generated mesh on the right in amber.
Figure 6: Draws of the anchor generator. Each object shows the reference mesh (GT) and three draws of its 643 latent voxel anchor at the reference vertex budget. Blue voxels differ from the other two draws.
Table 8: Generation time of MeshCarve per mesh in seconds, means over 24 Objaverse meshes on one RTX PRO 6000 Blackwell Max-Q, in total and per bin of reference vertex voxels at 5123 .
Anchor
CD ℓ2↓
CD ℓ1↓
HD ↓
NC ↑
Posterior mean
.00649
.00933
.0789
.8529
Posterior sample
.00674
.00964
.0803
.8516
Appendix
Table 9: End-to-end results on 185 Objaverse meshes.