Prior artisan mesh generation works largely predict face tokens autoregressively, which makes inference slow. Recent methods instead flow match continuous latents built by Variational AutoEncoders (VAEs), but reconstruction quality drops significantly when geometry and topology are jointly encoded, and further when the latent space is compressed. We present MeshCarve, a flow matching method that generates entirely in compact latent spaces, generating vertex positions and edge connections separately and sidestepping the difficulty of a joint compact latent. To shorten the token sequence, we propose a hierarchical sparse transformer backbone, instantiated as VertexVAE and EdgeVAE. Instead of encoding fields over the surface voxels, both VAEs anchor on discrete vertices in their latent spaces, which drastically reduces the token sequence length, and our spatial-aware compression shortens it further without costing reconstruction. VertexVAE directly encodes vertex occupancy. For connectivity, we propose vertex-link encoding, which turns arbitrary connectivity between vertices into fixed-length continuous per-vertex embeddings and recovers complex artistic topology faithfully. MeshCarve combines these VAEs with an anchor generator and flow matches on the shortened token sequences. It shows advantages over state-of-the-art autoregressive and flow matching methods on Objaverse and generalizes to Toys4K. To the best of our knowledge, it is among the first artisan mesh generation methods whose every generative stage runs in a spatially compressed latent, with a token sequence only a fraction of the most compressed previous autoregressive and flow matching works.
Figures & tables
Autoregressive
Diffusion / Flow Matching
Method
Tokens / face
Method
Token
Target
Tokens / face
MeshAnything ( Chen et al., 2025b )
9
LATO ( Zhao et al., 2026 )
coarse surface voxel
joint
6.5
MeshGPT ( Siddiqui et al., 2024 )
6
LATO.2 ( Long et al., 2026 )
coarse surface voxel / vertex
G+T
0.88 / 0.32
EdgeRunner ( Tang et al., 2025 )
4.2
MeshFlow ( Li et al., 2026b )
vertex
joint
0.52
MeshAnything V2 ( Chen et al., 2025c )
4.1
Nexus ( Wang et al., 2026b )
vertex
G+T
0.79 / 0.52
DeepMesh ( Zhao et al., 2025 )
2.5
MeshCraft ( He et al., 2025 )
face
joint
1.0
Table 1: Tokens per triangle face, lower is a shorter sequence for the generative transformer. Target: G = geometry, T = topology, joint = one latent for both. Two-stage methods list both stages, and the two stages of MeshCarve share the same anchor and hence one value.
Figure 2: Overview of MeshCarve. Top: VertexVAE and EdgeVAE share one sparse transformer backbone, reaching the compressed 643 latent through the spatial down projection of Eq. 1 and decoding the vertex voxels and their connections through the mirrored up projection. Bottom: end-to-end generation from a dense point cloud through the anchor generator and two flows.
Objaverse
Toys4K
Method
CD ℓ2↓
CD ℓ1↓
HD ↓
NC ↑
CD ℓ2↓
CD ℓ1↓
HD ↓
NC ↑
w/ GT vertices
LATO.2 T-Flow ( Long et al., 2026 )
.00595
.00824
.0709
.8775
.00483
.00689
.0490
.9245
MeshCarve edge flow (ours)
.00462
.00644
.0495
.9218
.00374
.00538
.0461
.9488
End-to-end generation
MeshAnything ( Chen et al., 2025b )
.02838
.03655
.1661
.7537
.03525
.04798
.1765
.7754
Table 2: Shape-conditioned artisan mesh generation on Objaverse and Toys4K. The top block gives the edge stage the reference vertices. Best in bold within each block.
Figure 3: Qualitative comparison on Objaverse and Toys4K. References are in clay, baselines in blue and MeshCarve in amber; all outputs are shown as generated with wireframes. Zoom in for better view.
Table 5
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
vertices / ref.
faces / ref.
faces / vertex
Reference
1.00
1.00
1.48
MeshAnything ( Chen et al., 2025b )
0.24
0.33
1.74
MeshAnythingV2 ( Chen et al., 2025c )
0.74
1.01
1.89
TreeMeshGPT ( Lionar et al., 2025 )
0.58
0.78
1.88
BPT ( Weng et al., 2025 )
0.43
0.61
1.96
DeepMesh ( Zhao et al., 2025 )
1.03
1.47
1.90
Appendix
Table 5: Vertex and face budgets relative to the reference on Toys4K, means per mesh. Bold: closest to the reference among the flow-based methods. Underlined: among the autoregressive methods.
Objaverse
Toys4K
Method
B
M
NM
Val.
W1
B
M
NM
Val.
W1
Reference
5.6
94.4
0.0
5.52
0.00
23.0
76.8
0.2
4.84
0.00
MeshAnything ( Chen et al., 2025b )
11.3
86.3
2.4
5.04
0.87
10.6
85.7
3.7
5.33
1.20
MeshAnythingV2 ( Chen et al., 2025c )
10.5
83.0
6.5
5.40
0.89
12.3
81.8
5.9
5.71
1.37
TreeMeshGPT ( Lionar et al., 2025 )
9.2
89.7
1.1
5.42
0.61
4.4
95.1
0.6
5.74
1.17
BPT ( Weng et al., 2025 )
9.6
84.2
6.2
5.72
0.87
5.3
92.1
2.6
5.86
1.28
Appendix
Table 6: Structural statistics of the generated meshes, means per mesh. B / M / NM: boundary, manifold and non-manifold edge shares in %. Val.: mean vertex valence. W1 : distance of the valence histogram to the reference’s (lower is closer).
Figure 4: Latent sequence length of each VAE against mesh size on 5,000 random training meshes, averaged within 1K-wide bins of the vertex-voxel count at 5123 . Every length is determined by the ground-truth mesh alone, so the figure isolates compression from generation. Surface voxels are those hit by a dense sample of the surface.
n
Objaverse
Toys4K
Method
Obj.
Toys
CD ℓ2↓
CD ℓ1↓
HD ↓
NC ↑
CD ℓ2↓
CD ℓ1↓
HD ↓
NC ↑
MeshAnything, cap 800
within the cap
15
131
.02838
.03655
.1661
.7537
.03525
.04798
.1765
.7754
MeshCarve, same meshes
15
131
.00615
.00872
.0693
.8950
.00563
.00809
.0649
.9296
full set
185
500
.03088
.04105
.1854
.7083
.03624
.05020
.1766
.7539
MeshAnything V2, cap 1,600
Appendix
Table 7: Autoregressive baselines inside and outside their published face caps, with MeshCarve on the same meshes inside each cap and on the full sets. n is the number of meshes scored per set.
Figure 5: Additional results of MeshCarve on Objaverse and Toys4K. In each pair the reference artisan mesh is on the left in ivory and the generated mesh on the right in amber.
Figure 6: Draws of the anchor generator. Each object shows the reference mesh (GT) and three draws of its 643 latent voxel anchor at the reference vertex budget. Blue voxels differ from the other two draws.
Table 8: Generation time of MeshCarve per mesh in seconds, means over 24 Objaverse meshes on one RTX PRO 6000 Blackwell Max-Q, in total and per bin of reference vertex voxels at 5123 .
Anchor
CD ℓ2↓
CD ℓ1↓
HD ↓
NC ↑
Posterior mean
.00649
.00933
.0789
.8529
Posterior sample
.00674
.00964
.0803
.8516
Appendix
Table 9: End-to-end results on 185 Objaverse meshes.
We present MeshFlow, a new method for generating artist-like 3D meshes. Current mesh generators often adopt Auto-Regressive (AR) next-token prediction, a natural choice given the discrete nature of mesh topology. However, AR methods scale poorly because the inference cost is quadratic in mesh size. They also require discretizing the vertex coordinates, which introduces quantization errors. To address these challenges, we introduce a Variational Autoencoder (VAE) that, supervised with a contrastive loss, represents both continuous vertex positions and discrete connectivity in a continuous latent space. This latent space is significantly more compact than prior token-based mesh representations. We then build a 3D generator based on a Rectified Flow transformer, generating all mesh vertices and edges in parallel. Our model generates meshes 18x faster than the fastest AR generator while also achieving excellent accuracy across standard mesh-generation metrics. Homepage: https://mesh-flow.github.io/, Code: https://github.com/facebookresearch/meshflow
Weiyu Li, Antoine Toisoul, Tom Monnier +4
Meta AI · Hong Kong University of Science and Technology
Polygonal meshes are the standard surface representation of modern 3D pipelines, and generating high-quality meshes with artist-style topology is essential for film, gaming, and interactive 3D applications. Mainstream approaches serialize a mesh into a token sequence and decode it autoregressively, which is slow at inference and sensitive to error accumulation, making them impractical for interactive asset creation. We present Meshy T2, a fast native mesh generation framework built on flow matching. At its core is a vertex-set mesh VAE that encodes a mesh into one continuous latent token per vertex and decodes vertices, edge connectivity, and face winding order in a single pass, preserving high-precision geometry and artist-authored topology without vertex quantization or welding. Generation proceeds as a coarse-to-fine cascade of two flow-matching models: an image-conditioned voxel flow first sketches the overall shape as a coarse occupancy scaffold, and a mesh flow then populates the scaffold with per-vertex latent tokens, conditioned on the image, the scaffold, and a requested vertex budget. This design delivers three practical capabilities: interactive generation speed through parallel flow-based synthesis; effective face-count control through the requested vertex budget; and native support for multi-part assets, whose components emerge directly from the generated connectivity. In our experiments, Meshy T2 achieves state-of-the-art geometric fidelity and completes end-to-end image-to-mesh generation within a median of 6 seconds, over an order of magnitude faster than autoregressive baselines. Code and weights will be available at https://github.com/meshy-dev/meshy-t2.
Autoregressive mesh generation has gained attention by tokenizing meshes into sequences and training models in a language-modeling fashion. However, existing approaches suffer from two fundamental limitations: (i) low tokenization efficiency, which yields long token sequences and prevents scaling to high-poly meshes, and (ii) absence of geometry-aware guidance, as generation is conditioned only on global shape embeddings rather than local surface cues. We introduce MeshWeaver, an autoregressive framework that treats mesh generation as a surface weaving process by directly predicting the next vertex instead of independent coordinates. At its core is a multi-level sparse-voxel encoder that injects geometric context into the generative process in three complementary ways: providing voxel features as vertex representations, guiding token prediction via cross-attention to voxel features, and serving as a structural scaffold that constrains generation around the input surface. Our hierarchical design enables coarse-to-fine vertex prediction in a single decoding step, while tightly coupling the generative model with 3D geometry. Extensive experiments demonstrate that MeshWeaver achieves a state-of-the-art compression ratio of 18%, can generate meshes with up to 16K faces, and significantly improves geometric fidelity over prior approaches.