Generating compact and geometrically faithful 3D meshes directly from point clouds remains a fundamental challenge. Point clouds are unordered and sparse, whereas meshes exhibit irregular structure and varying topology. As a result, many existing approaches rely on implicit representations followed by surface extraction or reconstruction. Although effective, these pipelines can produce dense or over-smoothed meshes, often requiring computationally expensive post-processing and simplification. We present OptimusMesh, a framework for direct compact triangle mesh generation from point clouds using sparse latent pivot conditioning. Our key idea is to compress 2,048 oriented input points into only 16 sparse latent pivots, reducing the geometric conditioning set by 128×. These pivots provide a compact structural representation shared across a two-stage autoregressive framework that first generates mesh vertices and then predicts triangular faces conditioned on the generated vertices and the same pivots. Compared with the evaluated recent point-cloud-conditioned autoregressive methods, which use 257 decoder-conditioning tokens, OptimusMesh uses only 16, yielding a 16.1× shorter conditioning sequence. Experiments show that OptimusMesh produces the most compact outputs among the compared recent autoregressive methods, using 25.7%--94.1% fewer faces while maintaining competitive geometric fidelity and distributional quality.
Figures & tables
Figure 2 : OptimusMesh pipeline. Given 2,048 oriented input points, the point-cloud encoder extracts K=16 sparse latent pivots with D=48 features, reducing the geometric conditioning sequence by 128× . The vertex model first generates V∈RN×3 , followed by a face model conditioned on the generated vertices and the same pivots to produce a compact triangle mesh with F≤800 .
Figure 3 : Sparse latent pivot encoder. The SLIDE-style point-cloud encoder progressively downsamples 2,048 oriented input points to K=16 sparse latent pivots with feature dimension D=48 .
Figure 4 : Qualitative comparison on ShapeNet. Representative point-cloud-conditioned mesh generation results from OptimusMesh and the evaluated baselines. Face counts (#F) are reported below each generated mesh. OptimusMesh produces compact meshes while preserving the overall structure of the input shapes.
Method
MMD ↓
COV ↑
1-NNA ↓
#V ↓
#F ↓
PSR
0.0849
0.4989
0.1203
5,980
11,563
SAP
0.0863
0.4855
0.1392
20,957
41,912
NKSR
0.0831
0.4922
0.1025
8,904
17,694
MeshAnything
0.0659
0.6347
0.1492
182
335
MeshAnythingV2
0.0689
0.6125
0.1192
463
858
FastMesh-V1K
0.0660
0.6659
0.1481
609
3,888
Table 1 : Distribution-level quality and mesh compactness on the ShapeNet test subset. Best results are shown in bold and second-best results are underlined .
Method
CD-L1 ↓
HD ↓
F1@2% ↑
NC ↑
MeshAnything
0.2123
0.3617
0.3881
0.5233
MeshAnythingV2
0.2059
0.3546
0.3985
0.5367
FastMesh-V1K
0.1878
0.3323
0.4151
0.5690
MeshRipple-10K
0.1994
0.3448
0.4245
0.5608
Ours
0.1421
0.2705
0.3570
0.5956
Table 2 : Paired geometric fidelity on 450 common ShapeNet test objects. Best results are shown in bold and second-best results are underlined .
Method
Input pts.
Cond. tokens ↓
Time (s) ↓
MeshAnything
4,096
257
35.61
MeshAnythingV2
8,192
257
63.40
FastMesh-V1K
8,192
257
8.58
MeshRipple-10K
16,384
257
379.50
Ours
2,048
16
15.75
Table 3 : Conditioning and inference efficiency of recent point-cloud-conditioned autoregressive mesh generators. Best results are shown in bold and second-best results are underlined .
Figure 5 : Qualitative results on Toys4K. Representative input point clouds and meshes generated by OptimusMesh across objects with diverse geometry and aspect ratios.
Condition
Tok. ↓
CD-L1 ↓
F1@2% ↑
NC ↑
#F ↓
Time ↓
32×48 pivots
32
0.1789
0.3012
0.5423
307.3
28.10
Dense 2048×48
2,048
0.1647
0.3026
0.5584
317.2
34.66
16×48 pivots
16
0.1417
0.3571
0.5741
249.5
15.81
Table 4 : Ablation of sparse and dense geometric conditioning.
Figure 6 : Additional meshes generated by OptimusMesh. Left: example output meshes generated by our method. Right: corresponding generated vertex sets before face prediction, illustrating how the predicted vertices are converted into final triangle meshes.
Point-cloud encoder
Mesh decoders
Dataset
Train
Val.
Test
Total
Train
Val.
Test
Total
ShapeNet
30,465
4,342
8,710
43,517
11,675
1,617
3,286
16,578
Objaverse
13,834
782
1,655
16,271
13,834
782
1,655
16,271
Toys4K
2,596
153
309
3,058
2,596
153
309
3,058
Total
46,895
5,277
10,674
62,846
28,105
2,552
5,250
35,907
Table 5 : Dataset composition for sparse-pivot encoder and mesh-decoder training.
Figure 7 : Visualization of sparse latent pivots. Representative objects are shown from multiple viewpoints. Orange markers denote the extracted pivot locations distributed across the input geometry.
Figure 8 : Representative failure cases of OptimusMesh. Examples include early termination, incomplete connectivity, and distorted local geometry.
Meshes are among the most common 3D scene representations, but directly generating meshes is challenging because the representation contains important symmetries, including permutation invariance of faces and vertices. MeshFlow learns to generate triangle meshes directly as triangle soups, avoiding the need to serialize meshes into long autoregressive sequences. We adopt equivariant optimal-transport flow matching models that respect the key symmetries of triangle soups: arbitrary permutations of faces and permutations of the vertices within each face. Toward this goal, we propose a simple yet effective modification to the Diffusion Transformer architecture, resulting in a scalable network capable of modeling a velocity field while maintaining the desired equivariance. We further introduce an optimal-transport-based training objective that improves convergence by eliminating supervision signals that violate these symmetries. MeshFlow achieves mesh quality comparable to state-of-the-art autoregressive mesh generators while providing about an 18× speedup during inference. Project page is at https://qiisun.github.io/MeshFlow/.
Qi Sun, Kiyohiro Nakayama, Jing Nathan Yan +6
City University of Hong Kong, Hong Kong · Stanford University, USA · Cornell Tech, USA +1
Autoregressive mesh generation has gained attention by tokenizing meshes into sequences and training models in a language-modeling fashion. However, existing approaches suffer from two fundamental limitations: (i) low tokenization efficiency, which yields long token sequences and prevents scaling to high-poly meshes, and (ii) absence of geometry-aware guidance, as generation is conditioned only on global shape embeddings rather than local surface cues. We introduce MeshWeaver, an autoregressive framework that treats mesh generation as a surface weaving process by directly predicting the next vertex instead of independent coordinates. At its core is a multi-level sparse-voxel encoder that injects geometric context into the generative process in three complementary ways: providing voxel features as vertex representations, guiding token prediction via cross-attention to voxel features, and serving as a structural scaffold that constrains generation around the input surface. Our hierarchical design enables coarse-to-fine vertex prediction in a single decoding step, while tightly coupling the generative model with 3D geometry. Extensive experiments demonstrate that MeshWeaver achieves a state-of-the-art compression ratio of 18%, can generate meshes with up to 16K faces, and significantly improves geometric fidelity over prior approaches.
Polygonal meshes are the standard surface representation of modern 3D pipelines, and generating high-quality meshes with artist-style topology is essential for film, gaming, and interactive 3D applications. Mainstream approaches serialize a mesh into a token sequence and decode it autoregressively, which is slow at inference and sensitive to error accumulation, making them impractical for interactive asset creation. We present Meshy T2, a fast native mesh generation framework built on flow matching. At its core is a vertex-set mesh VAE that encodes a mesh into one continuous latent token per vertex and decodes vertices, edge connectivity, and face winding order in a single pass, preserving high-precision geometry and artist-authored topology without vertex quantization or welding. Generation proceeds as a coarse-to-fine cascade of two flow-matching models: an image-conditioned voxel flow first sketches the overall shape as a coarse occupancy scaffold, and a mesh flow then populates the scaffold with per-vertex latent tokens, conditioned on the image, the scaffold, and a requested vertex budget. This design delivers three practical capabilities: interactive generation speed through parallel flow-based synthesis; effective face-count control through the requested vertex budget; and native support for multi-part assets, whose components emerge directly from the generated connectivity. In our experiments, Meshy T2 achieves state-of-the-art geometric fidelity and completes end-to-end image-to-mesh generation within a median of 6 seconds, over an order of magnitude faster than autoregressive baselines. Code and weights will be available at https://github.com/meshy-dev/meshy-t2.