3D content generation technology has significantly advanced the work of designers, as well as the 3D printing and gaming industries. However, it remains difficult to produce lightweight, editable, and topologically clean artistic content that is directly production-ready. To achieve this, we present TaoFlowForge, an artistic mesh foundation model that generates production-ready meshes. Specifically, TaoFlowForge decomposes the mesh generation process into vertices generation and their connectivity prediction, i.e., edges. We formulate vertices generation as a two-stage coarse-to-fine process and incorporate several effective loss functions to further enhance its performance. In the connectivity prediction stage, we propose a simple yet effective method for estimating the connectivity affinity between vertices and additionally predict per-vertex normals, which determines the correct orientation of faces. Besides, we construct a large-scale dataset combining hand-crafted 3D assets with public high-quality topology datasets. Based on this, a carefully designed data curation pipeline is employed to filter the raw dataset, retaining only high-quality topology data for model training. Our model is trained on the combined dataset and tested on both out-of-distribution hand-crafted set of 3D assets and public datasets. Under image-conditioned generation, TaoFlowForge outperforms autoregressive methods and achieves state-of-the-art results among open-source mesh topology generators. We will release all the code and weights together with a portion of our test dataset.
Figures & tables
Figure 2 : Overview of our proposed TaoFlowForge: Stage 1: Coarse-quantized vertices are generated at 643 resolution. Stage 2: Vertex positions are progressively refined from 643 to 5123 . Stage 3: Based on the generated vertices, we predict the final vertex connectivity to form mesh faces. Here, E and D denote the encoder and decoder, respectively, and F denotes the flow matching model in each stage, conditioned on image features cimg .
Figure 3 : Topology-aware co-occurrence priors of stage 2. Two soft-OR co-occupancy penalties complement per-cell cross-entropy with mesh connectivity constraints. Edge and face priors activate only when all connected vertices are jointly predicted empty, reinforcing structurally important high-degree vertices through accumulated constraints.
Figure 4 : Qualitative comparison with open-source polygon-based 3D mesh generation models. Given an image as input, our model generates the finest and most complete 3D meshes.
Figure 5 : Qualitative comparison with commercial polygon-based 3D mesh generation models. Given an image, our model generates more detailed and stable meshes with better alignment.
Method
CD ↓
HD ↓
ULIP-I ↑
Uni3D-I ↑
FD-Incep. ↓
FD-DINOv2 ↓
Toys4K
LATO.2 [ 20 ]
0.1553
0.3663
0.1674
0.3330
78.9649
616.4222
EdgeRunner [ 30 ]
0.3264
0.6145
0.1514
0.2347
30.9466
339.7754
TaoFlowForge (Ours)
0.0655
0.2451
0.1923
0.3509
18.1812
197.7314
TE-388
LATO.2 [ 20 ]
0.0991
0.2990
0.1509
0.2978
119.9100
931.6500
Table 1: Quantitative comparison on Toys4K and TE-388 under single-image conditioning.
Figure 6 : Qualitative ablation of late-stage progressive refinement, comparing continuous vertex generation with the late-stage refinement on structurally challenging examples.
Figure 7 : Additional image conditioned results. Our method generates meshes with both strong global structural stability and high-fidelity local details. Best viewed with zoom-in.
Stage 2 variant
CD ↓
HD ↓
ULIP-I ↑
Uni3D-I ↑
FD-Incep. ↓
FD-DINOv2 ↓
w/o. Geometry-aware Losses
0.0771
0.2734
0.1515
0.2501
52.6284
423.8286
w/o. Normalised Target
0.0576
0.2529
0.1647
0.2897
47.2389
390.2450
Ours
0.0415
0.2075
0.1829
0.3106
43.2544
363.7326
Table 2: Ablation of stage 2 normalisation and topology-aware losses.
Polygonal meshes are the standard surface representation of modern 3D pipelines, and generating high-quality meshes with artist-style topology is essential for film, gaming, and interactive 3D applications. Mainstream approaches serialize a mesh into a token sequence and decode it autoregressively, which is slow at inference and sensitive to error accumulation, making them impractical for interactive asset creation. We present Meshy T2, a fast native mesh generation framework built on flow matching. At its core is a vertex-set mesh VAE that encodes a mesh into one continuous latent token per vertex and decodes vertices, edge connectivity, and face winding order in a single pass, preserving high-precision geometry and artist-authored topology without vertex quantization or welding. Generation proceeds as a coarse-to-fine cascade of two flow-matching models: an image-conditioned voxel flow first sketches the overall shape as a coarse occupancy scaffold, and a mesh flow then populates the scaffold with per-vertex latent tokens, conditioned on the image, the scaffold, and a requested vertex budget. This design delivers three practical capabilities: interactive generation speed through parallel flow-based synthesis; effective face-count control through the requested vertex budget; and native support for multi-part assets, whose components emerge directly from the generated connectivity. In our experiments, Meshy T2 achieves state-of-the-art geometric fidelity and completes end-to-end image-to-mesh generation within a median of 6 seconds, over an order of magnitude faster than autoregressive baselines. Code and weights will be available at https://github.com/meshy-dev/meshy-t2.
We present MeshFlow, a new method for generating artist-like 3D meshes. Current mesh generators often adopt Auto-Regressive (AR) next-token prediction, a natural choice given the discrete nature of mesh topology. However, AR methods scale poorly because the inference cost is quadratic in mesh size. They also require discretizing the vertex coordinates, which introduces quantization errors. To address these challenges, we introduce a Variational Autoencoder (VAE) that, supervised with a contrastive loss, represents both continuous vertex positions and discrete connectivity in a continuous latent space. This latent space is significantly more compact than prior token-based mesh representations. We then build a 3D generator based on a Rectified Flow transformer, generating all mesh vertices and edges in parallel. Our model generates meshes 18x faster than the fastest AR generator while also achieving excellent accuracy across standard mesh-generation metrics. Homepage: https://mesh-flow.github.io/, Code: https://github.com/facebookresearch/meshflow
Weiyu Li, Antoine Toisoul, Tom Monnier +4
Meta AI · Hong Kong University of Science and Technology
Autoregressive Transformers dominate high-quality mesh generation by producing artist-worthy topologies, yet their inherent sequential decoding induces substantial computational overhead, falling orders of magnitude slower than parallel generative models. On the other hand, while continuous diffusion and flow-matching methods support efficient parallel synthesis across a variety of domains, they cannot be directly applied to meshes: mesh connectivity is inherently discrete and incompatible with standard continuous noise injection and denoising operations. To resolve this fundamental incompatibility, we introduce a compact topology embedder that projects discrete mesh vertex positions and normals into continuous per-vertex embeddings, where the original discrete adjacency information can be faithfully recovered via spacetime distance thresholding. After pretraining and freezing this embedder, any raw mesh can be fully converted into a continuous per-vertex state space unifying position, normal, and implicit topological attributes. Built upon this novel continuous mesh representation, we present PolyFlow, a Transformer-based flow-matching framework that achieves fully parallel vertex state denoising conditioned on extracted point-cloud features. During inference, our model completes generation rapidly via an ODE solver, and supports explicit, precise control over output mesh resolution by directly specifying the target vertex count. Extensive evaluations on the Toys4K benchmark demonstrate that PolyFlow surpasses state-of-the-art autoregressive baselines in both Chamfer Distance and Hausdorff Distance.