MeshOctave generates meshes via cascading resolution transitions
Organizations: Huazhong University of Science and Technology · Meshy AI · Independent Researcher · Technical University of Munich
Abstract
Generating compact, artist-style meshes with explicit topology typically relies on autoregressive models which incur prohibitive sequential per-token costs, or continuous flow models that depend on heuristic connectivity decoders. Next-scale generation paradigms offer a compelling alternative by enabling parallel intra-scale token prediction and coarse-to-fine refinement from global structure to local topology; yet, existing methods derive hierarchical scales via progressive mesh simplification and invert them sequentially. This eliminates intra-scale parallelism and scales generation steps linearly with face count. In this paper, we propose MeshOctave, which instead defines scale through dyadic spatial grid resolutions, framing coarsening as a deterministic collapse that merges vertices sharing a voxel cell and inherits connectivity. Its inverse operation, split-and-rewire, determines which octant sub-vertices are instantiated for each coarse face and resolves local connectivity using discrete structural tokens. These per-face operations require no serialization, each scale transition is modeled as an unordered set that adds one bit of coordinate precision, naturally supporting dynamic-length meshes and adaptive resolution refinement. We construct a scale-conditioned masked-uniform discrete diffusion model to learn split-and-rewire operation from resolution collapse hierarchies. MeshOctave outperforms strong baselines in geometric fidelity and topological validity by a non-trivial margin, while supporting adaptive resolution refinement and extending naturally to mesh subdivision tasks.
Figures & tables
| Metric | Mesh Any.V2 | Vertex Regen | AR Mesh | Fast Mesh | Mesh Silk. | BPT | Deep Mesh | Mesh Ripple | Mesh Flow | LATO | LATO.2 | MESH OCTAVE |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Type | AR | AR+ N | AR+ N | AR | AR | AR | AR | AR | Flow | Flow | Flow | DD+ N |
| CD-L2 | 0.1103 | 0.0874 | 0.0804 | 0.0647 | 0.0622 | 0.0610 | 0.0505 | 0.0468 | 0.0452 | 0.0430 | 0.0406 | 0.0392 |
| CD-L1 | 0.1539 | 0.1238 | 0.1146 | 0.0925 | 0.0895 | 0.0873 | 0.0729 | 0.0691 | 0.0672 | 0.0621 | 0.0595 | 0.0581 |
| HD | 0.2317 | 0.1875 | 0.1569 | 0.1151 | 0.1440 | 0.1111 | 0.0957 | 0.0948 | 0.0765 | 0.0741 | 0.0665 | 0.0645 |
| 0.6807 | 0.6655 | 0.6058 | 0.7413 | 0.7728 | 0.7976 | 0.8187 | 0.8141 | 0.8237 | 0.8273 | 0.8333 | 0.8478 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Hourglass level | Token Corresponds to | Anchor | Type |
| face, coordinate-level | |||
| face, vertex-level | subdivision / intra / inter aggregate | ||
| face, face-level | face feature |