cs.CVOct 8, 2026

HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution

Authors: Ziying Li, Shengchu Zhao, Huiang He, Yiyang Chen, Jianwen Huang, Bailin Li, Changhao Li, Jianhui Li, +16 more

Abstract

Image-to-3D generation has become increasingly capable of producing objects that closely resemble the input image, and an outstanding challenge is to reproduce the depicted object itself, including the specific geometry that defines it. Inscriptions, brand marks, and repeated structures are frequently distorted or lost, despite being critical to object identity. We present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for generating watertight triangle meshes at 204832048^{3} resolution. Twinkle3D advances high-fidelity geometry generation along four dimensions. First, while O-Voxel/FaithC offers high representational precision, it often suffers from poor surface quality and non-watertight geometry. We address both issues while retaining its 204832048^{3}-level precision. Second, we scale diffusion generation to sequences of up to 300K geometric tokens through a redesigned DiT architecture and large-scale distributed training optimizations, reducing training time per step from approximately ten minutes to ten seconds. Third, subsequent refinement cannot fully compensate for errors introduced during initial generation; we therefore strengthen both global shape and local detail in the initial generation stage, and the resulting single-stage model surpasses prior two-stage pipelines with 5123512^{3} refinement. Finally, we introduce a fine-grained image-3D cross-modal interaction mechanism that strengthens correspondence between visual evidence and geometric tokens, improving the recovery of object-specific structures. We evaluate geometric fidelity using alignment metrics derived from silhouettes and normal fields. Hi3D 3.0 outperforms four commercial systems across all reported metrics, recovering 82.1% of inscribed characters at 98.2% precision, compared with 21.7% recall for the strongest competitor.

Figures & tables

Explore similar work

CardsList
  1. Pixal3D: Pixel-Aligned 3D Generation from Images

    May 11, 2026Dong-Yang Li, Wang Zhao, Yuxin Chen +53D Asset Generation3D Scene Generation

  2. GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation

    Nov 28, 2025Yuhao Wan, Lijuan Liu, Jingzhi Zhou +6Video Diffusion Models3D Scene Generation