Image-to-3D generation has become increasingly capable of producing objects that closely resemble the input image, and an outstanding challenge is to reproduce the depicted object itself, including the specific geometry that defines it. Inscriptions, brand marks, and repeated structures are frequently distorted or lost, despite being critical to object identity. We present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for generating watertight triangle meshes at 20483 resolution. Twinkle3D advances high-fidelity geometry generation along four dimensions. First, while O-Voxel/FaithC offers high representational precision, it often suffers from poor surface quality and non-watertight geometry. We address both issues while retaining its 20483-level precision. Second, we scale diffusion generation to sequences of up to 300K geometric tokens through a redesigned DiT architecture and large-scale distributed training optimizations, reducing training time per step from approximately ten minutes to ten seconds. Third, subsequent refinement cannot fully compensate for errors introduced during initial generation; we therefore strengthen both global shape and local detail in the initial generation stage, and the resulting single-stage model surpasses prior two-stage pipelines with 5123 refinement. Finally, we introduce a fine-grained image-3D cross-modal interaction mechanism that strengthens correspondence between visual evidence and geometric tokens, improving the recovery of object-specific structures. We evaluate geometric fidelity using alignment metrics derived from silhouettes and normal fields. Hi3D 3.0 outperforms four commercial systems across all reported metrics, recovering 82.1% of inscribed characters at 98.2% precision, compared with 21.7% recall for the strongest competitor.
Figures & tables
Model
Global Alignment ↑
Detail Alignment ↑
Multi-view Plausibility ↑
Shape Aesthetics (Q-Judger) ↑
Rodin Gen 2.5
0.8787
0.2645
0.663
46.67
Tripo 3.1
0.8953
0.2826
0.722
50.16
Meshy 7
0.9013
0.3694
0.638
50.00
Hi3D 2.1
0.9038
0.2813
0.636
50.30
Hi3D 3.0
0.9184
0.3782
0.757
50.91
Table 1: Main results on the 33 -input benchmark. Hi3D 3.0 has the best value on all four metrics. Global alignment is silhouette IoU. Detail alignment is the appearance-independent normal bandpass score, computed on the high-confidence-pose subset. Multi-view plausibility is the vision-language completeness score in [0,1] . Shape aesthetics (Q-Judger) is the Qwen-Image-Bench score in [0,100] . All systems run at their highest geometry resolution. Best per column in bold .
Model
Recall ↑
Precision ↑
F1 ↑
Miss ↓
Halluc. ↓
Missed / Wrong / Extra
Hi3D 2.1
0.0%
0.0%
0.0%
99.2%
4.5%
522 / 4 / 25
Rodin Gen 2.5
4.6%
21.8%
7.5%
88.0%
8.2%
463 / 39 / 47
Tripo 3.1
7.4%
9.8%
8.4%
71.3%
32.0%
375 / 112 / 247
Meshy 7
21.7%
62.0%
32.1%
76.0%
9.9%
400 / 12 / 58
Hi3D 3.0
82.1%
98.2%
89.4%
16.5%
0.2%
87 / 7 / 1
Table 2: Inscription fidelity. Hi3D 3.0 recovers 82.1% of the 526 ground-truth characters at 98.2% precision; the best peer system recovers 21.7% . Measured with PaddleOCR on clay-shaded renderings against a shared ground truth. Recall is the share of ground-truth characters recovered. Precision is correct/(correct+wrong+extra) . F1 is their harmonic mean. Miss counts ground-truth characters OCR cannot locate, and Halluc. counts invented characters as a share of ground truth plus extras. Precision and F1 are derived from the counts in the last column. Peer systems run at their latest version and highest geometry resolution. Best per column in bold .
Figure 2: Text-bearing objects. The specification plates and the rim lettering are reproduced as geometry. Each pair shows the input photograph and the Twinkle3D output under normal shading, so nothing in the comparison depends on texture or material. Left: an industrial pump with a model number, pressure and temperature ratings, serial and lot plates, and a cast flow arrow, alongside machined bolt bosses, the flange bore, and the lifting eye. Right: a commemorative medal with lettering curved along its rim over a low-relief scene.
Figure 3: Inscription recovery. The white boxes in each view are enlarged beneath it. (a) The engraved name plate “Shadow” on the mask cabinet, from one frame of a close-up camera path, the same frame for every system, under clay shading. Only Hi3D 3.0 reproduces the word. Tripo 3.1 forms garbled letter-like strokes, the hallucination mode counted in Table 2 ; Meshy 7 leaves faint marks; Rodin Gen 2.5 returns a blank plate; and Hi3D 2.1 forms only a plain bar. (b) Carved calligraphy on a scholar’s diorama, from the same turntable frame for each system. Rodin Gen 2.5 and Hi3D 2.1 were not rendered for this input. The solid and dashed boxes cover two pairs of columns, enlarged left and right. Hi3D 3.0 forms every enlarged character with sharp, deeply cut strokes. Meshy 7 forms the same characters with shallower and softer strokes, and Tripo 3.1 forms pseudo-characters that do not read as the input. Outside the boxes, two characters in the Hi3D 3.0 result are malformed.
Hi3D 3.0 vs.
Structure (same or better)
Geometric Detail (same or better)
Hi3D 2.1
89.0%
94.7%
Meshy 7
78.0%
91.1%
Rodin Gen 2.5
89.8%
95.5%
Tripo 3.1
79.5%
88.1%
Table 3: Blind pairwise human study. Hi3D 3.0 is rated the same or better in 78.0 – 89.8% of structure judgements and 88.1 – 95.5% of detail judgements. Each entry is the share of judgements in which Hi3D 3.0 was rated the same as, or better than, the comparison system.
Figure 4: Autoencoder reconstruction. We pass the same input mesh through the autoencoders of Sparc3D [ 37 ] , TRELLIS 2.0 [ 77 ] , and Twinkle3D. Only Twinkle3D keeps the grille periodic and in phase, and Sparc3D alone tapers the arrow barbs into a cone. Columns are as labelled at the top throughout. White boxes mark the regions enlarged beneath each column. Both cases require cells that carry more than one surface sheet (Sec. 1 ).
Figure 5: One stage vs. two. We compare the standalone first stage of Twinkle3D with a first stage plus 5123 refinement in Sparc3D [ 37 ] and TRELLIS 2.0 [ 77 ] , on three inputs. Columns are as labelled at the top throughout. Twinkle3D runs one stage; both baselines run two. In every case the baselines fuse structures that should be separate, and the single-stage result keeps them distinct.
Figure 6: Five-system comparison with two close-ups per input. Each result is split diagonally, clay shading above and normal shading below. Beneath each view, the left close-up enlarges the solid white box and the right close-up the dashed white box; each box marks the same part of the object in every column, and all close-ups share one source resolution. (a) Front legs; the right-hand tree. Rodin Gen 2.5 returns the stag alone, without the niche, foliage, or base, so its right close-up shows only an antler. Hi3D 3.0, Meshy 7, and Tripo 3.1 form separate leaves and cones on the tree, which Hi3D 2.1 merges into lumpy clusters. (b) Skull and ribcage medallions on the left door. Hi3D 2.1 leaves both medallions blank; Hi3D 3.0 carves the eye sockets and teeth of the skull and separate ribs along the spine. (c) The lighthouse with the rim lettering “PORT OF”; the rim lettering “COMMERCE”. Meshy 7 and Hi3D 3.0 form both words and a windowed lighthouse. Tripo 3.1 misshapes “COMMERCE” and cuts away the inner field of the medal, leaving the lighthouse free-standing. Rodin Gen 2.5 forms “PORT OF” but garbles “COMMERCE” and forms no lighthouse. Hi3D 2.1 forms illegible glyphs and a plain tower. (d) Eye and forehead; side hair. Hi3D 3.0 forms separate tapered strands and a carved iris, where the other systems form clumps, ribbons, or smooth sheets; Meshy 7 also truncates the bust at the neck. (e) Beard; collar and buttons. Hi3D 3.0 and Meshy 7 form beard strands and curls, and Rodin Gen 2.5 forms no beard; all five systems reproduce the collar and buttons. Meshy 7 adds a plinth that is not in the input to both busts.
Figure 7: Close-up comparison on ornament and relief. Left: the input image. Right: one frame of the same close-up camera path for each system, the same frame for every system, under clay shading. The solid and dashed white boxes are enlarged beneath, left and right. (a) Winged figure: the wing feathers and the face with its crown band. Hi3D 3.0 separates individual feathers and carves the eyelids and the edges of the crown band. Tripo 3.1 forms smooth feather lobes and a clean but soft face, Hi3D 2.1 faceted feather plates, Meshy 7 noisy ridges and a face fused with the crown, and Rodin Gen 2.5 a nearly smooth wing and a melted crown band. (b) Qilin: the filigree band of a medallion and the C-scroll volute on its lower edge. Hi3D 3.0 fills the band with scroll relief and forms a spiral volute with nested curls. Tripo 3.1 forms radial grooves and a single tube-profile scroll, Meshy 7 an empty band and a soft curl, and Rodin Gen 2.5 and Hi3D 2.1 no scroll.
Figure 9: Twinkle3D results across six input categories. In each case the left column is the input image and the 2×2 block shows clay-shaded renderings of the generated mesh from complementary viewpoints.
BNRist, Department of Computer Science and Technology, Tsinghua University, China · Tencent ARC Lab, China · Victoria University of Wellington, New Zealand