Recent 3D generation models can produce accurate geometries while still struggling to reconstruct detailed textures. We propose a diffusion-based native 3D material generation model TaoTex, which faithfully recovers intricate textures through tailored strategies and improvements. First, we develop a data construction agent to create high-frequency textured 3D assets to bridge the data gap in public datasets. Training with these data significantly enhances the ability of TaoTex to recover challenging details such as text and patterns. Second, we design a multi-level feature fusion (MLFF) module to adaptively integrate local and global features of the conditional input, providing more complete texture cues for the diffusion model and thereby enhancing reconstruction fidelity. To alleviate VAE reconstruction errors, we adopt a latent-to-pixel space loss transition, further improving the pixel-level details and generation quality. Finally, we scale TaoTex to multi-view inputs by incorporating learnable viewpoint embeddings, achieving accurate and consistent material reconstruction across views. Extensive experiments demonstrate that our method significantly outperforms existing approaches in preserving texture details in both single- and multi-view settings.
Figures & tables
Figure 1: Native 3D materials generated by TaoTex. TaoTex is capable of generating highly faithful materials across a wide range of categories, including challenging details such as text and logos. The geometries shown in the figure above are generated by TRELLIS.2 ( Xiang et al., 2026 ) .
Figure 2: Texture generation and geometric cues. In the green box, geometric variations align well with texture changes, effectively guiding appearance generation. Conversely, the red box contains feature-sparse surfaces where geometry provides little guidance for synthesizing patterns and text.
Figure 3: Method overview. We advance native 3D material detail reconstruction across three fronts: a data agent for high-frequency asset curation, an MLFF module fusing multi-scale features, and a pixel-space loss countering latent compression. Coupled with learnable viewpoint embeddings, TaoTex robustly supports both single- and multi-view inputs.
Figure 4: Effect of the HFT data, MLFF and pixel space loss. The comparisons above validate the efficacy of each proposed component. Specifically, the HFT dataset is indispensable for reconstructing text and patterns; without it, the model fails to synthesize legible text. Furthermore, the MLFF module is crucial for faithfully reproducing textures from input images, while the pixel-space loss further sharpens subtle details, notably minute characters.
Figure 5: Effect of the viewpoint embeddings. Red boxes highlight texture misplacement caused by omitting viewpoint embeddings ( e.g. , left-view textures are erroneously mapped to the back), while arrows indicate severe texture fragmentation in its absence.
Conditioning View Reconstruction
Novel View Generation
Albedo
Rendering
Albedo
Rendering
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
CLIP-FID ↓
CLIP-I ↑
FID ↓
CLIP-FID ↓
CLIP-I ↑
MaterialMVP ( He et al., 2025 )
18.90
0.830
0.160
20.84
0.866
0.150
124.10
17.12
0.8777
102.88
13.19
0.9037
UniTEX ( Liang et al., 2026 )
21.84
0.873
0.101
22.81
0.896
0.097
111.82
16.33
0.8896
104.33
13.88
0.9074
Ink3D ( Han et al., 2026 )
19.39
0.837
0.155
21.24
0.873
0.134
143.89
21.08
0.8633
112.94
14.88
0.9023
TRELLIS.2 ( Xiang et al., 2026 )
21.92
0.879
0.096
23.46
0.907
0.090
113.86
15.21
0.9063
87.69
13.21
0.9150
Table 1: Quantitative comparison under single-view input. Bold : best, underlined : second best.
Albedo
Rendering
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
MaterialMVP ( He et al., 2025 )
20.18
0.873
0.117
22.40
0.900
0.103
TRELLIS.2 ( Xiang et al., 2026 )
19.91
0.852
0.162
21.76
0.883
0.153
Ours
25.48
0.919
0.051
27.30
0.941
0.045
Table 2: Quantitative comparison under multi-view input. Evaluated under four input views.
Figure 6: Single-view comparisons. Our method achieves superior reconstruction fidelity, especially on text and patterns, while synthesizing highly coherent textures across unobserved regions.
Figure 7: Multi-view comparisons. Given four-view inputs, our method produces accurate, consistent and detailed textures, whereas competing methods yield erroneous and incoherent patterns
Figure 8: Comparisons with commercial models. We use the geometry from Rodin Gen2.5 ( Rodin Team, 2026 ) as our geometric input and others use their own geometries. Overall, TaoTex achieves superior performance compared to existing commercial models in fine-grained reconstruction.
Albedo
Rendering
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
W/o HFT data
22.90
0.896
0.087
24.86
0.923
0.077
W/o MLFF
23.40
0.905
0.072
25.34
0.928
0.065
W/o viewpoint embeddings
24.47
0.911
0.062
26.47
0.935
0.054
W/o pixel-space loss
24.61
0.913
0.064
26.22
0.935
0.057
Ours
25.48
0.919
0.051
27.30
0.941
0.045
Table 3: Ablation study. All ablations are evaluated under the multi-view input setting.
Native 3D texture generation synthesizes colors directly in 3D space for a given geometry, conditioned on multi-view reference images. It is generally believed that training such models requires large-scale, high-quality real 3D asset data, whose acquisition remains a long-standing and challenging problem. In this work, we propose Tex-Zero, demonstrating that a high-fidelity native 3D texture generation framework can be trained without 3D assets. Our key observation is that only high-quality and fine-grained color information is essential for 3D texture training, while the required geometric information is less critical and can be manually constructed rather than obtained from real 3D assets. This finding makes it possible to transform abundant, high-quality 2D images into effective training samples for 3D texture generation. Specifically, we convert high-quality 2D images into 3D training samples by representing each image as a plane in 3D space and applying patch-wise random rotations and aggregation to construct complex geometric structures. Using these constructed image data, we train the Tex-Zero VAE, which can reconstruct real 3D assets with high quality despite never observing them during training. Building upon the Tex-Zero VAE, we train the Tex-Zero DiT also exclusively on the constructed image data, where the conditioning 2D multi-view images are transformed into planes in 3D space and also encoded by the Tex-Zero VAE, thereby reducing the representation gap and improving generation quality. Extensive experiments show that Tex-Zero generates high-fidelity 3D textures with fine-grained details solely using images as training data, offering a promising perspective on the data paradigm for scaling 3D texture generation.
Jiangshan Wang, Zeqiang Lai, Jiayi Guo +6
MMLab, CUHK · Tencent Hunyuan · Tsinghua University +3
We present a method for generating high quality materials for 3D objects entirely in texture space. We finetune a video diffusion transformer for text-guided material generation, multi-view material generation, and material upscaling. Our key insight is to use the known projection from image space to texture space, enabling the diffusion process to generalize across arbitrary geometries and texture parameterizations. This approach also avoids the view consistency issues inherent in video and multi-view diffusion models. Because texture space is two dimensional, we can reuse the strong priors of pretrained video diffusion models. We apply our method to high quality material reconstruction from posed photos captured under unknown lighting, as well as to text- and image guided material generation. Our method can scale to high resolutions (8K), 100+ input views, and neural material representations. In quantitative and qualitative evaluations we show state-of-the-art results for material generation and reconstruction.
Recent 3D generative models can synthesize high-quality geometry but often struggle to reproduce intricate textures from reference images, largely due to the scarcity of large-scale 3D training data with rich surface appearance. In contrast, visual generative models are trained on datasets several orders of magnitude larger and excel at modeling complex visual patterns. Motivated by this gap, we introduce Ink3D, a framework that bridges 3D generation with large-scale video generative models to synthesize extremely complex textures. Ink3D first reconstructs a white-mesh geometry using an off-the-shelf 3D generation model. It then employs OrbitPainter, a conditional video generative model, to produce dense orbit-scan videos capturing object appearance across viewpoints. To convert these views into coherent textures, we introduce TextureOptimizer, a neural baking module that integrates dense multi-view observations while mitigating geometry inconsistencies arising from video generation. By decoupling geometry and texture synthesis and leveraging large-scale pretrained video priors, Ink3D enables significantly richer and more faithful texture generation than prior approaches.
Yue Han, Chong Li, Zhening Liu +5
ZGCA & ZGCI · Zhejiang University · Microsoft Research +1