Native 3D texture generation synthesizes colors directly in 3D space for a given geometry, conditioned on multi-view reference images. It is generally believed that training such models requires large-scale, high-quality real 3D asset data, whose acquisition remains a long-standing and challenging problem. In this work, we propose Tex-Zero, demonstrating that a high-fidelity native 3D texture generation framework can be trained without 3D assets. Our key observation is that only high-quality and fine-grained color information is essential for 3D texture training, while the required geometric information is less critical and can be manually constructed rather than obtained from real 3D assets. This finding makes it possible to transform abundant, high-quality 2D images into effective training samples for 3D texture generation. Specifically, we convert high-quality 2D images into 3D training samples by representing each image as a plane in 3D space and applying patch-wise random rotations and aggregation to construct complex geometric structures. Using these constructed image data, we train the Tex-Zero VAE, which can reconstruct real 3D assets with high quality despite never observing them during training. Building upon the Tex-Zero VAE, we train the Tex-Zero DiT also exclusively on the constructed image data, where the conditioning 2D multi-view images are transformed into planes in 3D space and also encoded by the Tex-Zero VAE, thereby reducing the representation gap and improving generation quality. Extensive experiments show that Tex-Zero generates high-fidelity 3D textures with fine-grained details solely using images as training data, offering a promising perspective on the data paradigm for scaling 3D texture generation.
Figures & tables
Figure 1: Compared with previous methods, Tex-Zero trains the VAE and DiT without any real 3D assets and learns a unified latent space for 2D images and 3D textures. With only images as the training data, Tex-Zero achieves high-fidelity generation of real 3D assets.
Figure 2: Constructing 3D training data from 2D images. We treat an image as a colored plane in 3D space, divide it into patches, and independently rotate and arrange these patches to construct a 3D sample with complex geometric structures.
Figure 3: Overview of Tex-Zero. The Tex-Zero VAE encodes target textures and multi-view image conditions into a unified latent space. Conditioned on the image and geometry features, the Tex-Zero DiT generates high-fidelity 3D textures. Although the entire framework is trained exclusively on 2D images, it can directly generate textures for real 3D assets at inference time.
Figure 4: Tex-Zero VAE Reconstruction Results. Although trained without any 3D assets, the Tex-Zero VAE accurately encodes and reconstructs real 3D assets with fine-grained details.
Model
Training Data
3D Asset Reconstruction
2D Image Reconstruction
LPIPS ↓
PSNR-PC ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
Existing methods
FLUX VAE
Image
–
–
–
–
0.0209
39.63
0.933
NaTex VAE
3D Asset
0.0374
31.79
40.94
0.980
0.1977
34.36
0.918
TRELLIS.2 VAE
3D Asset
0.0272
33.94
42.17
0.988
0.1487
33.81
0.915
Sparse VAE
Table 1: VAE Reconstruction on 3D Textures and 2D Images. facb denotes a spatial downsampling factor of a and b latent channels.
Figure 5: Visual Comparison between Tex-Zero and Baselines. Tex-Zero can generate high-fidelity textures with fine-grained details.
Method
Six-view
Front-view
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
NaTex
0.0754
27.74
0.949
0.0669
27.49
0.947
TRELLIS.2
–
–
–
0.1187
21.38
0.900
Tex-Zero
0.0340
35.68
0.983
0.0294
35.31
0.985
Table 2: Quantitive Results for Texture Generation. We compare Tex-Zero with two representative baselines, NaTex ( Lai et al., 2025 ) and TRELLIS.2 ( Xiang et al., 2025a ) .
Training Data
VAE-f16c32
DiT
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
3D
0.0343
41.84
0.981
0.0302
36.23
0.980
2D + 3D
0.0140
42.85
0.988
0.0216
37.68
0.983
Table 3: Joint Training with 2D and 3D Data. Combining 2D images with textured 3D assets improves both VAE reconstruction and DiT generation.
Figure 6: Ablation of the strategy for data preprocessing.
Number of Patches
Random Rotation
Aggregation
LPIPS ↓
PSNR-PC ↑
PSNR ↑
SSIM ↑
1
No
No
0.2160
11.34
18.16
0.777
1
Yes
No
0.1107
20.22
28.38
0.904
4
Yes
No
0.0591
26.83
35.05
0.963
4
Yes
Yes
0.0513
27.81
36.18
0.966
16
Yes
Yes
0.0355
28.97
37.16
0.973
Table 4: Ablation of data preprocessing. We progressively introduce global rotation, patch-wise rotation, and spatial aggregation to explore their effect.
Image Encoder
Training Data
Unified Latent
LPIPS ↓
PSNR ↑
SSIM ↑
DINO
Image
No
0.1228
25.67
0.886
Separate VAE
Image
No
0.0838
29.65
0.955
Shared VAE
3D Asset
Yes
0.0835
30.42
0.966
Tex-Zero VAE
Image
Yes
0.0481
32.54
0.970
Table 5: Effects of different conditioning strategies on DiT generation. Unified 2D-3D VAE with 2D training data achieves the best results.
Figure 7: Visualization of DiT generation results with different conditioning strategies. Our image-trained Tex-Zero with the 2D-3D unified latent space best preserves fine-grained details from the reference images, achieving high-fidelity texture generation.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Input / output channels
3 / 3
Resolution levels
5
Channel widths
[128,256,512,512,512]
Encoder residual blocks per level
2
Decoder residual blocks per level
3
Spatial downsampling factor
16
Appendix
Table 6: Architecture of the f16c16 Tex-Zero VAE. Channel widths are listed from the finest to the coarsest resolution level.
Setting
Value
Texture / normal latent channels
16 / 16
Concatenated input channels
32
Output channels
16
Image-condition latent channels
16
Hidden width
1024
Attention heads
8
Appendix
Table 7: Architecture and conditioning settings of the Tex-Zero DiT.
Setting
VAE
DiT
Optimizer
AdamW
AdamW
Base learning rate
10−4
10−4
Adam betas
(0.9,0.99)
(0.9,0.99)
Adam epsilon
10−6
10−6
Weight decay
10−2
10−2
LR warm-up updates
1
100
Appendix
Table 8: Optimization settings. The micro-batch size is specified per training process, before gradient accumulation.
Recent 3D generative models can synthesize high-quality geometry but often struggle to reproduce intricate textures from reference images, largely due to the scarcity of large-scale 3D training data with rich surface appearance. In contrast, visual generative models are trained on datasets several orders of magnitude larger and excel at modeling complex visual patterns. Motivated by this gap, we introduce Ink3D, a framework that bridges 3D generation with large-scale video generative models to synthesize extremely complex textures. Ink3D first reconstructs a white-mesh geometry using an off-the-shelf 3D generation model. It then employs OrbitPainter, a conditional video generative model, to produce dense orbit-scan videos capturing object appearance across viewpoints. To convert these views into coherent textures, we introduce TextureOptimizer, a neural baking module that integrates dense multi-view observations while mitigating geometry inconsistencies arising from video generation. By decoupling geometry and texture synthesis and leveraging large-scale pretrained video priors, Ink3D enables significantly richer and more faithful texture generation than prior approaches.
Yue Han, Chong Li, Zhening Liu +5
ZGCA & ZGCI · Zhejiang University · Microsoft Research +1
Recent 3D generation models can produce accurate geometries while still struggling to reconstruct detailed textures. We propose a diffusion-based native 3D material generation model TaoTex, which faithfully recovers intricate textures through tailored strategies and improvements. First, we develop a data construction agent to create high-frequency textured 3D assets to bridge the data gap in public datasets. Training with these data significantly enhances the ability of TaoTex to recover challenging details such as text and patterns. Second, we design a multi-level feature fusion (MLFF) module to adaptively integrate local and global features of the conditional input, providing more complete texture cues for the diffusion model and thereby enhancing reconstruction fidelity. To alleviate VAE reconstruction errors, we adopt a latent-to-pixel space loss transition, further improving the pixel-level details and generation quality. Finally, we scale TaoTex to multi-view inputs by incorporating learnable viewpoint embeddings, achieving accurate and consistent material reconstruction across views. Extensive experiments demonstrate that our method significantly outperforms existing approaches in preserving texture details in both single- and multi-view settings.
We present Seed3D 2.0, an advanced 3D content generation system built on Seed3D 1.0, with substantial improvements across generation fidelity, simulation-ready capabilities, and application coverage. For geometry, a coarse-to-fine two-stage pipeline decouples global structure learning from high-frequency detail recovery, while a locality-aware VAE achieves higher spatial compression and more efficient decoding. For texture and material generation, we replace the cascaded pipeline of Seed3D 1.0 with a unified PBR model that directly generates multi-view albedo and metallic-roughness maps, enhanced by Mixture-of-Experts scaling and VLM-based semantic conditioning for improved material precision and visual fidelity. Beyond single-object generation, Seed3D 2.0 introduces a simulation-ready model suite comprising scene layout planning, part-aware decomposition, and training-free articulation generation, enabling coherent scene construction and part-level physical interaction across physics and graphics engines. A large-scale human preference study against five recent commercial models shows that Seed3D 2.0 achieves consistent win rates of 69.0% to 89.9% in textured 3D asset generation. Seed3D 2.0 is available on https://exp.volcengine.com/ark/vision?_vtm_=0.0.c70961.d701978.0&mode=vision&modelId=doubao-seed3d-2-0-260328&tab=Gen3D