This paper aims to generate materials for 3D meshes from text descriptions. Unlike existing methods that synthesize texture maps, we propose to generate segment-wise procedural material graphs as the appearance representation, which supports high-quality rendering and provides substantial flexibility in editing. Instead of relying on extensive paired data, i.e., 3D meshes with material graphs and corresponding text descriptions, to train a material graph generative model, we propose to leverage the pre-trained 2D diffusion model as a bridge to connect the text and material graphs. Specifically, our approach decomposes a shape into a set of segments and designs a segment-controlled diffusion model to synthesize 2D images that are aligned with mesh parts. Based on generated images, we initialize parameters of material graphs and fine-tune them through the differentiable rendering module to produce materials in accordance with the textual description. Extensive experiments demonstrate the superior performance of our framework in photorealism, resolution, and editability over existing methods. Project page: https://zju3dv.github.io/MaPa
Figures & tables
Figure 1 . Examples from MaPa Gallery , which facilitates photo-realistic 3D rendering by generating materials for daily objects.
Figure 2 . Illustration of our pipeline. Our pipeline primarily consists of four steps: a) Segment-controlled image generation. First, we decompose the input mesh into various segments, project these segments onto 2D images, and then generate the corresponding images using the segment-controlled ControlNet. b) Material grouping. We group segments that share the same material and have similar appearance into a material group. c) Material graph selection and optimization. For each material group, we select an appropriate material graph based on generated images and then optimize this material graph. d) Iterative material recovery. We render additional views of the input mesh with the optimized material graphs, inpaint the missing regions in these rendered images, and repeat steps b) and c) until all segments are assigned with material graphs.
Figure 3 . Downstream editing. We perform material editing on generated material. The user can edit the material using textual prompts through the GPT-4 and a set of predifined APIs.
Dataset
Methods
FID ↓
KID ↓
Overall quality ↑
Fidelity ↑
Chair
TEXTure
94.6
0.044
2.43
2.41
Text2tex
102.1
0.048
2.35
2.51
Fantasia3D
113.7
0.055
1.98
2.08
Ours
88.3
0.037
4.33
3.35
ABO
TEXTure
118.0
0.027
2.65
2.41
Text2tex
109.8
0.020
3.10
2.94
Table 1 . Comparison results. We compare our method with three strong baselines. Our approach yields superior quantitative results and attains the highest ratings in user studies.
Figure 4 . Qualitative comparisons. The results generated by our method and all the baselines are rendered in the same CG environment for comparison. The prompts for the three objects are: "a photo of a wooden bedside table," "a photo of a toy rocket," and "a photo of a brand-new sword."
Figure 5 . Diversity of our generated material. We show the diversity of results synthesized by our framework with the same prompt: “A photo of a toy airplane”. Images in the first row are generated by diffusion models, and models in the second row are our painted meshes.
Figure 6 . Appearance transfer. Our method can also take image prompt as input, transferring the appearance of reference images to input objects.
Figure 7 . Results of downstream editing. We perform text-driven editing on generated material through the GPT-4 and a set of predifined APIs.
Figure 8 . Failure case. Because of the unusual light effects generated by diffusion (silver metallic with yellow light), the albedo estimation network fails to accurately estimate the albedo, leading to dissimilar material prediction.
Figure 9 . Results on watertight meshes. We decompose the mesh into several segments using the graph cut algorithm and generate materials for each segment. The text prompts for those two examples are “A photo of a modern chair with brown legs” and "A black and white eyeglass", respectively.
Figure 10 . Ablation study of “w/o Albedo” and “w/o Grouping”. Example with text prompt “A dark brown leather sofa”. “w/o Albedo” exhibits color disparity between the rendering result and generated image, and fails to merge segments with similar color. “w/o Grouping” has similar rendering results to our method.
Figure 11 . Ablation study of our material graph optimization. We show the ablation study of our material graph optimization and the effectiveness of residual light. The baseline “w/o Opt” and “w/o Res” are optimized to minimize the loss of the generated image (a) and the rendered image. The generated image (b) and (c) are used to prove that our method can yield diverse results even with a small material graph dataset.