This paper aims to generate materials for 3D meshes from text descriptions. Unlike existing methods that synthesize texture maps, we propose to generate segment-wise procedural material graphs as the appearance representation, which supports high-quality rendering and provides substantial flexibility in editing. Instead of relying on extensive paired data, i.e., 3D meshes with material graphs and corresponding text descriptions, to train a material graph generative model, we propose to leverage the pre-trained 2D diffusion model as a bridge to connect the text and material graphs. Specifically, our approach decomposes a shape into a set of segments and designs a segment-controlled diffusion model to synthesize 2D images that are aligned with mesh parts. Based on generated images, we initialize parameters of material graphs and fine-tune them through the differentiable rendering module to produce materials in accordance with the textual description. Extensive experiments demonstrate the superior performance of our framework in photorealism, resolution, and editability over existing methods. Project page: https://zju3dv.github.io/MaPa
Figures & tables
Figure 1 . Examples from MaPa Gallery , which facilitates photo-realistic 3D rendering by generating materials for daily objects.
Figure 2 . Illustration of our pipeline. Our pipeline primarily consists of four steps: a) Segment-controlled image generation. First, we decompose the input mesh into various segments, project these segments onto 2D images, and then generate the corresponding images using the segment-controlled ControlNet. b) Material grouping. We group segments that share the same material and have similar appearance into a material group. c) Material graph selection and optimization. For each material group, we select an appropriate material graph based on generated images and then optimize this material graph. d) Iterative material recovery. We render additional views of the input mesh with the optimized material graphs, inpaint the missing regions in these rendered images, and repeat steps b) and c) until all segments are assigned with material graphs.
Figure 3 . Downstream editing. We perform material editing on generated material. The user can edit the material using textual prompts through the GPT-4 and a set of predifined APIs.
Dataset
Methods
FID ↓
KID ↓
Overall quality ↑
Fidelity ↑
Chair
TEXTure
94.6
0.044
2.43
2.41
Text2tex
102.1
0.048
2.35
2.51
Fantasia3D
113.7
0.055
1.98
2.08
Ours
88.3
0.037
4.33
3.35
ABO
TEXTure
118.0
0.027
2.65
2.41
Text2tex
109.8
0.020
3.10
2.94
Table 1 . Comparison results. We compare our method with three strong baselines. Our approach yields superior quantitative results and attains the highest ratings in user studies.
Figure 4 . Qualitative comparisons. The results generated by our method and all the baselines are rendered in the same CG environment for comparison. The prompts for the three objects are: "a photo of a wooden bedside table," "a photo of a toy rocket," and "a photo of a brand-new sword."
Figure 5 . Diversity of our generated material. We show the diversity of results synthesized by our framework with the same prompt: “A photo of a toy airplane”. Images in the first row are generated by diffusion models, and models in the second row are our painted meshes.
Figure 6 . Appearance transfer. Our method can also take image prompt as input, transferring the appearance of reference images to input objects.
Figure 7 . Results of downstream editing. We perform text-driven editing on generated material through the GPT-4 and a set of predifined APIs.
Figure 8 . Failure case. Because of the unusual light effects generated by diffusion (silver metallic with yellow light), the albedo estimation network fails to accurately estimate the albedo, leading to dissimilar material prediction.
Figure 9 . Results on watertight meshes. We decompose the mesh into several segments using the graph cut algorithm and generate materials for each segment. The text prompts for those two examples are “A photo of a modern chair with brown legs” and "A black and white eyeglass", respectively.
Figure 10 . Ablation study of “w/o Albedo” and “w/o Grouping”. Example with text prompt “A dark brown leather sofa”. “w/o Albedo” exhibits color disparity between the rendering result and generated image, and fails to merge segments with similar color. “w/o Grouping” has similar rendering results to our method.
Figure 11 . Ablation study of our material graph optimization. We show the ablation study of our material graph optimization and the effectiveness of residual light. The baseline “w/o Opt” and “w/o Res” are optimized to minimize the loss of the generated image (a) and the rendered image. The generated image (b) and (c) are used to prove that our method can yield diverse results even with a small material graph dataset.
Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define coverage and physically based rendering (PBR) channels, making dependencies between patterns, color, and relief explicit. A standalone interpreter evaluates the program into material maps, while the source retains named fields and layer parameters for subsequent authoring. Without task-specific fine-tuning, our pipeline uses parser-guided repair and preview-based critique to revise material designs, then searches noise seeds while keeping each candidate's remaining source fixed. On a curated benchmark of 141 prompts evaluated with six backbones, our best-performing configuration achieves higher mean scores than three diffusion baselines on all four flat-layout prompt-alignment metrics. Its initial programs already exceed all three baselines on mean BLIPScore, before critique or seed search. Retained programs have a median length of 21 lines when pooled across backbones. In a blind four-way comparison involving 30 participants and 20 prompts, our renders receive 59.2% of choices, compared with 19.3% for the most-preferred baseline. Compact executable programs thus offer a way to generate prompt-aligned materials while retaining their construction as part of the asset.
Anson Y. Lam, Shuqing Li, Michael R. Lyu
Department of Computer Science and Engineering The Chinese University of Hong Kong Hong Kong, China
Recently, diffusion-based material transfer methods rely on image fine-tuning or complex architectures with auxiliary networks but face challenges such as text dependency, additional computational costs, and feature misalignment. To address these limitations, we propose \textbf{DealMaTe}, using \underline{\textbf{de}}pth, norm\underline{\textbf{a}}l, and \underline{\textbf{l}}ighting images for \underline{\textbf{ma}}terial \underline{\textbf{t}}ransf\underline{\textbf{e}}r. DealMaTe is a simplified diffusion framework that eliminates text guidance and reference networks. We design a lightweight 3D information injection method, Multi-Dim 3D Shader LoRA, which, without modifying the base model weights, enables compatible control conditions and achieves harmonious and stable results. Additionally, we optimize the attention mechanism with Shader Causal Mutual Attention and key-value (KV) caching to reduce inference latency caused by multiple conditions, improve computational efficiency, and achieve high-quality material transfer results with low architectural complexity. Extensive experiments covering a wide variety of objects and lighting conditions consistently demonstrate that DealMaTe achieves remarkable high-fidelity material transfer under arbitrary input materials. The code is available at https://github.com/haha-lisa/DealMaTe.
Nisha Huang, Yizhou Lin, Jie Guo +3
Tsinghua University, China · Pengcheng Laboratory, China · National Cheng-Kung University, Taiwan +2
Recent diffusion-based methods for material transfer rely on image fine-tuning or complex architectures with assistive networks, but face challenges including text dependency, extra computational costs, and feature misalignment. To address these limitations, we propose MaTe, a streamlined diffusion framework that eliminates textual guidance and reference networks. MaTe integrates input images at the token level, enabling unified processing via multi-modal attention in a shared latent space. This design removes the need for additional adapters, ControlNet, inversion sampling, or model fine-tuning. Extensive experiments demonstrate that MaTe achieves high-quality material generation under a zero-shot, training-free paradigm. It outperforms state-of-the-art methods in both visual quality and efficiency while preserving precise detail alignment, significantly simplifying inference prerequisites.
Nisha Huang, Henglin Liu, Yizhou Lin +5
Tsinghua University · PengCheng Laboratory · Lenovo Research +1