Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define coverage and physically based rendering (PBR) channels, making dependencies between patterns, color, and relief explicit. A standalone interpreter evaluates the program into material maps, while the source retains named fields and layer parameters for subsequent authoring. Without task-specific fine-tuning, our pipeline uses parser-guided repair and preview-based critique to revise material designs, then searches noise seeds while keeping each candidate's remaining source fixed. On a curated benchmark of 141 prompts evaluated with six backbones, our best-performing configuration achieves higher mean scores than three diffusion baselines on all four flat-layout prompt-alignment metrics. Its initial programs already exceed all three baselines on mean BLIPScore, before critique or seed search. Retained programs have a median length of 21 lines when pooled across backbones. In a blind four-way comparison involving 30 participants and 20 prompts, our renders receive 59.2% of choices, compared with 19.3% for the most-preferred baseline. Compact executable programs thus offer a way to generate prompt-aligned materials while retaining their construction as part of the asset.
Figures & tables
Figure 1: Selected text-to-material outputs from MatLoom , rendered in the flat layout. Each render comes from a compact executable material program.
Figure 2: Overview of the MatLoom representation using the tile program of Listing 1 . A program defines an optional View , reusable fields, and a bottom-to-top Material stack. Shared fields couple coverage with color, roughness, or height, explicit noise seeds pin a realization, scalar and color channels use the over rule of Eq. 1 , and height uses a separate maximum rule before normals and approximate ambient occlusion are derived. The shared tileMask (orange) drives both the glaze layer’s coverage and its height, and explicit noise seed values (blue) pin one stochastic realization.
Figure 3: A visual walkthrough of material authoring. A text request becomes a layered program, whose preview, channel statistics, and source inform critique and revision. Stage III depicts one candidate’s seed sweep. The full search applies this sweep to the distinct initial, selected, and final programs (Section 4.3 ). Each sweep includes its original program and changes only noise seeds. The program excerpt, intermediate appearances, stars, and 68/100 score are schematic illustrations of information flow, not a measured trajectory or evidence of monotonic improvement.
Flat
Staged
Method
BLIPScore
CLIPScore
VQAScore
Judge
BLIPScore
CLIPScore
VQAScore
Judge
IntrinsiX ( Kocsis et al., 2025 )
29.94
24.29
43.61
56.10
20.96
21.80
41.97
49.60
MatFuse ( Vecchio et al., 2024 )
8.27
20.12
30.61
36.36
8.58
19.25
30.56
33.63
StableMaterials ( Vecchio, 2026 )
28.57
25.66
45.36
57.41
21.33
23.59
45.53
54.29
MatLoom (gemma-4-26b-a4b-it)
33.56
25.46
47.12
47.48
22.07
22.48
47.43
40.17
MatLoom (qwen3.6-35b-a3b)
35.64
25.70
46.36
45.80
21.93
22.55
46.60
37.47
Table 1: Main comparison on the 141 -prompt benchmark under four alignment metrics (mean over three recorded runs, higher is better). Flat and Staged give the two rendering layouts. Our rows use the full three-stage pipeline with the named generator and the same model as critic, with text-only critique for DeepSeek and GLM. Best per layout half in bold .
Figure 4: Qualitative comparison. Each cell stacks the flat layout over the staged layout. The examples were selected to expose structure and channel behavior, not to estimate typical performance or isolate the representation.
Flat
Staged
Stage
BLIPScore
CLIPScore
VQAScore
Judge
BLIPScore
CLIPScore
VQAScore
Judge
Round 0
48.31
27.70
53.34
61.35
30.95
24.37
52.67
52.98
Round 5
51.86
28.25
52.93
67.61
33.23
24.91
53.43
56.82
Selected
54.17
28.71
53.97
66.56
35.50
25.25
54.07
57.20
Final (full)
56.06
28.80
54.71
67.01
36.14
25.30
54.26
57.34
Table 2: Stage-wise scores for gemini-3.6-flash , averaged over 423 runs. Round 0 precedes critique, Round 5 is the last revision, Selected maximizes the quick scorer over the trajectory, and Final additionally searches seeds of the first, selected, and last programs. Best per metric in bold .
Figure 5: Explicit edits to one material program. We manually change roughness, a shared grout parameter, or base-color offsets while preserving the remaining source and all three noise seeds. This is one illustrative case, not an automated-edit benchmark.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Work
User conditioning
Generated representation
Execution dependence
Relevant distinction
Raster generators ( Vecchio et al., 2024 ; Vecchio, 2026 ; Kocsis et al., 2025 )
Text; additional modalities vary
PBR raster maps
Renderer consuming the maps
Material appearance without an explicit generative program
Conditional MatFormer ( Hu et al., 2023 )
Text, image, or partial graph
Procedural node graph
Substance ecosystem
Prior text-conditioned procedural synthesis
VLMaterial ( Li et al., 2025 )
Image
Python program constructing a shader graph
Blender API
Learned image-to-program synthesis
MultiMat ( Belouadi et al., 2026 )
Image or unconditional
CompactSBS, a compact YAML graph program
Substance Designer
Intermediate visual feedback and incremental validation
MatLayerNet
Text
Layer plans, per-layer parameters, and library-derived masks over PBR maps
MetaGPT multi-agent pipeline with a curated mask-generator library
Prior language-guided substrate, texture, and aging layers ( Cai et al., 2026 )
Closest text-to-procedural system, using process retrieval
Appendix
Table 3: Task and representation boundaries. User conditioning is distinct from textual program encoding: MultiMat consumes compact textual programs during synthesis without evaluating natural-language-conditioned generation. Language standards are representation precedents, not learned generation systems. These properties do not imply a quality ranking, and matched generation and editing comparisons remain necessary.
Budget K
Mean gain
% of K=1000
Runs at best (%)
1
+0.25
18
0
10
+0.68
48
1
50
+0.96
67
6
100
+1.07
75
9
250
+1.21
85
22
500
+1.34
94
50
Appendix
Table 4: Stage III’s sweep budget replayed from the stored per-variant quick-scorer scores of the flagship’s 423 runs. The gain is the quick-scorer margin over the run’s best pre-polish program (median 31.5 ), in quick-scorer points rather than evaluation metrics, non-negative because the selection keeps the original unless a variant beats it. “Runs at best” is the share of runs whose truncated sweep already contains the full sweep’s winner.
Flat
Staged
Round
BLIPScore
CLIPScore
VQAScore
Judge
BLIPScore
CLIPScore
VQAScore
Judge
0
48.31
27.70
53.34
61.35
30.95
24.37
52.67
52.98
1
+ 1.30
+ 0.13
− 0.76
+ 3.71
+ 0.57
+ 0.27
+ 0.02
+ 1.80
2
+ 1.74
+ 0.22
− 0.41
+ 4.19
+ 1.38
+ 0.40
+ 0.22
+ 2.28
3
+ 1.89
+ 0.31
− 0.40
+ 3.87
+ 1.66
+ 0.46
+ 0.37
+ 3.32
4
+ 2.95
+ 0.38
− 0.75
+ 3.32
+ 2.02
+ 0.46
− 0.02
+ 3.04
Appendix
Table 5: Refinement rounds for gemini-3.6-flash (the Flat and Staged column groups give both layouts, mean over all runs). Round 0 is the initial program’s absolute score. Rounds 1 – 5 are signed deltas from it. Best round per layout half in bold .
Flat
Staged
Program
BLIPScore
CLIPScore
VQAScore
Judge
BLIPScore
CLIPScore
VQAScore
Judge
First
48.31
27.70
53.34
61.35
30.95
24.37
52.67
52.98
Selected
54.17
28.71
53.97
66.56
35.50
25.25
54.07
57.20
Last
51.86
28.25
52.93
67.61
33.23
24.91
53.43
56.82
Oracle
64.80
30.01
60.71
79.24
44.48
26.31
60.42
69.07
Win rate
42.1
41.4
38.1
30.5
42.8
43.3
36.4
35.0
Appendix
Table 6: Quick-scorer selection over the trajectory of gemini-3.6-flash in both layouts. Selected is the MobileCLIP2 argmax over six programs, last uses round 5 , and oracle uses each evaluation metric to select its own best program. The oracle is an upper bound unavailable to the deployed selector. Win rate counts strictly higher scores than last, with ties not treated as losses, and Δ is the mean signed difference.
Flat
Staged
Polished
BLIPScore
CLIPScore
VQAScore
Judge
BLIPScore
CLIPScore
VQAScore
Judge
First
+ 2.22
+ 0.17
+ 0.49
+ 0.64
+ 0.23
+ 0.08
+ 0.33
+ 0.27
Selected
+ 1.54
+ 0.12
+ 0.22
+ 0.13
+ 0.17
+ 0.04
+ 0.09
+ 0.59
Last
+ 1.63
+ 0.18
+ 0.36
+ 0.11
+ 1.09
+ 0.09
+ 0.25
+ 0.90
Mean
+ 1.80
+ 0.16
+ 0.36
+ 0.30
+ 0.50
+ 0.07
+ 0.22
+ 0.59
Appendix
Table 7: Signed mean deltas from seed search ( 1000 variants, original retained) applied to the first, selected, and last trajectory programs, with a Mean row averaging those 3 conditions. All retained mean deltas are positive, but no per-run or statistical improvement guarantee follows.
Backbone
Median total (min)
Median polish share (%)
Polish p90 (min)
gemini-3.6-flash
4.6
68
5.6
gpt-5.6-luna
6.3
43
5.2
gemma-4-26b-a4b-it
7.7
33
3.6
glm-5.2
10.3
27
4.7
qwen3.6-35b-a3b
12.8
20
3.6
deepseek-v4-flash-0731
20.0
18
7.2
Appendix
Table 8: Recorded authoring latency per backbone over 423 runs each, with 3 stochastic runs per prompt. The polish share is the fraction of total wall time spent in Stage III seed search, and p90 is the 90th percentile of that stage’s duration.
Flat
Staged
Configuration
BLIPScore
CLIPScore
VQAScore
Judge
BLIPScore
CLIPScore
VQAScore
Judge
Full (ours)
51.23
27.53
50.01
56.76
32.90
24.48
49.95
47.60
w/o playbook
− 4.99
− 0.52
− 1.47
− 0.08
− 4.66
− 0.54
− 1.30
− 1.87
w/o few-shot
− 1.06
+ 0.32
+ 1.60
+ 0.69
− 2.53
+ 0.18
+ 2.06
+ 2.40
w/o both
+ 0.52
+ 0.31
+ 0.78
+ 0.21
+ 0.17
+ 0.14
+ 0.61
+ 0.64
w/o render critique
− 6.84
− 1.03
+ 0.08
− 2.42
− 4.27
− 0.55
− 0.10
− 1.64
Appendix
Table 9: Exploratory component configurations in both layouts. The first row contains the full reference’s absolute means for one authoring run per prompt, and subsequent rows are signed deltas. Diagnostics denotes the material-statistics panel.
Flat
Staged
Critic
BLIPScore
CLIPScore
VQAScore
Judge
BLIPScore
CLIPScore
VQAScore
Judge
Self (gpt)
51.23
27.53
50.01
56.76
32.90
24.48
49.95
47.60
Gemini
+ 3.56
+ 0.99
+ 1.74
+ 5.08
+ 1.77
+ 0.83
+ 2.01
+ 2.99
DeepSeek
− 0.40
− 0.13
+ 1.22
− 2.80
− 0.04
+ 0.09
+ 1.17
+ 0.65
GLM
− 8.57
− 0.51
+ 0.87
+ 7.16
− 2.20
− 0.14
+ 2.23
+ 10.20
Qwen
− 5.75
− 0.53
− 1.22
− 4.71
− 4.37
− 0.34
+ 1.75
− 2.44
Appendix
Table 10: Exploratory cross-critic configurations with the gpt-5.6-luna generator. The reference contains absolute means, and subsequent rows contain signed deltas. Critic modalities differ as described above. Largest observed delta per column is in bold .
Figure 6: Three illustrative CMA-ES outputs from separate runs for “a red brick wall” in the flat layout, ordered by the authors’ visual assessment rather than by optimization time. Quick-scorer gains over each run’s pre-polish program are +2.0 , +3.7 , and +0.7 , with the largest gain belonging to the middle image. The selected images illustrate scorer disagreement and do not establish a general failure rate.
Method
Share of forced choices (%)
Mean fidelity ( 1 – 7 )
MatLoom (ours)
59.2[55.8,62.7]
4.98[4.73,5.21]
StableMaterials
19.3[16.5,22.3]
3.60[3.31,3.89]
IntrinsiX
15.3[13.3,17.2]
3.01[2.77,3.26]
MatFuse
6.2[4.3,8.2]
2.27[2.07,2.49]
Appendix
Table 11: Blind preference study over 30 participants and 20 prompts. Brackets in both columns give 95% participant-cluster bootstrap intervals. The forced-choice column shares sum to 100% over the 600 trials.
Prompt source
MatLoom
StableMaterials
IntrinsiX
MatFuse
Rating gap
GenProc
40.7
35.3
12.0
12.0
+0.35
MatSynth
62.0
8.7
26.7
2.7
+1.95
StableMaterials
57.3
18.7
17.3
6.7
+1.13
text2fabric
76.7
14.7
5.3
3.3
+2.06
Appendix
Table 12: Descriptive user-study results by prompt source. Each row pools five prompts and 150 choices from the same 30 participants. The four method columns give shares of forced choices in percent. Rating gap is the mean paired fidelity rating difference ( MatLoom minus StableMaterials) on the 1 – 7 scale.
Method
Backbone
Steps
Guidance
Height map
MatFuse
latent diffusion
50
5.0
✗
StableMaterials
SD-class LDM + LCM
4
10.0 (LCM)
✓
IntrinsiX
FLUX.1-dev + LoRA
28
3.5
✗
MatLoom (ours)
LLM program
n/a
n/a
✓
Appendix
Table 13: Recorded baseline generation settings at 512×512 map resolution and three run labels per prompt. Height-map availability does not imply matched rendering: StableMaterials uses bump shading while MatLoom uses geometric displacement.
Method
Flat CLIP-IQA
Staged CLIP-IQA
IntrinsiX
46.51
24.08
MatFuse
19.39
10.68
StableMaterials
52.17
21.27
MatLoom (gemma-4-26b-a4b-it)
47.51
28.95
MatLoom (qwen3.6-35b-a3b)
44.65
28.42
MatLoom (deepseek-v4-flash-0731)
49.68
29.00
Appendix
Table 14: CLIP-IQA image-quality diagnostic, averaged across 141 prompts and three runs per method, with higher values indicating better predicted image quality. Best mean per layout is in bold . This metric does not measure prompt alignment.
BLIPScore
CLIPScore
VQAScore
Judge
Source ( n )
ours
SM
ours
SM
ours
SM
ours
SM
Hu et al. ( 30 )
45.12
30.30
27.53
26.12
63.50
58.39
71.43
72.59
MatSynth ( 11 )
73.72
28.46
30.83
25.05
59.04
42.69
76.30
38.82
StableMaterials ( 50 )
48.24
34.34
26.83
25.17
43.98
36.64
56.35
59.40
text2fabric ( 50 )
66.55
21.80
31.09
26.00
59.21
46.85
72.97
50.39
Appendix
Table 15: Flat-layout means by benchmark source for the flagship ( MatLoom , final program) and StableMaterials (SM). Best mean per pair is in bold . These descriptive comparisons have no per-source significance claim.
Idiom
Use
Multi-octave layering
A low-frequency fBm for broad form plus a high-frequency one for fine grain, reused wherever each scale is needed.
Anisotropic striation
fBm / Worley with base_freq_x = base_freq_y elongates features along an axis (bark, brushed metal).
Coordinate-domain warping
Add a noise to a coordinate term before a Sin pattern to turn banding into turbulent stratification, and drive height from the same warped pattern.
Micro-gloss grain
A high-frequency, low-amplitude fBm added only to roughness for skin, bark, or stone microstructure.
Substrate-matrix-first
Continuous substrate in a bottom Layer(1) , discrete features as upper layers whose thresholded masks let the substrate show through.
Appendix
Table 16: The organic-texture playbook’s 5 idioms, as they appear in the system prompt.
Figure 8: Selected refinement trajectory for “Giraffe skin with large polygonal patches.” The figure-selection procedure filters for BLIPScore/judge gains and ranks candidates by visual change. This example’s flat BLIPScore rises from 2.8 to 99.8 , and its round- 5 judge score is 16 points above round 0 . The polished final output is omitted because it contains a smudge accepted by the quick scorer. This illustration does not estimate the frequency of successful or monotonic refinement.
This paper aims to generate materials for 3D meshes from text descriptions. Unlike existing methods that synthesize texture maps, we propose to generate segment-wise procedural material graphs as the appearance representation, which supports high-quality rendering and provides substantial flexibility in editing. Instead of relying on extensive paired data, i.e., 3D meshes with material graphs and corresponding text descriptions, to train a material graph generative model, we propose to leverage the pre-trained 2D diffusion model as a bridge to connect the text and material graphs. Specifically, our approach decomposes a shape into a set of segments and designs a segment-controlled diffusion model to synthesize 2D images that are aligned with mesh parts. Based on generated images, we initialize parameters of material graphs and fine-tune them through the differentiable rendering module to produce materials in accordance with the textual description. Extensive experiments demonstrate the superior performance of our framework in photorealism, resolution, and editability over existing methods. Project page: https://zju3dv.github.io/MaPa
Shangzhan Zhang, Sida Peng, Tao Xu +7
Zhejiang University · Ant Group · Shenzhen University
Procedural material creation underpins applications in digital content creation, visual effects, and 3D asset design. Achieving high-quality results requires more than reproducing node graphs -- it demands understanding the process by which experts construct materials. We formulate procedural material generation as retrieval-time process reasoning over expert demonstrations, elevating process to a first-class representation beyond graph-only synthesis. Concretely, we represent expert workflows as process traces: textual records of construction steps, parameters, and design intent. To instantiate this idea, we use a pretrained LLM-based ProcessSynthesizer to synthesize a process trace aligned with a user's intent and a pretrained LLM-based Compiler to ground the process trace into an executable Blender material graph. Because procedural expertise is most naturally conveyed through demonstrations, we leverage tutorial videos as a source of process knowledge and extract textual, LLM-compatible traces using automated video analysis tools. In an expert study with five Blender artists (avg. 7.5 years of experience), materials generated by reflecting expert demonstrations were found to produce workflows requiring fewer edits, and more closely match professional design strategies than methods operating solely on static artifacts. A user study with 150 participants further shows that our approach achieves superior generation and editing performance compared to prior procedural systems. All code, models, and data will be available at https://materialapprentice.github.io
Kunal Gupta, Gaurav Joshi, Yen-Ru Chen +3
University of California San Diego, La Jolla, CA, USA
Rapid identification of candidate materials with target properties has become a key task in materials science. Machine learning has emerged as an alternative to physics-based simulation, offering a faster and cheaper way to filter materials based on their stability and other target properties, reducing the number of candidates that reach the costly synthesis stage. Recently, Large Language Models (LLMs) have been applied to this role, but these models are parameter-heavy and computationally expensive both during training and at inference time, making them unsuitable for high-throughput tasks. This inefficiency stems from both the large over-parameterization of language models and the difficulty of framing material generation as a sequence learning problem. In this paper, we present PRISMat, a cost-effective, permutation-invariant model, which addresses these limitations. We show that PRISMat, despite taking less time for inference, is able to outperform LLMs in generating crystal slabs conditioned on critical materials' surface properties. In targeted material discovery, we achieve mean absolute errors of 0.188 eV/A2 and 2.79 eV for cleavage energy and work function tasks, respectively, reducing the error of the next best model by 4×.
Claire Schlesinger, Circe Hsu, Peter Schindler +1
Khoury College of Computer Sciences Northeastern University Boston, MA 02115 · College of Engineering Northeastern University Boston, MA 02115