We present a method for generating high quality materials for 3D objects entirely in texture space. We finetune a video diffusion transformer for text-guided material generation, multi-view material generation, and material upscaling. Our key insight is to use the known projection from image space to texture space, enabling the diffusion process to generalize across arbitrary geometries and texture parameterizations. This approach also avoids the view consistency issues inherent in video and multi-view diffusion models. Because texture space is two dimensional, we can reuse the strong priors of pretrained video diffusion models. We apply our method to high quality material reconstruction from posed photos captured under unknown lighting, as well as to text- and image guided material generation. Our method can scale to high resolutions (8K), 100+ input views, and neural material representations. In quantitative and qualitative evaluations we show state-of-the-art results for material generation and reconstruction.
Figures & tables
Figure 2: Texture Space Observations. We leverage known geometry with a non-overlapping texture parameterization. Given camera poses and geometry, we project each view back into texture space, which results in partially covered texture space observations. By formulating our diffusion process in the 2D texture space domain, we can leverage powerful image- and video diffusion model priors to synthesize high quality materials.
Figure 3: Method Overview. Our pipeline leverages diffusion transformers in texture space to synthesize high quality materials. We support multiple input modalities: 1) one or more views projected into texture space, which can be photographs, rendered images, or frames generated by diffusion models 2) a text prompt describing the material 3) low-resolution material maps. In all cases, we condition the model on world space positions, normals and a text prompt. Our finetuned DiT successfully demodulates the (unknown) lighting, and creates clean albedo, roughness, and metallicity maps with high frequency details. The materials can be directly applied in standard content creation tools.
Figure 4: Reconstruction - Real-World. Material reconstructions from multi-view observations (photographs with known poses) on two examples from the DTC dataset ( Dong et al., 2025 ) . For optimization-based differentiable path tracing (DiffPT) and our method, we leverage known geometry, while LSRM jointly reconstructs geometry and materials. In the right part we show the reconstructed materials under two novel lighting conditions.
Real recon (136 img)
Synthetic recon (256 img)
Synthetic relight (2k img)
Geom.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Known
DiffPT
30.67
0.959
0.0307
27.50
0.929
0.0739
24.90
0.903
0.0911
(ref.)
Ours
29.03
0.947
0.0331
31.24
0.960
0.0370
30.46
0.950
0.0401
Recon.
LSRM
27.75
0.933
0.0388
22.82
0.872
0.1210
21.82
0.862
0.1241
(LSRM)
Ours (LSRM geo)
27.01
0.936
0.0423
24.10
0.890
0.0896
23.71
0.885
0.0927
Table 1: Multi-view to materials. Comparison on synthetic materials and posed real photographs. Synthetic metrics are averaged over 32 extracted materials rendered from eight views. For the relighting metrics, we render each view with eight different HDR probes. Real metrics are averaged over eight examples from the DTC dataset ( Dong et al., 2025 ) . Materials are reconstructed from 17 views in 2K resolution. We compare against DiffPT ( Hasselgren et al., 2022 ) and LSRM ( Li et al., 2026b ) . DiffPT overfits (reconstruction w/ single lighting) but fails to generalize (relighting).
Figure 5: Reconstruction - Generated Frames. Material reconstructions from imperfect views, with views synthesized by an off-the-shelf depth-conditioned video model: Wan2.2-VACE-Fun-A14B ( Jiang et al., 2025 ) . Compared to an optimization-based approach, our method robustly reconstructs materials also from the inconsistent views from the video model.
Figure 6: Generation - Single-View. Materials from our generative method, conditioned on a single image and text description. The visualized view is rotated slightly from the conditioning view (leftmost insets). We use reference geometry and the same conditioning view for all methods.
Single-view to material
Text to material
Method
CLIP-FID ( ↓ )
CMMD ( ↓ )
LPIPS ( ↓ )
Method
CLIP-FID ( ↓ )
CMMD ( ↓ )
LPIPS ( ↓ )
Hunyuan3D 2.1
3.419
0.0437
0.0520
VideoMatGen
2.973
0.0236
0.0527
VideoMatGen
4.725
0.0330
0.0639
Trellis.2
2.227
0.0184
0.0395
VideoMat
4.376
0.0247
0.0652
Ours
1.520
0.0081
0.0325
Ours
3.338
0.0230
0.0542
Table 2: Material generation. Left: single-view to material generation. Right: text-guided material generation. All metrics evaluated on 32 scenes × 8 views × 8 probes.
Figure 7: TEXGen . TEXGen ( Yu et al., 2024 ) only produces albedo textures. Metrics for diffuse-only renderings from 32 scenes × 8 views × 8 probes from single-view guidance.
Figure 8: Ablation – Inference Scaling. With our inference augmentations, we can scale our results to 8K texture resolution and increase the number of input views from 17 to 117. This enables capturing high-frequency details and accurate specular predictions. We provide average metrics over the first four samples of the real dataset, showing preserved fidelity with slightly improved LPIPS.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1: Effect of CRF . We show the effect of the CRF adjustment on the LSRM ( Li et al., 2026b ) prediction of the Gargoyle sample from the DTC dataset ( Dong et al., 2025 ) . Even though the material patterns are mostly captured properly by the baseline, color is shifted due to unknown lighting condition and potential training bias. CRF adjustment helps to rule out this bias and provide a more fair comparison.
Real recon (136 img)
Synthetic recon (256 img)
Synthetic relight (2k img)
Geom.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Known
DiffPT
27.01
0.959
0.0308
21.99
0.908
0.0854
21.14
0.903
0.0911
(ref.)
Ours
22.66
0.930
0.0430
30.17
0.958
0.0327
29.89
0.954
0.0350
Recon.
LSRM
24.56
0.926
0.0445
20.08
0.848
0.1296
19.80
0.847
0.1330
(LSRM)
Ours (LSRM geo)
22.22
0.923
0.0496
23.61
0.884
0.0860
23.43
0.889
0.0880
Appendix
Table S1: Results Without CRF Adjustment . We provide quantitative results for Table 1 without adjusting the CRF. The tendency is the same, however some metrics become heavily biased due to the decomposition ambiguity.
Method
Resolution
2K recon
8K super-res
Projection
Combined time
Peak VRAM
17-view
2K
130 s
-
-
2 min 10 s
22 GiB
117-view
2K
2610 s
-
-
43 min 30 s
48 GiB
17-view
8K
130 s
267 s
10 s
6 min 47 s
22 GiB
117-view
8K
2610 s
279 s
10 s
48 min 19 s
48 GiB
Appendix
Table S2: Runtime cost . Runtime cost and Peak VRAM usage. Scores are averages over four examples from the DTC dataset, measured on an NVIDIA GB300 GPU.
Figure S2: Material components .
Method
Base color
Roughness
Metallicity
siPSNR ( ↑ )
PSNR ( ↑ )
PSNR ( ↑ )
PSNR ( ↑ )
Ours
29.19
27.00
24.32
18.24
DiffPT ( Hasselgren et al., 2022 )
25.08
20.21
17.98
13.23
LSRM ( Li et al., 2026b )
23.17
15.46
13.18
14.73
Appendix
Table S3: Material components . We compute metrics on g-buffer renderings for 32 scenes × 17 views from multi-view guidance in our BlenderVault test set. We report scale-invariant PSNR for the base color renderings, and standard PSNR scores for base color, roughness, and metallicity g-buffer renderings.
Figure S3: Material generations from casually captured photographs, using 17 views from the DTC dataset Dong et al. (2025) . We expect known camera poses and geometry in this test. Top : Input views (two selected views per example). Bottom : Our extracted materials (two seeds per example) re-rendered in novel lighting in Blender.
Figure S4: Generated basecolor textures from single-view guidance. Overall, our method generates more consistent albedo maps than TEXGen, demodulates lighting from the conditional view (TEXGen uses a demodulated conditional view), and outputs full PBR materials (not shown here).
Figure S5: Reconstruction - Real-World. Material reconstructions from multi-view observations (photographs with known poses) on three examples from the DTC dataset ( Dong et al., 2025 ) . For optimization-based differentiable path tracing (DiffPT) and our method, we leverage known geometry, while LSRM jointly reconstructs geometry and materials.
Basecolor
(Rgh, Met)
Method
PSNR ( ↑ )
LPIPS ( ↓ )
PSNR ( ↑ )
LPIPS ( ↓ )
Our
34.7
0.061
38.0
0.025
PBR-SR ( Chen et al., 2025a )
30.9
0.151
40.6
0.030
ESRGAN ( Wang et al., 2018 )
32.6
0.146
37.9
0.041
Appendix
Table S4: Texture upscaler. Results for 4× texture upscaling comparing our method with popular alternatives. Results are quality metric averages for all textures in our synthetic dataset (32 objects).
Figure S6: Attention visualization . We visualize the attention activations for a query point (red circle) for three variants: Top : The RoPE of Wan 2.1, combining pixel xy -coordinates, puv , and the frame id fid . Middle : The Wan 2.1 RoPE and G-buffer world position and normal guides, Gbuf . Bottom : Our 3D-aware RoPE + Gbuf . We show the attention activations for five layers. We note that our 3D-aware RoPE has stronger activations at similar 3D locations.
Figure S7: Attention visualization. We show the attention activations for all layers for two query points on the lion king DTC example. The RoPE of Wan 2.1 combines pixel xy -coordinates, puv , and the frame id fid . We add a G-buffer wpos and normals, Gbuf , as conditions (middle column). In the rightmost column, we show our 3D-aware RoPE + Gbuf . Please zoom on the PDF image to see details. For reference, the left column shows the world space distance between two points in texture space. Note that our embedding attends strongly to points closer in world space over all layers of the network.
RoPE
Inputs
Gbuf
CLIP-FID ↓
CMMD ↓
LPIPS ↓
Wan 2.1
puv+fid
✗
4.036
0.0405
0.1005
Wan 2.1
puv+fid
✓
3.677
0.0289
0.0971
3D-aware
pxyz+fid
✓
3.561
0.0237
0.0957
Appendix
Table S5: RoPE embedding ablation. We ablate our proposed 3D-aware rotary positional embedding against the standard Wan 2.1 RoPE on the text to material pipeline in Figure 2 of the main paper, evaluated on the same 32 scenes × 8 views × 8 probes with material textures generated at a resolution of 512×512 texels.
Figure S8: Stability. Impact of RoPE embedding on a single text to material example with varying random seed. As can be seen, the 3D-aware RoPE encoding better captures the global structure of the object, and produces more consistent materials. The small insets show the base color, roughness, and metallicity maps for each generation.
Figure S9: Generation - Text to material. Materials from our generative method, conditioned on a text prompt and the input geometry (the UV mask, world space positions, and surface normals). We compare against VideoMat ( Munkberg et al., 2025 ) and VideoMatGen ( Hasselgren et al., 2026 ) . Without any image guidance, our model still produces semantically meaningful materials for all examples.