Generating high-quality UV textures for 3D meshes remains challenging. Multi-view projection pipelines suffer from occlusion and view inconsistency, and recent methods that generate textures directly in UV space still rely on auxiliary modules to supply 3D information, leaving the attention mechanism tied to UV-grid positions rather than to the underlying surface geometry. This mismatch limits coherence across seams and disconnected UV islands. We propose DirectUV, an image-conditioned UV texture diffusion framework that operates in the latent UV space of a pretrained image VAE, in which a Diffusion Transformer denoises the UV latent given a single input image and a coarse UV map. At its core, Surface-Aware Positional Encoding (SAPE) replaces the standard 2D-grid positional encoding with encodings derived from per-token 3D surface coordinates obtained via UV-to-surface correspondence. As positional encodings define the distance metric used by attention, SAPE enables tokens to interact according to 3D positional proximity derived from surface correspondence rather than UV-grid distance, restoring coherence across seams and disconnected islands. A multi-level extension further assigns different attention heads to progressively finer subdivisions of the same latent UV patch, allowing the model to reason about surface structure at multiple granularities. Experiments show that DirectUV produces sharper and more globally consistent textures than other baselines, with the largest improvements in occluded and view-unseen regions where projection-based methods leave gaps or stretched textures.
Figures & tables
Figure 1 : Overview of DirectUV. Top (training). The reference image is encoded by CLIP [ 24 ] into a pooled feature that drives AdaLN modulation, and by DINOv2 [ 25 ] into patch tokens that feed the joint-attention stream. A coarse UV map is encoded by the frozen Flux VAE and channel-concatenated with the noisy UV latent to form the DiT input. The CCM supplies per-head RoPE coordinates for Multi-Level SAPE. The model is trained with the rectified flow-matching objective. Bottom (inference). The reference image drives both an off-the-shelf multi-view generator, whose outputs are projected to a coarse UV map, and the CLIP/DINOv2 encoders. Under this conditioning, N Euler steps denoise z1 to z0 , which is decoded by the frozen Flux VAE and wrapped onto the mesh as the final texture.
Figure 3 : Qualitative comparison with state-of-the-art texture generation methods. (i) Finer details: our results preserve sharper high-frequency patterns, e.g., the line strokes in the number block. (ii) Occluded regions: for heavily occluded interiors such as the well and the trash bin, our method produces coherent and plausible content, while others yield blurry or incomplete textures. (iii) Robustness to multi-view inconsistency: by denoising directly in latent UV space, our method avoids being misled by conflicting views. In the fan case, UniTEX incorrectly transfers the back-cover texture onto the front blades, whereas ours preserves the correct appearance.
Method
PSNR ↑
SSIM ↑
FID ↓
KID ↓
TEXGen [ 10 ]
24.45
0.9417
44.09
34.32
FlexPainter [ 6 ]
23.21
0.9463
47.04
39.25
Hunyuan3D-2.1 [ 13 ]
24.74
0.9450
39.872
33.34
UniTEX [ 11 ]
24.23
0.9504
42.11
39.21
Ours
25.13
0.9527
39.17
30.03
Table 1 : Quantitative comparison on the GSO dataset. Best results are shown in bold, and second-best results are underlined.
Figure 4 : Qualitative ablation of Surface-Aware Positional Encoding. Without SAPE , texels that are adjacent on the 3D surface but split into different UV islands cannot interact properly, leading to severe appearance divergence in some regions, as seen in the chocobo case where the black texture from the legs is incorrectly generated on the neck. Removing only the multi-level design preserves cross-island coherence but loses spatial precision, yielding blurry textures .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1 : The keyboard underside appears white in the coarse UV (center) despite being covered by the bottom-view projection, but is correctly recovered as dark metallic casing in the final texture (right), consistent with the reference image. DirectUV’s learned prior overrides the inaccurate coarse UV values rather than copying them.
Figure 2 : Texture generation results under varying coarse UV completeness at inference time ( 0 , 1 , 2 , and 4 projected views). The reference image and mesh are identical across all cases.
Figure 3 : Texture generation on meshes produced by image-to-3D shape generators, demonstrating DirectUV in a complete image-to-textured-mesh pipeline. Each block shows the input reference image alongside the generated texture rendered from multiple viewpoints.
Figure 4 : Failure cases on heavily fragmented UV layouts. The UV map (top row) is split into a large number of small islands and slivers with substantial unused border space, leading to blurred or color-leaked regions in the wrapped texture (bottom row).
UV parameterization is a fundamental step in 3D content creation, yet producing production-ready UV layouts remains challenging due to the gap between geometric distortion objectives and the stylistic preferences of professional artists. While classical methods optimize handcrafted energy functions, artist-authored UVs exhibit structural patterns such as straightened seams, axis-aligned islands, and flexible interior deformation, properties that are difficult to explicitly formulate. In this work, we present DreamUV, an end-to-end learning framework that formulates UV unwrapping as a generative Flow Matching problem. Rather than predicting a single optimal parameterization, DreamUV learns a mesh-conditioned transport process that maps noise samples to a distribution of artist-like UV layouts. To reflect real-world authoring practices, we introduce a boundary-aware training strategy that prioritizes seam geometry, and a Model-in-the-Loop Finetuning(MITL) scheme that explicitly accounts for discretization errors during sampling and stabilizes transport dynamics under heterogeneous supervision. We evaluate DreamUV on a large-scale dataset of professionally authored UV layouts. Experiments demonstrate that our method produces significantly straighter boundaries and tighter axis-aligned islands than both classical and learning-based baselines, while maintaining competitive distortion metrics. Qualitative results and a user study with professional artists further confirm that DreamUV generates UV layouts that are not only valid, but aligned with practical production requirements.
Quanyuan Ruan, Jiabao Lei, Xingyi Du +1
South China University of Technology · School of Data Science, The Chinese University of Hong Kong, Shenzhen · Lightspeed
We propose OMGTex, an end-to-end diffusion-based framework for reconstructing high-quality and editable facial UV textures from multi-style facial images. Existing texture reconstruction methods face two major limitations: (1) Fragility due to reliance on 3D geometry priors, which are difficult to estimate accurately, especially under facial occlusions or in stylized domains; and (2) A lack of semantic disentanglement, inhibiting region-specific texture editing and style transfer. Our work addresses both challenges simultaneously. Our core innovation is a geometry-free pipeline that directly maps a 2D face image to its corresponding editable UV texture. We introduce two key techniques: First, to address the challenge of UV misalignment common in diffusion generation, we introduce a gradient-guided refinement strategy at inference time, which explicitly corrects structural consistency. Second, we leverage the inherent semantic distribution capability of diffusion models and design a novel training paradigm to enhance this tendency, enabling semantic-aware editing of facial texture. Furthermore, to address the data scarcity in multi-style texture reconstruction, we construct CANVAS, the first comprehensive paired texture reconstruction dataset covering realistic and diverse stylized domains. To the best of our knowledge, OMGTex is the first geometry-free inference framework that achieves robust, style-consistent, and editable facial texture reconstruction across diverse domains. Our method achieves state-of-the-art performance on multiple facial texture benchmarks.
Zitong Xiao, Yuda Qiu, Zisheng Ye +1
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen · Guangdong Provincial Key Laboratory of Future Networks of Intelligence · FNii-Shenzhen
Reconstructing high-fidelity, relightable 3D avatars from a single in-the-wild image is a challenging ill-posed problem, primarily hindered by the scarcity of high-quality PBR data and the complexity of disentangling illumination from intrinsic materials. In this paper, we present a data-efficient framework that leverages the robust priors of a unified pre-trained diffusion backbone to sequentially address texture completion, delighting, and material decomposition. Unlike existing methods that rely on fragmented pipelines or extensive proprietary datasets, we utilize cascaded Low-Rank Adaptations (LoRAs) to adapt the strong generative prior of the diffusion model for each sub-task in UV space. Specifically, we first employ an Inpainting LoRA to complete missing UV textures caused by occlusion, leveraging the model's semantic understanding to generate semantically and photometrically coherent details. Subsequently, a Light-Homogenization LoRA and a novel Cross-Intrinsic Attention mechanism are introduced to remove baked-in lighting and collaboratively synthesize pixel-aligned PBR maps (Albedo, Normal, Roughness, Specular, and Displacement). To ensure physical plausibility, we impose a UV-space differentiable BRDF shading loss during the decomposition stage, forcing the generative process to adhere to the rendering equation without the artifacts typical of rasterization-based supervision. Extensive experiments demonstrate that our method, trained on fewer than 100 real 3D scans, generates comprehensive, 4K-resolution PBR assets with superior realism and generalization compared to state-of-the-art methods, and all training code and model weights will be released upon acceptance.