DiDE:Direct Injection with Color-Texture DEcoupling for 3D Stylization
Authors: Tao Wu, Alexandra Gomez-Villa, Senmao Li, Yaxing Wang, Joost van de Weijer, Kai Wang
Organizations: Computer Vision Center, Universitat Autònoma de Barcelona, Spain · Mohamed bin Zayed University of Artificial Intelligence, United Arab Emirates · Jilin University, China · City University of Hong Kong (Dongguan), China
Recent advances in rectified flow-based image-to-3D generative models have enabled high-fidelity 3D asset generation. Building on this, a growing line of work has exploited these strong 3D priors for training-free stylization, transferring visual attributes from a reference image onto a generated 3D asset. However, existing methods enforce an all-or-nothing paradigm: color and texture are transferred jointly, with no mechanism to control them independently -- a limitation we formalize as Disentangled 3D Stylization(Disen3D). To address this, we propose DiDE, the first training-free framework for Disen3D. Key to our approach is the observation that the structured latent space of image-to-3D models is overcomplete with respect to texture: texture information occupies only a small subset of the style-significant channels, leaving a free subspace available for independent color encoding. DiDE exploits this via a channel partition mechanism that processes a content image, a texture reference, and a color reference through dedicated branches and composes both style signals interference-free at every self-attention layer, preserving content geometry throughout. Experiments on Disen3D-Bench, our newly collected multi-reference benchmark, show that DiDE consistently outperforms 2D and 3D stylization baselines in color fidelity, texture transfer, and content preservation.
Figures & tables
Figure 1 : Our training-free framework DiDE enables the generation of stylized 3D assets, with flexible control over the independent modulation of color or texture.
Figure 2
Figure 4 : Overview of DiDE . Left: Our framework takes three inputs — a content image, a texture reference, and a color reference — processed through three independent branches: the content branch ( blue ), the texture branch ( red ), and the color branch ( green ). Right: At each DiT block, the texture and color branches produce independent self-attention outputs. A channel selection operator S partitions the feature channels into texture-allocated (high-variance) and color-allocated (low-variance) subsets, and the two attention outputs are composed via the binary mask Mc into a single fused representation SAfusion , which is then injected into the residual stream of the content branch through its gated update.
Figure 5 : Comparison with existing 3D stylization methods and adapted baselines for the Disen3D problem. The first two rows show results with a single style reference, while the remaining rows show results with disentangled texture and color references.
Method
CLIP-I ↑
Geometry
Color
Texture
CD ×102↓
HD ×102↓
MS-SWD ↓
C-Hist ↓
VLM ↑
LPIPS ↓
SSIM ×103↑
VLM ↑
SADis Qin et al. (2025)
0.73
18.17
41.59
18.38
0.99
37.17
0.65
116.10
18.00
StyleSculptor Qu et al. (2025)
0.74
17.91
8.38
21.26
0.89
35.44
0.69
124.61
23.22
MorphAny3D Sun et al. (2026)
0.71
198.52
44.50
18.31
0.91
40.05
0.65
117.21
19.72
DiDE (Ours)
0.75
15.38
7.87
20.40
0.87
70.94
0.68
125.19
30.72
Table 1 : Quantitative comparisons with existing style transfer methods. The best and second-best results are highlighted in bold and underlined , respectively. CD and HD denote Chamfer Distance and Hausdorff Distance, respectively.
Figure 6 : Ablation on grayscale conversion and Voronoi tessellation preprocessing.
Figure 7 : DiDE applied to the TRELLIS.2 backbone, showing generalization across model variants.
Figure 8 : Ablation on channel selection variance source: texture vs. color branch.
λ
Color
Texture
MS-SWD ↓
C-Hist ↓
LPIPS ↓
SSIM ×103↑
0.5
20.24
0.92
0.68
83.36
1.0
19.47
0.89
0.69
83.34
1.2
19.25
0.89
0.69
83.45
1.5
19.32
0.90
0.70
83.30
2.0
19.38
0.90
0.70
83.40
Table 2 : Ablation on color branch weight λ .
Method
Geometry ↑
Texture ↑
Color ↑
MorphAny3D
4.4%
1.2%
2.3%
SADiS
12.1%
3.2%
8.7%
StyleSculptor
31.8%
37.4%
16.6%
Ours
51.7%
58.2%
72.4%
Table 3: User study results on geometry, texture, and color preference. Higher is better.
Figure 9 : Failure case under texture-reference brightness shifts.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10 : Ablation on λ .
Figure 11 : Ablation on k .
Figure 12 : Additional qualitative comparisons
Figure 13 : DINOv2 cosine similarity between rendered outputs and texture references as a function of the number of PCA components retained from TRELLIS texture-significant features. Similarity stabilizes after as few as 10 components, confirming the low intrinsic dimensionality of texture within the feature space.
With the growth of gaming, animation, and virtual reality industries, the demand for efficient generation of stylized 3D assets is rapidly increasing. However, existing approaches still struggle to jointly preserve style fidelity, geometric consistency, and generation efficiency, as most of them still rely on indirect 2D-to-3D stylization pipelines. This motivates a native 3D stylization framework that can explicitly disentangle style from geometry while remaining efficient. To this end, we propose DreamStyle3D, an efficient framework for stylized 3D asset generation built on a Decoupled Dual Cross-Attention mechanism. Our method explicitly separates geometric and stylistic features to enable efficient style injection while preserving structural consistency, and further adopts a lightweight training strategy to enhance style consistency and model generalization. In addition, we build an automated data pipeline and construct a dataset of about 15K content-style-stylized triplets for training and evaluation. Extensive experiments demonstrate that our DreamStyle3D can generate high-fidelity, geometrically consistent stylized 3D assets within 10 seconds, substantially improving efficiency while maintaining superior style quality and offering a new solution for 3D content creation. The project is available at https://github.com/NK-JittorCV/nk-3D/tree/main/models/DreamStyle3D.
Kai Wang, Ziheng Ouyang, Xuying Zhang +2
VCIP, Nankai University Tianjin, China · Zhongguancun Academy Beijing, China · Joy Future Academy, JD Beijing, China +1
3D asset generation plays a pivotal role in fields such as gaming and virtual reality, enabling the rapid synthesis of high-fidelity 3D objects from a single or multiple images. Building on this capability, enabling style-controllable generation naturally emerges as an important and desirable direction. However, existing approaches typically rely on style images that lie within or are similar to the training distribution of 3D generation models. When presented with out-of-distribution (OOD) styles, their performance degrades significantly or even fails. To address this limitation, we introduce \textbf{DiLAST}: 2D Diffusion-based Latent Awakening for 3D Style Transfer. Specifically, we leverage a pretrained 2D diffusion model as a teacher to provide rich and generalizable style priors. By aligning rendered views with the target style under diffusion-based guidance, our method optimizes the structured 3D latent representations for stylization. We observe that this limitation stems not from insufficient model capacity, but from the underutilization of structured 3D latents, which are inherently expressive. Despite being trained on comparatively limited data, 3D generation models can leverage 2D diffusion guidance to steer denoising toward specific directions in latent space, thereby producing diverse, OOD styles. Extensive experiments across diverse data and multiple 3D generation backbones demonstrate the effectiveness and plug-and-play nature of our approach.
Text-guided style editing of 3D assets is essential for adapting existing objects to diverse visual aesthetics in digital content creation. Despite rapid progress in 3D shape modeling, faithfully stylizing an existing asset remains challenging when the desired stylization involves fine-grained structural ornamentation, which requires the model to preserve the source geometry and object identity, while coherently integrating new style-specific details. We propose \textbf{OrnaStyler}, a zero-shot framework for text-guided ornament-aware 3D stylization. Built upon rectified flow-based generative modeling, OrnaStyler introduces an inversion-guided editing strategy that recovers content-aware latent representations at both geometry and appearance levels in a staged manner to facilitate faithful editing. Our core idea is to explicitly model the spatial configuration of stylistic elements, thereby mitigating the fundamental tension between content preservation and style expression in the voxel space. Specifically, at the geometry level, we manipulate voxel representations through flow inversion to synthesize ornament-enhanced structures while preserving the spatial identity of the source asset. Then, at the appearance level, we introduce an adjacency-aware feature inpainting mechanism to harmonize newly generated ornaments with the original content, yielding coherent geometry-appearance integration. Our approach operates solely in the inference phase and enables selective editing over geometric augmentation or appearance stylization. Extensive experiments on both generated and real-world 3D assets against prior methods demonstrate that OrnaStyler achieves state-of-the-art editing performance in terms of content preservation, style fidelity, and overall visual realism. Code is available at: https://github.com/tomohiro0427/OrnaStyler
Tomohiro Aizawa, Shigeru Kuriyama, Chunzhi Gu
University of Fukui, Japan · Toyohashi University of Technology, Japan · AI Lab, CyberAgent, Inc., Japan