DiDE:Direct Injection with Color-Texture DEcoupling for 3D Stylization
Organizations: Computer Vision Center, Universitat Autònoma de Barcelona, Spain · Mohamed bin Zayed University of Artificial Intelligence, United Arab Emirates · Jilin University, China · City University of Hong Kong (Dongguan), China
Abstract
Recent advances in rectified flow-based image-to-3D generative models have enabled high-fidelity 3D asset generation. Building on this, a growing line of work has exploited these strong 3D priors for training-free stylization, transferring visual attributes from a reference image onto a generated 3D asset. However, existing methods enforce an all-or-nothing paradigm: color and texture are transferred jointly, with no mechanism to control them independently -- a limitation we formalize as Disentangled 3D Stylization(Disen3D). To address this, we propose DiDE, the first training-free framework for Disen3D. Key to our approach is the observation that the structured latent space of image-to-3D models is overcomplete with respect to texture: texture information occupies only a small subset of the style-significant channels, leaving a free subspace available for independent color encoding. DiDE exploits this via a channel partition mechanism that processes a content image, a texture reference, and a color reference through dedicated branches and composes both style signals interference-free at every self-attention layer, preserving content geometry throughout. Experiments on Disen3D-Bench, our newly collected multi-reference benchmark, show that DiDE consistently outperforms 2D and 3D stylization baselines in color fidelity, texture transfer, and content preservation.
Figures & tables
| Method | CLIP-I | Geometry | Color | Texture | |||||
|---|---|---|---|---|---|---|---|---|---|
| CD | HD | MS-SWD | C-Hist | VLM | LPIPS | SSIM | VLM | ||
| SADis Qin et al. (2025) | 0.73 | 18.17 | 41.59 | 18.38 | 0.99 | 37.17 | 0.65 | 116.10 | 18.00 |
| StyleSculptor Qu et al. (2025) | 0.74 | 17.91 | 8.38 | 21.26 | 0.89 | 35.44 | 0.69 | 124.61 | 23.22 |
| MorphAny3D Sun et al. (2026) | 0.71 | 198.52 | 44.50 | 18.31 | 0.91 | 40.05 | 0.65 | 117.21 | 19.72 |
| DiDE (Ours) | 0.75 | 15.38 | 7.87 | 20.40 | 0.87 | 70.94 | 0.68 | 125.19 | 30.72 |
| Color | Texture | |||
|---|---|---|---|---|
| MS-SWD | C-Hist | LPIPS | SSIM | |
| 0.5 | 20.24 | 0.92 | 0.68 | 83.36 |
| 1.0 | 19.47 | 0.89 | 0.69 | 83.34 |
| 1.2 | 19.25 | 0.89 | 0.69 | 83.45 |
| 1.5 | 19.32 | 0.90 | 0.70 | 83.30 |
| 2.0 | 19.38 | 0.90 | 0.70 | 83.40 |
| Method | Geometry | Texture | Color |
|---|---|---|---|
| MorphAny3D | 4.4% | 1.2% | 2.3% |
| SADiS | 12.1% | 3.2% | 8.7% |
| StyleSculptor | 31.8% | 37.4% | 16.6% |
| Ours | 51.7% | 58.2% | 72.4% |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.