InStyle: Instant Appearance Stylization of 3D Shapes
Organizations: University of Amsterdam · Bosch Center for AI
Abstract
3D stylization is central to game development, virtual reality, and digital arts, where the demand for diverse assets calls for scalable methods that support fast, high-fidelity manipulation. Existing text-to-3D stylization methods typically distill from 2D image editors, requiring time-intensive per-asset optimization and exhibiting multi-view inconsistency due to the limitations of current text-to-image models, which makes them impractical for large-scale production. In this paper, we introduce GaussianBlender, a pioneering feed-forward framework for text-driven 3D stylization that performs edits instantly at inference. Our method learns structured, disentangled latent spaces with controlled information sharing for geometry and appearance from spatially-grouped 3D Gaussians. A latent diffusion model then applies text-conditioned edits on these learned representations. Comprehensive evaluations show that GaussianBlender not only delivers instant, high-fidelity, geometry-preserving, multi-view consistent stylization, but also surpasses methods that require per-instance test-time optimization - unlocking practical, democratized 3D stylization at scale.
Figures & tables
| Method | Text Alignment | Structural Consistency | Visual Quality |
| IN2N | 16.30% | 20.24% | 21.68% |
| IGS2GS | 14.67% | 17.48% | 20.21% |
| GaussCtrl | 20.10% | 15.84% | 17.48% |
| GaussianEditor | 15.21% | 4.37% | 6.01% |
| Ours | 33.69% | 42.05% | 34.60% |
| LightGaussian | Input 3DGS | w/o disent. | w/ disent. (Ours) | |
| PSNR | 35.2117 | 32.7136 | 33.6635 | 34.3299 |
| (1) w/o | 0.227 | 0.183 |
| (2) w/o feat. sharing | 0.221 | 0.171 |
| (3) Ours (full) | 0.260 | 0.223 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Evaluation Benchmark | “Make its colours look like rainbow”, “Make it Starry Night Van Gogh painting style”, “Make it wooden”, “Make it in cyberpunk style”, “Make it pop art neon duotone”, “Make it marble”, “Make it Barbie style”, “Make it golden”, “Make it botanical themed”, “Make it Halloween themed” |
| Training (single-prompt) | Each prompt in the evaluation benchmark is used to train a separate model. |
| Training (multi-prompt) | “Make its colours look like rainbow”, “Make it Starry Night Van Gogh painting style”, “Make it wooden”, “Make it in cyberpunk style”, “Make it pop art neon duotone”, “Make it marble”, “Make it Barbie style”, “Make it golden”, “Make it botanical themed”, “Make it Halloween themed”, “Make it Hello Kitty themed”, “Make it look like made of steel”, “Make it The Scream Edvard Munch painting”, “Make it in sunset palette”, “Make it Pirates of the Caribbean style”, “Make it look like a statue”, “Make it Garfield themed”, “Make it look like made of bronze”, “Make it Dalmatian style”, "Make it gothic style", "Make it ocean palette", "Make it graffiti street art style", “Make it watercolor painting style”, "Turn it into red", “Turn it into orange”, "Turn it into pink", "Turn it into blue", "Turn it into yellow", "Turn it into green", "Turn it into black" |
| Criterion | |
| Text Alignment | Which stylized result follows the prompt most closely? Consider whether the style change is too subtle (almost no visible change), or too drastic (over-saturated). |
| Structural Consistency | Which edited result best preserves the original asset’s 3D structure while applying only the requested appearance stylization? The object’s shape should not change after stylization (e.g., no changes to the object’s physical structure; for characters, the face, identity and features, must remain unchanged). |
| Visual Quality | Which edited result looks best overall? |
| Structure Dist. | |||
| MV-L1 | 0.251 | 0.210 | 0.0064 |
| DMD + MV-L1 (Ours) | 0.260 | 0.223 | 0.0050 |
| Structure Dist. | |||
| (1) | 0.260 | 0.223 | 0.0050 |
| (2) | 0.263 | 0.226 | 0.0089 |
| (3) | 0.268 | 0.232 | 0.0172 |