High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. However, preserving fine detail across the physically based rendering (PBR) modalities needed for relighting remains challenging. To address this, we propose Luce, a 3D representation that unifies geometry and PBR materials within a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for albedo, metallic-roughness, and surface normals. A variational autoencoder compresses this representation into a unified material-aware latent space. A rectified-flow transformer generates this latent from a single image using multi-layer features from a pretrained image encoder that preserve both semantic context and fine spatial detail. The latent is then decoded into relightable PBR Gaussians and an optional textured mesh with a tangent-space normal map. On Toys4K, Luce achieves state-of-the-art single-image-to-3D generation, improving FID by 28% over the strongest baseline. We further evaluate Luce on a benchmark of AI-generated images depicting diverse subjects and materials, where it improves the CLIP image-alignment score over the best baseline (0.8519 vs. 0.8299). Luce generates relightable, geometrically accurate, and materially faithful assets that preserve fine details such as text, logos, and inscriptions.
Figures & tables
Figure 1: Our method Luce generates relightable 3D assets from a single image. Each asset consists of Gaussians encoding three modalities of a physically based rendering (PBR) material: albedo, metallic-roughness, and surface normals. These modalities enable standard PBR shading, rendering the Gaussians under novel illumination. From left to right, the columns show: the input condition image; the generated PBR modalities stacked vertically—albedo (top), metallic-roughness (middle, metallic in Red, roughness in Green channel), and surface normals (bottom); and the generated 3D asset rendered under three environment maps, previewed as small strips above each render (top row).
Figure 2: Legible text on generated 3D assets. Luce compared with LiTo ( Chang et al., 2026 ) , TRELLIS 2 ( Xiang et al., 2026 ) , and TRELLIS ( Xiang et al., 2025 ) on single-image-to-3D generation; each row shows the input condition image and one rendered view of each method’s generated 3D asset. LiTo generates assets aligned to the input view, whereas other methods generate in a canonical orientation. Luce keeps surface text and markings legible where baselines distort them.
Figure 3: Overview of Luce. (Top) Representation and SLatVAE. Given a 3D PBR asset, here a Toys4K ( Stojanov et al., 2021 ) sample, we render multiview images and fit per-modality Gaussian splats (albedo, metallic-roughness, normals) on a sparse voxel grid. A structured latent VAE (SLatVAE) encodes this representation into a compact, diffusible latent and decodes it back to PBR Gaussians. (Bottom) Generation pipeline. Given a single input image, a sparse-structure flow first predicts the sparse voxel layout. Conditioned on multi-layer DINOv2 features, SLatFlow then generates a PBR Gaussian latent at each active voxel. The structured latent decodes into relightable PBR Gaussians and also serves as input to the mesh decoder, which yields textured meshes with tangent-space normal maps and a complete PBR material set.
Figure 4: Effect of multi-layer DINOv2 conditioning. We compare Luce trained with multi-layer DINOv2 features (layers 6, 12, 18, 24) against single-layer DINOv2 features (layer 24 only). Multi-layer conditioning preserves fine spatial detail from the condition image, including legible text and logos on the generated 3D surface. For each variant, the larger shaded render is shown next to a cascade of the per-modality decomposition (albedo, metallic-roughness, surface normals).
Method
Params (B)
Time (s)
Toys4K ( N=412 )
AI-generated images ( N=130 )
FID ↓
KID ↓
FID dino↓
KID dino↓
CLIP ↑
SigLIP2 ↑
ULIP ↑
Uni3D-L ↑
CLIP ↑
SigLIP2 ↑
ULIP ↑
Uni3D-L ↑
TRELLIS GS ( Xiang et al., 2025 )
1.70
4.78
30.75
0.212
0.109
0.0013
0.8898
0.9164
—
—
0.8299
0.8339
—
—
TRELLIS mesh ( Xiang et al., 2025 )
1.80
19.67
32.39
0.263
0.145
0.0039
0.8829
0.9092
0.1672
0.3747
0.8000
0.7958
0.1247
0.3278
3DTopia-XL ( Chen et al., 2025 )
1.02
39.61
83.23
2.726
0.542
0.0306
0.7644
0.7810
0.1418
0.2519
0.6481
0.6458
0.1108
0.2197
TRELLIS 2 ( Xiang et al., 2026 )
7.48
199.55
29.22
0.165
0.129
0.0022
0.8895
0.9161
0.1634
0.3635
0.8110
0.8166
0.1207
0.3280
LiTo ( Chang et al., 2026 )
1.84
3.75
29.76
0.208
0.128
0.0025
0.8909
0.9082
—
—
0.8234
0.8240
—
—
Table 1: Image-to-3D generation. We evaluate on two benchmarks: Toys4K ( N=412 , left) and our 130 AI-generated images ( N=130 , right). Toys4K reports FID/KID with Inception and DINO ( Oquab et al., 2024 ) backbones (KID and KID dino×100 ); both benchmarks report CLIP ( Radford et al., 2021 ) , SigLIP2 ( Tschannen et al., 2025 ) , ULIP ( Xue et al., 2024 ) , and Uni3D-L ( Zhou et al., 2024 ) . Time(s) is the mean warm wall-clock time per generation on one H100 (Appendix C ). Luce GS renders the decoded PBR Gaussians directly via deferred shading; Luce mesh rows render a textured mesh extracted from the same latent, with and without tangent-space normal map transfer. Best in bold , second-best underlined ; shaded rows are ours. “—” marks metrics that do not apply to that render path; ULIP and Uni3D-L are mesh-only.
Figure 5: Image-to-3D generation comparison. Luce preserves legible text and fine surface details on generated 3D assets. For each method, the larger shaded render is shown next to a cascade of the per-modality decomposition (albedo, metallic-roughness, surface normals) when available.
Condition Image PBR GS Render
Condition Image PBR GS Render
Figure 6: Additional generation examples. Luce on four diverse inputs: each cell shows, from left to right, the condition image, the per-modality PBR Gaussians (albedo, metallic-roughness, normals), and a shaded render under the environment map previewed above it.
Method
Color
Albedo
Metallic-Roughness
Normal
Params (B)
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
TRELLIS GS ( Xiang et al., 2025 )
27.0
0.912
0.124
—
—
—
—
—
—
—
—
—
0.17
TRELLIS mesh ( Xiang et al., 2025 )
26.3
0.907
0.118
—
—
—
—
—
—
26.9
0.914
0.113
0.26
3DTopia-XL ( Chen et al., 2025 )
23.9
0.877
0.188
25.4
0.893
0.191
22.6
0.910
0.191
23.3
0.876
0.165
0.03
TRELLIS 2 ( Xiang et al., 2026 )
34.5
0.961
0.074
40.7
0.976
0.051
42.7
0.986
0.038
32.1
0.940
0.093
1.66
LiTo ( Chang et al., 2026 )
31.9
0.944
0.110
—
—
—
—
—
—
27.6
0.912
0.116
0.29
Table 2: Reconstruction quality on a PBR subset of Toys4K ( N=338 ). We report per-modality PSNR, SSIM, and LPIPS for color rendering (combined appearance under fixed illumination), albedo (diffuse color), metallic-roughness, and normal maps (surface orientation). For Luce, the GS row renders the decoded PBR Gaussians directly via deferred shading (no mesh); the mesh rows render a textured mesh from the same latent, with and without baked tangent-space normals. Best in bold , second-best underlined ; shaded rows are ours. “—” indicates that the method does not produce that modality. Params is the autoencoder parameter count in billions.
Figure 7: Reconstruction quality on Toys4K. Columns are ground truth and methods; each block shows the modalities that each method reconstructs (albedo, metallic-roughness, normals). Luce preserves fine detail (text, highlights, texture); the green box marks the region magnified below.
Figure 8: Tangent-space normal map transfer. Luce bakes tangent-space normals from decoded normal Gaussians onto the extracted mesh, recovering fine surface detail at zero polygon-count cost. Each green box marks the region magnified below.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
SLatVAE
SLatFlow
Architecture
Transformer blocks
16
30
Model channels
1536
1536
I/O block channels
–
[1024]
Attention heads
12
12
MLP ratio
4
6
Appendix
Table 3: SLatVAE and SLatFlow hyperparameters. Architecture diagram in Fig. 9 .
Figure 9: SLatVAE architecture. The encoder compresses the sparse multimodal Gaussian cloud into a compact latent grid. The decoder reconstructs a denser set of Gaussians per voxel per modality, supervised through per-modality differentiable rendering. The same latent also serves as input to a mesh decoder for textured mesh extraction. Full hyperparameters in Table 3 .
Figure 10: PBR shaded rendering. The PBR modalities (albedo, metallic-roughness, normals) enable shaded rendering under novel environment maps ( Poly Haven, 2026 ) via standard PBR compositing (Eq. 1 ), without mesh extraction or UV unwrapping. We compare our deferred PBR renderer on the GS fit against the textured mesh rendered with Blender EEVEE ( Blender Online Community, 2024 ) as a reference; the environment maps are listed in Appendix D .
DINOv2 conditioning
Toys4K ( N=412 )
AI-generated images ( N=130 )
FID ↓
KID ↓
FID dino↓
KID dino↓
CLIP ↑
SigLIP2 ↑
CLIP ↑
SigLIP2 ↑
Luce, single-layer (layer 24)
25.21
0.101
0.121
0.0035
0.8977
0.9164
0.8081
0.8143
Luce, multi-layer (layers 6, 12, 18, 24)
20.99
0.033
0.100
0.0028
0.9062
0.9230
0.8519
0.8508
Appendix
Table 4: Multi-layer DINOv2 conditioning ablation. Single-layer (layer 24) versus multi-layer (layers 6, 12, 18, 24) image conditioning for SLatFlow; both rows use the GS render path, so ULIP and Uni3D-L (mesh-only metrics in our setup) are omitted. KID is reported ×100 ; FID dino and KID dino use a DINOv2 backbone. Both rows are Luce; the shaded row is our default.
Method
Sparse-structure stage
SLat
3DTopia-XL ( Chen et al., 2025 )
—
25 (DDIM)
TRELLIS ( Xiang et al., 2025 )
25
25
LiTo ( Chang et al., 2026 )
—
20 (Heun)
TRELLIS 2 ( Xiang et al., 2026 )
12
12 (shape) + 12 (texture)
Luce (ours)
12
10
Appendix
Table 5: Inference sampling steps per method. SLat = structured-latent stage. “—” indicates a single-stage method without a separate structure flow.
Condition Image PBR GS Render
Condition Image PBR GS Render
Appendix
Figure 11: Additional generation results. Additional Luce examples across diverse asset categories. Each cell shows, from left to right, the condition image, the per-modality PBR Gaussians (albedo, metallic-roughness, normals), and a shaded render under the environment map previewed above it.
Figure 12: Per-sample baseline comparison (1/8). Rows: methods (top label = condition); Luce rows are ours. Columns: four evaluation views (yaws 300∘ , 30∘ , 120∘ , 210∘ ) matching the views used to compute the CLIP and SigLIP2 scores in Table 1 . Per-row metrics show CLIP and SigLIP2, plus ULIP and Uni3D-L for the mesh rows, on the depicted sample. For methods with explicit material decomposition, three small thumbnails are stacked to the right of each shaded render, top to bottom: albedo, metallic-roughness (metallic=R, roughness=G, B unused), and surface normals. LiTo has only the normal thumbnail and TRELLIS GS none; TRELLIS mesh has only a base-color texture, so its uniform metallic-roughness thumbnail is the default material it is rendered with (Appendix D ).
Figure 13: Per-sample baseline comparison (2/8). Same layout as Figure 12 .
Figure 14: Per-sample baseline comparison (3/8). Same layout as Figure 12 .
Figure 15: Per-sample baseline comparison (4/8). Same layout as Figure 12 .
Figure 16: Per-sample baseline comparison (5/8). Same layout as Figure 12 .
Figure 17: Per-sample baseline comparison (6/8). Same layout as Figure 12 .
Figure 18: Per-sample baseline comparison (7/8). Same layout as Figure 12 .
Figure 19: Per-sample baseline comparison (8/8). Same layout as Figure 12 .
BNRist, Department of Computer Science and Technology, Tsinghua University, China · Tencent ARC Lab, China · Victoria University of Wellington, New Zealand