Recent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens. However, compared with two-stage methods that provide explicit positional guidance, these models must implicitly infer token positions throughout denoising, limiting their generation quality. We observe that, despite the absence of explicit positional conditioning, VecSet tokens retain recoverable spatial correspondences. Building on this observation, we propose Position Forcing, a position-based self-conditioning framework. During denoising, Position Forcing recovers token positions from the current clean latent estimate, quantizes them at progressively finer resolutions according to the denoising stage, and feeds the resulting positional encodings back into the diffusion Transformer. This progressively refined positional feedback provides spatial guidance at a granularity appropriate to each denoising stage, guiding shape generation along a coarse-to-fine trajectory and substantially improving generation quality without a separate position generation stage. Experiments demonstrate that Position Forcing achieves strong performance among single-stage 3D generative methods and outperforms several competitive multi-stage approaches.
Figures & tables
Figure 1 : Our observation and positional conditioning designs. (I) Latents can be decoded into shape geometry and token positions. (II) (a) VecSet uses no explicit positional conditioning; (b) VoxSet and two-stage pipelines use positions established beforehand. (c) Our naive design recovers positions from the current noisy state. (d) Position Forcing recovers positions from the predicted clean state and applies progressive quantization to provide coarse-to-fine spatial guidance.
Figure 2 : Method overview. (a) Position VAE learns latent tokens that preserve geometry–position correspondences, enabling both shape geometry and token-associated spatial positions to be decoded. (b) During denoising, Position-Forced DiT recovers token positions from the predicted clean latents and injects them as positional guidance into subsequent denoising steps through RoPE. (c) The recovered positions are progressively quantized from coarse to fine as denoising proceeds.
Noise Level (%)
Training Strategy
0
10
20
30
Frozen VAE
67.4
51.3
23.0
6.7
Joint Finetuning
99.9
99.7
92.5
67.3
Table 1 : Effect of joint training on position recovery. Accuracy (%) measures the fraction of predicted positions that fall into the same cells as their corresponding query positions on a 1283 grid.
Method
Latent Size
CD ↓
F1 ↑
TripoSG
64×4096
17.96
92.33
[ Li et al., 2025b ]
64×8192
16.61
94.19
Hunyuan3D-2.1
64×4096
9.08
89.43
[ Hunyuan3D et al., 2025 ]
64×8192
8.28
90.99
64×20480
7.62
92.06
Position Forcing
64×4096
8.28
90.74
Table 2 : Quantitative comparisons of geometry reconstruction.
Model
Single Stage
ULIP-T ↑
ULIP-I ↑
Uni3D-T ↑
Uni3D-I ↑
TRELLIS [ Xiang et al., 2025 ]
✗
0.076
0.126
0.249
0.311
TRELLIS 2 [ Xiang et al., 2026 ]
✗
0.077
0.124
0.245
0.317
Direct3D-S2 [ Wu et al., 2026b ]
✗
0.074
0.122
0.247
0.314
Hi3DGen [ Ye et al., 2025 ]
✗
0.065
0.113
0.252
0.301
CraftsMan 1.5 [ Li et al., 2024 ]
✓
0.074
0.129
0.237
0.298
UniLat3D [ Wu et al., 2026a ]
✓
0.071
0.119
0.252
0.307
Table 3 : Comparison with existing 3D generation methods. Best results are highlighted in bold.
Figure 3 : Qualitative comparison of image-conditioned 3D generation. The input images are shown in the top row. Compared with existing methods, our approach produces more faithful 3D geometry with coherent global structures and fine-grained details across diverse object categories. Methods marked with * adopt a two-stage generation framework.
Figure 4 : Qualitative and quantitative ablation of Position Forcing. Columns (a)–(g) share the same configurations across qualitative and quantitative results. Query, Pred, and x0 denote ground-truth query positions, predicted positions, and positions recovered from the predicted clean latent, respectively. Prog. denotes progressive position quantization. Configuration (g) uses the Stage-I VAE instead of the jointly trained VAE. Bold and underlining denote the best and second-best results.
Figure 5 : Positional consistency during sampling. We compare cell accuracy along the sampling trajectory for positions recovered from predicted clean latents ( P0∣t ) and noisy latents ( Pt ). Accuracy is the percentage of tokens whose intermediate and final decoded positions fall into the same cell, using progressive resolution R(t) or fixed resolution Rmax for both.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Generation trajectory visualization. We use a fixed 50-step sampling schedule and visualize the predicted clean sample x^0 at different generation steps. The intermediate predictions show that the model progressively forms the global structure in the early steps and then refines local details in later steps, eventually converging to the final output.
Figure 7 : Additional generation results. Generated geometries are shown with their corresponding input images inset at the bottom right.
Figure 8 : Multi-view qualitative comparison of image-conditioned 3D generation. Position Forcing produces smoother surfaces and more geometrically plausible details, particularly in rear views, with fewer surface artifacts and structural distortions. Methods marked with * adopt a two-stage generation framework.
Autoregressive multimodal large language models (MLLMs) enable 3D generation but struggle to scale to high-resolution shapes due to inadequate 3D tokenizations. Compact set-based representations discard deterministic spatial ordering, leading to ambiguous sequence prediction, while uniform or octree-based voxel grids preserve ordering at the cost of severe redundancy and excessively long sequences. This structural trade-off limits stable and efficient autoregressive 3D generation. We present SuperVoxelGPT, a representation-first framework that resolves this tension through adaptive and deterministically ordered supervoxel tokenization. Given a prompt, we first predict a coarse geometric saliency distribution and construct a shape-adaptive supervoxel partition using saliency-guided centroidal Voronoi tessellation, allocating fine-grained cells to complex regions and larger cells to smooth regions. Conditioned on this prompt and ordered supervoxel layout, we introduce a SuperVoxelVAE and fine-tune a pretrained MLLM to autoregressively generate supervoxel tokens. Experiments using Trellis-500K data show that SuperVoxelGPT reduces token sequence length to 12.8% of uniform voxel tokenization while achieving state-of-the-art generation quality and an average 10x speedup over prior methods.
Yuan Li, Congyi Zhang, Xifeng Gao +1
University of Texas at Dallas, USA · LightSpeed, USA
Recent image-to-3D generation models built on flow-matching diffusion Transformers (DiT) can produce high-fidelity meshes, yet their post-training strategy remains largely unexplored. There exist several critical bottlenecks in reinforcement learning: the inherent difficulty of defining comprehensive rewards for 3D geometric quality, and the gradient interference that arises when jointly optimizing heterogeneous objectives. Inspired by the practicability of on-policy distillation (OPD) in large language models and image generation, we propose \textbf{Flow3D-OPD}, a two-stage post-training framework that introduces multi-teacher distillation into 3D geometry generation. In the first stage, we utilize the semi-policy to enhance the foundational capability of the pretrained model and then design an agentic verifier for 3D geometric quality evaluation. Based on the verifier, we could cultivate domain-specialized teacher models via direct preference optimization (DPO). In the second stage, we consolidate heterogeneous expertise into a unified student model through on-policy distillation with hard task-routing sampling and gradient accumulation, which could mitigate the gradient interference in joint optimization. Without relying on elaborate modifications, our straightforward yet effective design achieves consistent improvements across all geometric quality dimensions and surpasses all teacher models in the average metric. Extensive experiments demonstrate that our approach provides an effective paradigm for reinforcement learning in 3D generation.
Most recent advances in 3D generative modeling rely on diffusion or flow-matching formulations. We instead explore a fully autoregressive alternative and introduce GaussianGPT, a transformer-based model that directly generates 3D Gaussians via next-token prediction, thus facilitating full 3D scene generation. We first compress Gaussian primitives into a discrete latent grid using a sparse 3D convolutional autoencoder with vector quantization. The resulting tokens are serialized and modeled using a causal transformer with 3D rotary positional embedding, enabling sequential generation of spatial structure and appearance. Unlike diffusion-based methods that refine scenes holistically, our formulation constructs scenes step-by-step, naturally supporting completion, outpainting, controllable sampling via temperature, and flexible generation horizons. This formulation leverages the compositional inductive biases and scalability of autoregressive modeling while operating on explicit representations compatible with modern neural rendering pipelines, positioning autoregressive transformers as a complementary paradigm for controllable and context-aware 3D generation.
Nicolas von Lützow, Barbara Rössle, Katharina Schmid +1