cs.CVOct 7, 2026

Position Forcing: Self-Conditioning 3D Generation

Authors: Ziheng Ouyang, Zeqiang Lai, Jiarui Chen, Jiangshan Wang, Yuhao Wan, Jingbo Gong, Xiangyu Yue, Hengshuang Zhao, +2 more

Organizations: VCIP, Nankai University · Tencent Hunyuan · MMLab, CUHK · Fudan University · Shanghai Innovation Institute · HKU

Abstract

Recent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens. However, compared with two-stage methods that provide explicit positional guidance, these models must implicitly infer token positions throughout denoising, limiting their generation quality. We observe that, despite the absence of explicit positional conditioning, VecSet tokens retain recoverable spatial correspondences. Building on this observation, we propose Position Forcing, a position-based self-conditioning framework. During denoising, Position Forcing recovers token positions from the current clean latent estimate, quantizes them at progressively finer resolutions according to the denoising stage, and feeds the resulting positional encodings back into the diffusion Transformer. This progressively refined positional feedback provides spatial guidance at a granularity appropriate to each denoising stage, guiding shape generation along a coarse-to-fine trajectory and substantially improving generation quality without a separate position generation stage. Experiments demonstrate that Position Forcing achieves strong performance among single-stage 3D generative methods and outperforms several competitive multi-stage approaches.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation

    May 28, 2026Yuan Li, Congyi Zhang, Xifeng Gao +1VoxelAutoregressive Generation

  2. GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation

    Mar 27, 2026Nicolas von Lützow, Barbara Rössle, Katharina Schmid +13D Scene Generation3D Gaussian