cs.CVSep 29, 2026

NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generation

Authors: Jiawei Zhang, Shuhao Liu, Rong Huang, Yuancheng Li, Zhihui Li, Xiaojun Chang, Changlin Li

Organizations: North China Electric Power University · University of Science and Technology of China · Stanford University

Abstract

One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail to fully exploit the representational capacity of the tokenizer, resulting in suboptimal performance in both image reconstruction and generation. In this work, we introduce NesTok, a nested self-alignment framework tailored to dynamic visual tokenizers. NesTok introduces cross-length training, which jointly optimizes reconstruction across token lengths while using the full-length sequence to guide shorter counterparts, enabling shorter token sequences to approach the reconstruction quality of full-length sequences. On ImageNet, NesTok improves substantially over standard training and achieves an rFID score of 0.98. On downstream image generation, it achieves the state-of-the-art gFID score of 1.46 on ImageNet 256×\times256 among existing variable-length autoregressive image generation methods. Code will be available at https://github.com/Jiawei804/NesTok.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer

    May 1, 2026Wenda Chu, Bingliang Zhang, Jiaqi Han +4Autoregressive Image GenerationVisual Tokenizers

  2. Balancing Image Compression and Generation with Bootstrapped Tokenization

    Jun 4, 2026Haozhe Chi, Jinghan Li, Hao Jiang +4Visual TokenizersAutoregressive Image Generation

  3. ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters

    May 6, 2026Philippe Hansen-Estruch, Jiahui Chen, Vivek Ramanujan +9Visual TokenizersVision Transformer