One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail to fully exploit the representational capacity of the tokenizer, resulting in suboptimal performance in both image reconstruction and generation. In this work, we introduce NesTok, a nested self-alignment framework tailored to dynamic visual tokenizers. NesTok introduces cross-length training, which jointly optimizes reconstruction across token lengths while using the full-length sequence to guide shorter counterparts, enabling shorter token sequences to approach the reconstruction quality of full-length sequences. On ImageNet, NesTok improves substantially over standard training and achieves an rFID score of 0.98. On downstream image generation, it achieves the state-of-the-art gFID score of 1.46 on ImageNet 256×256 among existing variable-length autoregressive image generation methods. Code will be available at https://github.com/Jiawei804/NesTok.
Figures & tables
Figure 1: Codebook statistics and AR generation performance on ImageNet-1K. (a): Normalized entropy of the code distribution at each token position, computed over the training dataset and averaged within groups of 16 positions. (b): Average codebook utilization across token positions and gFID. (c): Sampling efficiency with different methods. (d): Sampling efficiency for different token lengths with different methods.
Figure 2: Overview of the nested self-aligned training pipeline. We train a 1D variable-length tokenizer using two components: (1) cross-length sampling, which jointly optimizes reconstruction from a full-length sequence and a sampled shorter sequence at each iteration; and (2) a nested self-alignment loss, which aligns latent representations across the two lengths to promote cross-length consistency and improve reconstruction.
Method
Tokenizer
Generator
w/o guidance
w/ guidance
Type
#Params
#Tokens
rFID ↓
Type
#Params
gFID ↓
IS ↑
gFID ↓
IS ↑
2D Tokenization
DiT-XL/2 ( Peebles and Xie, 2023 )
SD-VAE
84M
256
0.62
Diff.
675M
9.62
121.5
2.27
278.2
REPA-XL/2 ( Yu et al., 2024d )
KL
84M
1024
0.62
Diff.
675M
5.90
157.8
1.42
305.7
Lightning-DiT-XL ( Yao et al., 2025 )
KL
84M
1024
0.28
Diff.
675M
2.17
205.6
1.35
295.3
MAR-L ( Li et al., 2024 )
KL
66M
256
0.87
MAR Diff.
479M
2.60
221.4
1.78
296.0
Table 1: System-level comparison of different tokenizers and generation models on ImageNet 256 × 256. ↓ and ↑ indicate whether lower or higher values are better. We categorize tokenizers into three groups: 2D tokenizers, fixed-length 1D tokenizers, and variable-length 1D tokenizers.
Figure 3: Training curves. (a) Training loss. (b) Training accuracy (%). (c) gFID without classifier-free guidance across training steps and model sizes.
Figure 4: Examples of image generation with NesTok-XL on ImageNet 256×256 .
Figure 5: Visualization of variable-length generation on ImageNet 256×256 resolution. Images are generated using NesTok-XL with 32-256 tokens.
Figure 6: gFID w/ and w/o cfg across different methods and model sizes.
Figure 7: Visualization of variable-length reconstruction on ImageNet 256×256 resolution. 1D denotes the 1D Tokenizer with 256 tokens.
Training Setting
AR Training
Evaluation
Loss ↓
Acc. ↑
rFID ↓
gFID ↓
32
64
128
256
32
64
128
256
1D Tokenizer
5.37
7.9%
-
-
-
0.83
-
-
-
2.33
+Nested Dropout
2.37
46.2%
4.28
2.88
2.07
1.93
5.98
4.64
3.12
2.46
+Cross-length Sample
4.70
11.3%
3.70
2.14
1.43
0.98
5.76
4.16
2.36
1.93
+Nested Self-Alignment Loss
4.68
11.4%
3.49
2.10
1.39
0.98
5.48
3.85
2.18
1.92
Table 2: Ablation study of different training settings. We report generation results using LlamaGen-Base (177M) trained for 200 epochs.
Method
Tokens
Params
gFID
images/s
ReTok-L
256
343M
2.27
10.11
One-D-Piece-L
256
318M
2.35
42.67
DetailFlow
256
326M
2.75
19.69
Semanticist-L
32
343M
2.57
1.03
NesTok-L
256
318M
1.50
16.35
Table 3: Comparison of sampling speed. Throughput (images/s) measured on a single H800 GPU with a batch size of 128, averaged over ten sampling runs using the official implementation of each method.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Visualization of variable-length generation on ImageNet 256×256 resolution. Images are generated using NesTok-XL with 8-256 tokens and without classifier-free guidance.
Figure 9: Visualization of variable-length generation on ImageNet 256×256 resolution. Images are generated using NesTok-XL with 8-256 tokens and without classifier-free guidance.
Method
32 Tokens
64 Tokens
128 Tokens
256 Tokens
One-D-Piece
3.23
2.10
1.42
1.08
ReTok-S-B
4.72
2.66
1.56
1.01
FlexTok d18-d28
1.45
1.37
1.20
1.08
Semanticist (DiT-XL)
1.40
1.07
0.86
0.72
DetailFlow
64.59
21.61
6.24
0.77
NesTok
3.49
2.10
1.39
0.98
Appendix
Table 4: Reconstruction performance (rFID ↓ ) at different token lengths on ImageNet 256×256 .
Figure 10: Visualization of class-condition image generation on 256×256 resolution.
Method
32 Tokens
64 Tokens
128 Tokens
256 Tokens
w/o cfg
cfg
w/o cfg
cfg
w/o cfg
cfg
w/o cfg
cfg
One-D-Piece-L
8.30
5.27
8.28
3.07
10.91
2.56
13.01
2.47
ReTok-L
7.45
6.76
4.69
4.54
3.79
3.18
3.78
2.66
DetailFlow-32
67.59
60.34
28.19
19.63
12.88
6.77
6.43
2.63
NesTok-L
4.82
3.78
2.87
2.34
2.28
1.71
2.08
1.50
Appendix
Table 5: Comparison of generation FID (gFID ↓ ) at different token lengths on ImageNet 256×256 .
Autoregressive image modeling relies on visual tokenizers to compress images into compact latent representations. We design an end-to-end training pipeline that jointly optimizes reconstruction and generation, enabling direct supervision from generation results to the tokenizer. This contrasts with prior two-stage approaches that train tokenizers and generative models separately. We further investigate leveraging vision foundation models to improve 1D tokenizers for autoregressive modeling. Our autoregressive generative model achieves strong empirical results, including a state-of-the-art FID score of 1.48 without guidance on ImageNet 256x256 generation.
Wenda Chu, Bingliang Zhang, Jiaqi Han +4
ByteDance Seed · California Institute of Technology · Stanford University
Despite progress in image tokenization, standard methods encode redundant information by mixing all granularities within each token, thus redundancy persists between tokens. The mix of information of different granularity also complicates the training of generators. This paper introduces SelfBootTok, a method that resolves this by cleanly decomposing information into global and local token groups. Through self-bootstrapped learning, the model predicts local details exclusively from global tokens, shifting the burden of visual details from the generator to the tokenizer. Consequently, our generator is far more efficient, requiring only global tokens and reducing computation by approximately 40%, while delivering superior reconstruction and generation. Moreover, this paradigm scales elegantly: by leveraging more data or parameters to self-supervise local representation learning, SelfBootTok achieves a new state-of-the-art gFID score of 1.56 using only 64 tokens.
Haozhe Chi, Jinghan Li, Hao Jiang +4
1Peking University · 2Central Media Technology Institute, Huawei
Vision Transformer (ViT) autoencoders have emerged as compelling tokenizers for images, offering improved reconstruction over convolutional tokenizers. However, existing ViT tokenizers cannot explore this landscape as performance degrades outside training resolutions, and reliance on adversarial losses prevents stable scaling. ViTok (Hansen-Estruch et al., 2025) found that the compression ratio r mediates a reconstruction-generation trade-off where lower r means better reconstructions but harder generations, so improving tokenizer reconstruction is key to more Pareto-optimal tokenizers. We introduce ViTok-v2, which addresses these limitations with native resolution support via NaFlex for generalization across resolutions and aspect ratios, and a novel DINOv3 perceptual loss that replaces both LPIPS and GAN objectives for stable training at any scale. ViTok-v2 is trained on about 2B images and scaled to 5B parameters, the largest image autoencoder to date. ViTok-v2 matches or exceeds state-of-the-art reconstruction at 256p and outperforms all baselines at 512p and above. In joint scaling experiments with flow matching generators, we show that scaling both the autoencoder and the generator advances the Pareto frontier of this trade-off.
Philippe Hansen-Estruch, Jiahui Chen, Vivek Ramanujan +9
University of Texas, Austin · University of Washington · 3Stanford University +3