One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail to fully exploit the representational capacity of the tokenizer, resulting in suboptimal performance in both image reconstruction and generation. In this work, we introduce NesTok, a nested self-alignment framework tailored to dynamic visual tokenizers. NesTok introduces cross-length training, which jointly optimizes reconstruction across token lengths while using the full-length sequence to guide shorter counterparts, enabling shorter token sequences to approach the reconstruction quality of full-length sequences. On ImageNet, NesTok improves substantially over standard training and achieves an rFID score of 0.98. On downstream image generation, it achieves the state-of-the-art gFID score of 1.46 on ImageNet 256×256 among existing variable-length autoregressive image generation methods. Code will be available at https://github.com/Jiawei804/NesTok.
Figures & tables
Figure 1: Codebook statistics and AR generation performance on ImageNet-1K. (a): Normalized entropy of the code distribution at each token position, computed over the training dataset and averaged within groups of 16 positions. (b): Average codebook utilization across token positions and gFID. (c): Sampling efficiency with different methods. (d): Sampling efficiency for different token lengths with different methods.
Figure 2: Overview of the nested self-aligned training pipeline. We train a 1D variable-length tokenizer using two components: (1) cross-length sampling, which jointly optimizes reconstruction from a full-length sequence and a sampled shorter sequence at each iteration; and (2) a nested self-alignment loss, which aligns latent representations across the two lengths to promote cross-length consistency and improve reconstruction.
Method
Tokenizer
Generator
w/o guidance
w/ guidance
Type
#Params
#Tokens
rFID ↓
Type
#Params
gFID ↓
IS ↑
gFID ↓
IS ↑
2D Tokenization
DiT-XL/2 ( Peebles and Xie, 2023 )
SD-VAE
84M
256
0.62
Diff.
675M
9.62
121.5
2.27
278.2
REPA-XL/2 ( Yu et al., 2024d )
KL
84M
1024
0.62
Diff.
675M
5.90
157.8
1.42
305.7
Lightning-DiT-XL ( Yao et al., 2025 )
KL
84M
1024
0.28
Diff.
675M
2.17
205.6
1.35
295.3
MAR-L ( Li et al., 2024 )
KL
66M
256
0.87
MAR Diff.
479M
2.60
221.4
1.78
296.0
Table 1: System-level comparison of different tokenizers and generation models on ImageNet 256 × 256. ↓ and ↑ indicate whether lower or higher values are better. We categorize tokenizers into three groups: 2D tokenizers, fixed-length 1D tokenizers, and variable-length 1D tokenizers.
Figure 3: Training curves. (a) Training loss. (b) Training accuracy (%). (c) gFID without classifier-free guidance across training steps and model sizes.
Figure 4: Examples of image generation with NesTok-XL on ImageNet 256×256 .
Figure 5: Visualization of variable-length generation on ImageNet 256×256 resolution. Images are generated using NesTok-XL with 32-256 tokens.
Figure 6: gFID w/ and w/o cfg across different methods and model sizes.
Figure 7: Visualization of variable-length reconstruction on ImageNet 256×256 resolution. 1D denotes the 1D Tokenizer with 256 tokens.
Training Setting
AR Training
Evaluation
Loss ↓
Acc. ↑
rFID ↓
gFID ↓
32
64
128
256
32
64
128
256
1D Tokenizer
5.37
7.9%
-
-
-
0.83
-
-
-
2.33
+Nested Dropout
2.37
46.2%
4.28
2.88
2.07
1.93
5.98
4.64
3.12
2.46
+Cross-length Sample
4.70
11.3%
3.70
2.14
1.43
0.98
5.76
4.16
2.36
1.93
+Nested Self-Alignment Loss
4.68
11.4%
3.49
2.10
1.39
0.98
5.48
3.85
2.18
1.92
Table 2: Ablation study of different training settings. We report generation results using LlamaGen-Base (177M) trained for 200 epochs.
Method
Tokens
Params
gFID
images/s
ReTok-L
256
343M
2.27
10.11
One-D-Piece-L
256
318M
2.35
42.67
DetailFlow
256
326M
2.75
19.69
Semanticist-L
32
343M
2.57
1.03
NesTok-L
256
318M
1.50
16.35
Table 3: Comparison of sampling speed. Throughput (images/s) measured on a single H800 GPU with a batch size of 128, averaged over ten sampling runs using the official implementation of each method.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Visualization of variable-length generation on ImageNet 256×256 resolution. Images are generated using NesTok-XL with 8-256 tokens and without classifier-free guidance.
Figure 9: Visualization of variable-length generation on ImageNet 256×256 resolution. Images are generated using NesTok-XL with 8-256 tokens and without classifier-free guidance.
Method
32 Tokens
64 Tokens
128 Tokens
256 Tokens
One-D-Piece
3.23
2.10
1.42
1.08
ReTok-S-B
4.72
2.66
1.56
1.01
FlexTok d18-d28
1.45
1.37
1.20
1.08
Semanticist (DiT-XL)
1.40
1.07
0.86
0.72
DetailFlow
64.59
21.61
6.24
0.77
NesTok
3.49
2.10
1.39
0.98
Appendix
Table 4: Reconstruction performance (rFID ↓ ) at different token lengths on ImageNet 256×256 .
Figure 10: Visualization of class-condition image generation on 256×256 resolution.
Method
32 Tokens
64 Tokens
128 Tokens
256 Tokens
w/o cfg
cfg
w/o cfg
cfg
w/o cfg
cfg
w/o cfg
cfg
One-D-Piece-L
8.30
5.27
8.28
3.07
10.91
2.56
13.01
2.47
ReTok-L
7.45
6.76
4.69
4.54
3.79
3.18
3.78
2.66
DetailFlow-32
67.59
60.34
28.19
19.63
12.88
6.77
6.43
2.63
NesTok-L
4.82
3.78
2.87
2.34
2.28
1.71
2.08
1.50
Appendix
Table 5: Comparison of generation FID (gFID ↓ ) at different token lengths on ImageNet 256×256 .