We introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation. Compared to traditional approaches using 2D grids or 1D token sequences, we propose a hierarchical quadtree structure, bridging the gap between 2D spatial binding and 1D sequence-level flexibility. The QuadTok tokenizer dynamically allocates representational capacity to visually intricate areas while leaving homogeneous regions at a coarse resolution. Compared with a fixed 256-token grid, our ImageNet-trained tokenizer saves approximately 10% of tokens on ImageNet and 9% when transferred zero-shot to the COCO dataset, while maintaining comparable reconstruction fidelity. Furthermore, the natural causality introduced by the tree structure seamlessly enables autoregressive image generation. Conditioned on a quadtree topology supplied before generation, our 947M GPT-style generative model achieves a 2.08 gFID on the ImageNet 256×256 benchmark. Additionally, leveraging the strong spatial correlation preserved by the quadtree structure, the QuadTok generator enables zero-shot spatially controlled image generation capabilities. Code: https://github.com/myc634/QuadTok.
Figures & tables
Figure 2: (a) Overview of QuadTok. An image is serialized into a 1D token sequence from a quadtree via breadth-first traversal. ViT features are aggregated into quadtree tokens, quantized, and decoded hierarchically. (b) Hierarchical decoding. Coarse tokens recover global structure; fine tokens refine local regions. (c) Region-wise Complexity Guidance. Complementary probing trees estimate per-region reconstruction benefit to build a content-adaptive quadtree.
Figure 2
Type
Tokenizer
Codebook
ImageNet-1K
COCO
#Tokens
rFID ↓
PSNR ↑
#Tokens
rFID ↓
PSNR ↑
2D
Taming VQ-GAN ( Esser et al., 2021 )
16384
256
4.98
19.40
256
19.29
19.57
MaskGIT VQ-GAN ( Chang et al., 2022 )
1024
256
2.28
−
–
−
−
LlamaGen ( f=16 ) ( Sun et al., 2024 )
16384
256
2.19
20.79
256
8.11
20.42
LlamaGen ( f=8 ) ( Sun et al., 2024 )
16384
1024
0.59
24.45
1024
4.19
24.20
1D
TiTok-S ( Yu et al., 2024b )
4096
128
1.71
17.80
128
9.22
17.27
Table 1: Image reconstruction on ImageNet-1K and COCO at 256×256 . #Tokens reports dataset-specific mean sequence lengths for QuadTok. f denotes the LlamaGen downsampling factor. High-resolution results are provided in Appendix D . – denotes unavailable values.
Tokens
Model
Size
gFID ↓
IS ↑
Prec. ↑
Rec. ↑
1D
TiTok-S-128–MaskGIT ( Yu et al., 2024b )
287M
1.97
281.80
–
–
TiTok-S-128–AR ∗ ( Yu et al., 2024b )
775M
4.90
191.71
0.77
0.56
One-D-Piece-L ∗ ( Miwa et al., 2025b )
775M
2.99
235.10
0.81
0.59
FlexTok d18-d28 ( Bachmann et al., 2025a )
1.33B
1.86
–
–
–
GigaTok-B-L ( Xiong et al., 2025 )
1.4B
2.03
238.52
0.80
0.63
GigaTok-XL-XXL ( Xiong et al., 2025 )
1.4B
1.98
256.76
0.81
0.62
Table 2: Class-conditional generation on ImageNet evaluated at 256×256 . These are system-level comparisons with different training and sampling settings. Size denotes generator parameter count. ‡ : native-256 LlamaGen results. xAR and FlowAR use continuous 2D latents.
Figure 5: Zero-shot spatial layout control. Highlighted regions specify where the input quadtree is refined. Generated subjects align with these regions using the frozen ImageNet-trained model.
Target
Topology
Grad-CAM
Box
Mask
FID-50cls
Left
Random
29.6
47.5
51.1
35.03
Prescribed
58.3 (+28.7)
51.8 (+4.3)
59.0 (+7.9)
35.15 (+0.12)
Right
Random
70.4
51.0
48.5
35.03
Prescribed
87.9 (+17.5)
58.1 (+7.1)
55.7 (+7.2)
34.88 (-0.15)
Top
Random
33.9
40.7
42.9
35.03
Prescribed
55.4 (+21.5)
47.0 (+6.3)
47.8 (+4.9)
34.97 (-0.06)
Table 3: Zero-shot spatial control with a frozen ImageNet-trained QuadTok generator. Prescribed vs. random topologies at 192 tokens, 1,000 images each. Grad-CAM, Box, and Mask: % of subjects in the target half-plane ( ↑ ). FID-50cls ( ↓ ): against 2,500 class-matched real images.
Table 7
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Type
Tokenizer
#Tokens
Codebook size
rFID ↓
PSNR ↑
2D
MaskGIT VQ-GAN ( Chang et al., 2022 )
1024
1024
1.97
–
LlamaGen † ( Sun et al., 2024 )
1024
16384
0.70
23.03
1D
TiTok-L-64 ( Yu et al., 2024b )
64
4096
1.77
–
TiTok-B-128 ( Yu et al., 2024b )
128
4096
1.52
–
Quadtree
QuadTok
987
16384
0.84
24.96
Appendix
Table 6: ImageNet-1K reconstruction with 512×512 inputs. #Tokens is the mean for QuadTok. † : LlamaGen resizes reconstructions to 256×256 for evaluation. MaskGIT and TiTok values follow Table 2 of the TiTok paper. – denotes unreported values.
Model
Size
gFID ↓
IS ↑
Precision ↑
Recall ↑
DiT-XL/2 ( Peebles and Xie, 2023 )
675M
3.04
240.82
0.84
0.54
CAT ( Shen et al., 2025 )
≈ 431M
4.38
181.03
0.76
0.48
TiTok-B-128 ( Yu et al., 2024b ) -MaskGIT
177M
2.13
261.2
–
–
VAR- d36 -s ( Tian et al., 2024 )
2.3B
2.63
303.2
–
–
QuadTok
344M
2.734
268.01
0.835
0.537
Appendix
Table 7: Class-conditional ImageNet generation at 512×512 . Size denotes generator parameters. CAT uses a continuous VAE with DiT ( c=8 , original Table 7); TiTok uses MaskGIT with 64 sampling steps. VAR uses next-scale prediction, while QuadTok uses causal autoregressive decoding. Results are system-level comparisons; – denotes unreported metrics.
Parameters
p=0.25
p=0.50
p=0.75
Approx. tokens
128
192
255
111M
1.81–2.40 ×
1.14–1.58 ×
1.20 ×
343M
1.62–1.94 ×
1.09–1.20 ×
1.09 ×
775M
1.57–1.91 ×
1.15–1.21 ×
1.10 ×
Appendix
Table 8: Generation speedup relative to LlamaGen. The LlamaGen timing baseline uses randomly initialized generator weights in the 256×256 configuration and generates 256 tokens. At p=0.25 and 0.50 , ranges span batch sizes 1/32 and CFG scales 1/2. The p=0.75 column reports the near-equal-token comparison. Ratios measure autoregressive generation, excluding image decoding; values above 1 indicate faster generation. The 775M comparison is timing-only ( Appendix E ).
Scale
Method
Tokens
Gen. (s)
+ decode (s)
Speedup
111M
LlamaGen
256.0
5.440
5.551
1.00 ×
111M
QuadTok ( p=0.25 )
126.6
2.524
2.560
2.16 ×
111M
QuadTok ( p=0.50 )
189.2
3.528
3.570
1.54 ×
111M
QuadTok ( p=0.75 )
255.0
4.533
4.586
1.20 ×
343M
LlamaGen
256.0
8.404
8.512
1.00 ×
343M
QuadTok ( p=0.25 )
126.6
4.656
4.692
1.80 ×
Appendix
Table 9: Absolute generation latency at batch size 32 and CFG scale 2. Means over 50 timed batches, in seconds per batch. Gen. denotes autoregressive sampling; + decode includes image decoding. Results cover p=0.25 , 0.50 , and 0.75 under the protocol in Appendix E . Speedup is the ratio of LlamaGen to QuadTok generation latency.
Tokenizer
Tokens/image
Batch 32
Batch 128
LlamaGen VQ-16
256
2.927
2.885
TiTok L-32
32
1.480
1.352
TiTok B-64
64
0.586
0.512
TiTok S-128
128
0.305
0.277
QuadTok
234.3 / 233.2
8.795
6.186
Appendix
Table 10: Pretokenization latency at 256×256 . Mean milliseconds per image on one A100-80GB GPU. Token counts for QuadTok are measured on the timed batches (32 / 128); fixed encoders use a constant budget. Timing includes adaptive selection for QuadTok and ends at discrete codes, before reconstruction decoding.
Guidance Metric
Threshold ( τ )
#Tokens ↓
rFID ↓
PSNR ↑
LMSE
0.0005
230
1.67
20.52
LLPIPS
0.03
267
1.46
20.24
0.05
230
1.46
20.37
0.06
195
1.75
19.99
Combined ( LMSE×10+LLPIPS )
0.06
230
1.58
20.37
Appendix
Table 11: Ablation of complexity guidance metric and threshold τ . We compare MSE, LPIPS, and their linear combination as probing metrics; for LPIPS we sweep the expansion threshold τ , and report MSE and the combination at a matched ∼ 230-token budget on ImageNet-1K. LPIPS at τ=0.05 achieves the best perceptual quality (rFID 1.46 ) at 230 tokens, and is adopted as the default.
One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail to fully exploit the representational capacity of the tokenizer, resulting in suboptimal performance in both image reconstruction and generation. In this work, we introduce NesTok, a nested self-alignment framework tailored to dynamic visual tokenizers. NesTok introduces cross-length training, which jointly optimizes reconstruction across token lengths while using the full-length sequence to guide shorter counterparts, enabling shorter token sequences to approach the reconstruction quality of full-length sequences. On ImageNet, NesTok improves substantially over standard training and achieves an rFID score of 0.98. On downstream image generation, it achieves the state-of-the-art gFID score of 1.46 on ImageNet 256×256 among existing variable-length autoregressive image generation methods. Code will be available at https://github.com/Jiawei804/NesTok.
Jiawei Zhang, Shuhao Liu, Rong Huang +4
North China Electric Power University · University of Science and Technology of China · Stanford University
Autoregressive image modeling relies on visual tokenizers to compress images into compact latent representations. We design an end-to-end training pipeline that jointly optimizes reconstruction and generation, enabling direct supervision from generation results to the tokenizer. This contrasts with prior two-stage approaches that train tokenizers and generative models separately. We further investigate leveraging vision foundation models to improve 1D tokenizers for autoregressive modeling. Our autoregressive generative model achieves strong empirical results, including a state-of-the-art FID score of 1.48 without guidance on ImageNet 256x256 generation.
Wenda Chu, Bingliang Zhang, Jiaqi Han +4
ByteDance Seed · California Institute of Technology · Stanford University
In this work, we explore the largely unexplored direction of building a generalist image tokenizer directly on top of a frozen vision foundation model (VFM). To build this tokenizer, we utilize a frozen VFM as the encoder and introduce two key innovations: (1) a region-adaptive quantization framework to eliminate spatial redundancy in standard 2D grid features, and (2) a semantic reconstruction objective that aligns the decoded outputs with the VFM's representations to preserve semantic fidelity. Grounded in these designs, we propose VFMTok, a generalist visual tokenizer capable of operating seamlessly in both discrete and continuous latent spaces. VFMTok achieves substantial improvements in synthesis quality while drastically enhancing token efficiency. For discrete autoregressive (AR) generation, it accelerates model convergence by \textbf{3 times} and achieves a state-of-the-art gFID of \textbf{1.36} on ImageNet class-conditional synthesis. Similarly, for continuous-space generation, integrating VFMTok with a denoising model yields an exceptional gFID of \textbf{1.25}. Furthermore, because the latent space inherently captures rich spatial semantics, VFMTok enables high-fidelity class-conditional synthesis without classifier-free guidance (\textbf{w/o CFG}) across both generative paradigms, significantly accelerating inference speed. Beyond these remarkable empirical results, we systematically investigate the underlying mechanisms of our approach. We discover that the specific self-supervised learning objectives utilized during VFM pre-training dictate its effectiveness as a tokenizer. Specifically, a VFM jointly optimized with global contrastive learning and latent masked image modeling provides the optimal representations for image tokenization. These insights establish a strong foundation and offer valuable guidance for the design of future image tokenizers.
Anlin Zheng, Qi Han, Xin Wen +5
University of Hong Kong, Pokfulam, Hong Kong · StepFun