We introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation. Compared to traditional approaches using 2D grids or 1D token sequences, we propose a hierarchical quadtree structure, bridging the gap between 2D spatial binding and 1D sequence-level flexibility. The QuadTok tokenizer dynamically allocates representational capacity to visually intricate areas while leaving homogeneous regions at a coarse resolution. Compared with a fixed 256-token grid, our ImageNet-trained tokenizer saves approximately 10% of tokens on ImageNet and 9% when transferred zero-shot to the COCO dataset, while maintaining comparable reconstruction fidelity. Furthermore, the natural causality introduced by the tree structure seamlessly enables autoregressive image generation. Conditioned on a quadtree topology supplied before generation, our 947M GPT-style generative model achieves a 2.08 gFID on the ImageNet 256×256 benchmark. Additionally, leveraging the strong spatial correlation preserved by the quadtree structure, the QuadTok generator enables zero-shot spatially controlled image generation capabilities. Code: https://github.com/myc634/QuadTok.
Figures & tables
Figure 2: (a) Overview of QuadTok. An image is serialized into a 1D token sequence from a quadtree via breadth-first traversal. ViT features are aggregated into quadtree tokens, quantized, and decoded hierarchically. (b) Hierarchical decoding. Coarse tokens recover global structure; fine tokens refine local regions. (c) Region-wise Complexity Guidance. Complementary probing trees estimate per-region reconstruction benefit to build a content-adaptive quadtree.
Figure 2
Type
Tokenizer
Codebook
ImageNet-1K
COCO
#Tokens
rFID ↓
PSNR ↑
#Tokens
rFID ↓
PSNR ↑
2D
Taming VQ-GAN ( Esser et al., 2021 )
16384
256
4.98
19.40
256
19.29
19.57
MaskGIT VQ-GAN ( Chang et al., 2022 )
1024
256
2.28
−
–
−
−
LlamaGen ( f=16 ) ( Sun et al., 2024 )
16384
256
2.19
20.79
256
8.11
20.42
LlamaGen ( f=8 ) ( Sun et al., 2024 )
16384
1024
0.59
24.45
1024
4.19
24.20
1D
TiTok-S ( Yu et al., 2024b )
4096
128
1.71
17.80
128
9.22
17.27
Table 1: Image reconstruction on ImageNet-1K and COCO at 256×256 . #Tokens reports dataset-specific mean sequence lengths for QuadTok. f denotes the LlamaGen downsampling factor. High-resolution results are provided in Appendix D . – denotes unavailable values.
Tokens
Model
Size
gFID ↓
IS ↑
Prec. ↑
Rec. ↑
1D
TiTok-S-128–MaskGIT ( Yu et al., 2024b )
287M
1.97
281.80
–
–
TiTok-S-128–AR ∗ ( Yu et al., 2024b )
775M
4.90
191.71
0.77
0.56
One-D-Piece-L ∗ ( Miwa et al., 2025b )
775M
2.99
235.10
0.81
0.59
FlexTok d18-d28 ( Bachmann et al., 2025a )
1.33B
1.86
–
–
–
GigaTok-B-L ( Xiong et al., 2025 )
1.4B
2.03
238.52
0.80
0.63
GigaTok-XL-XXL ( Xiong et al., 2025 )
1.4B
1.98
256.76
0.81
0.62
Table 2: Class-conditional generation on ImageNet evaluated at 256×256 . These are system-level comparisons with different training and sampling settings. Size denotes generator parameter count. ‡ : native-256 LlamaGen results. xAR and FlowAR use continuous 2D latents.
Figure 5: Zero-shot spatial layout control. Highlighted regions specify where the input quadtree is refined. Generated subjects align with these regions using the frozen ImageNet-trained model.
Target
Topology
Grad-CAM
Box
Mask
FID-50cls
Left
Random
29.6
47.5
51.1
35.03
Prescribed
58.3 (+28.7)
51.8 (+4.3)
59.0 (+7.9)
35.15 (+0.12)
Right
Random
70.4
51.0
48.5
35.03
Prescribed
87.9 (+17.5)
58.1 (+7.1)
55.7 (+7.2)
34.88 (-0.15)
Top
Random
33.9
40.7
42.9
35.03
Prescribed
55.4 (+21.5)
47.0 (+6.3)
47.8 (+4.9)
34.97 (-0.06)
Table 3: Zero-shot spatial control with a frozen ImageNet-trained QuadTok generator. Prescribed vs. random topologies at 192 tokens, 1,000 images each. Grad-CAM, Box, and Mask: % of subjects in the target half-plane ( ↑ ). FID-50cls ( ↓ ): against 2,500 class-matched real images.
Table 7
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Type
Tokenizer
#Tokens
Codebook size
rFID ↓
PSNR ↑
2D
MaskGIT VQ-GAN ( Chang et al., 2022 )
1024
1024
1.97
–
LlamaGen † ( Sun et al., 2024 )
1024
16384
0.70
23.03
1D
TiTok-L-64 ( Yu et al., 2024b )
64
4096
1.77
–
TiTok-B-128 ( Yu et al., 2024b )
128
4096
1.52
–
Quadtree
QuadTok
987
16384
0.84
24.96
Appendix
Table 6: ImageNet-1K reconstruction with 512×512 inputs. #Tokens is the mean for QuadTok. † : LlamaGen resizes reconstructions to 256×256 for evaluation. MaskGIT and TiTok values follow Table 2 of the TiTok paper. – denotes unreported values.
Model
Size
gFID ↓
IS ↑
Precision ↑
Recall ↑
DiT-XL/2 ( Peebles and Xie, 2023 )
675M
3.04
240.82
0.84
0.54
CAT ( Shen et al., 2025 )
≈ 431M
4.38
181.03
0.76
0.48
TiTok-B-128 ( Yu et al., 2024b ) -MaskGIT
177M
2.13
261.2
–
–
VAR- d36 -s ( Tian et al., 2024 )
2.3B
2.63
303.2
–
–
QuadTok
344M
2.734
268.01
0.835
0.537
Appendix
Table 7: Class-conditional ImageNet generation at 512×512 . Size denotes generator parameters. CAT uses a continuous VAE with DiT ( c=8 , original Table 7); TiTok uses MaskGIT with 64 sampling steps. VAR uses next-scale prediction, while QuadTok uses causal autoregressive decoding. Results are system-level comparisons; – denotes unreported metrics.
Parameters
p=0.25
p=0.50
p=0.75
Approx. tokens
128
192
255
111M
1.81–2.40 ×
1.14–1.58 ×
1.20 ×
343M
1.62–1.94 ×
1.09–1.20 ×
1.09 ×
775M
1.57–1.91 ×
1.15–1.21 ×
1.10 ×
Appendix
Table 8: Generation speedup relative to LlamaGen. The LlamaGen timing baseline uses randomly initialized generator weights in the 256×256 configuration and generates 256 tokens. At p=0.25 and 0.50 , ranges span batch sizes 1/32 and CFG scales 1/2. The p=0.75 column reports the near-equal-token comparison. Ratios measure autoregressive generation, excluding image decoding; values above 1 indicate faster generation. The 775M comparison is timing-only ( Appendix E ).
Scale
Method
Tokens
Gen. (s)
+ decode (s)
Speedup
111M
LlamaGen
256.0
5.440
5.551
1.00 ×
111M
QuadTok ( p=0.25 )
126.6
2.524
2.560
2.16 ×
111M
QuadTok ( p=0.50 )
189.2
3.528
3.570
1.54 ×
111M
QuadTok ( p=0.75 )
255.0
4.533
4.586
1.20 ×
343M
LlamaGen
256.0
8.404
8.512
1.00 ×
343M
QuadTok ( p=0.25 )
126.6
4.656
4.692
1.80 ×
Appendix
Table 9: Absolute generation latency at batch size 32 and CFG scale 2. Means over 50 timed batches, in seconds per batch. Gen. denotes autoregressive sampling; + decode includes image decoding. Results cover p=0.25 , 0.50 , and 0.75 under the protocol in Appendix E . Speedup is the ratio of LlamaGen to QuadTok generation latency.
Tokenizer
Tokens/image
Batch 32
Batch 128
LlamaGen VQ-16
256
2.927
2.885
TiTok L-32
32
1.480
1.352
TiTok B-64
64
0.586
0.512
TiTok S-128
128
0.305
0.277
QuadTok
234.3 / 233.2
8.795
6.186
Appendix
Table 10: Pretokenization latency at 256×256 . Mean milliseconds per image on one A100-80GB GPU. Token counts for QuadTok are measured on the timed batches (32 / 128); fixed encoders use a constant budget. Timing includes adaptive selection for QuadTok and ends at discrete codes, before reconstruction decoding.
Guidance Metric
Threshold ( τ )
#Tokens ↓
rFID ↓
PSNR ↑
LMSE
0.0005
230
1.67
20.52
LLPIPS
0.03
267
1.46
20.24
0.05
230
1.46
20.37
0.06
195
1.75
19.99
Combined ( LMSE×10+LLPIPS )
0.06
230
1.58
20.37
Appendix
Table 11: Ablation of complexity guidance metric and threshold τ . We compare MSE, LPIPS, and their linear combination as probing metrics; for LPIPS we sweep the expansion threshold τ , and report MSE and the combination at a matched ∼ 230-token budget on ImageNet-1K. LPIPS at τ=0.05 achieves the best perceptual quality (rFID 1.46 ) at 230 tokens, and is adopted as the default.