Most image generation models rely on uniform tokenization, allocating the exact same computational budget to equally-sized image patches. This static paradigm cannot adapt to different resource constraints at inference time, and yields suboptimal quality-cost tradeoff by devoting the same effort to both plain backgrounds and intricate details. We propose BudgetPix, an adaptive tokenization framework that dynamically allocates compute based on visual complexity and spatial layout, enabling flexible computational budgeting at inference time. BudgetPix comprises three key components: (1) an adaptive encoder that maps a fixed-size image to a variable-length token sequence using an entropy-guided quadtree alongside a multi-scale patch embedder; (2) a scale-aware decoder reconstructs fixed-resolution images from multi-scale token sets; and (3) a flexible training and sampling schedule that enables pixel-space denoisers to operate across variable token counts. BudgetPix seamlessly integrates with existing pixel-space diffusion architectures, enabling a single checkpoint to be operated at a wide range of compute budgets. Evaluated on text-to-image generation, BudgetPix matches the fidelity of MiniT2I-L at 5122 and PixelDiT at 10242 using just 25% of the original compute budget. In class-conditional generation using a MeanFlow backbone, BudgetPix requires merely 60% of the full compute budget to produce images with near-zero quality degradation, observing a marginal 0.8-point increase in FID. Comprehensive assessments by human and VLM judges confirm that BudgetPix establishes a significantly improved quality-efficiency tradeoff over prior budget-adaptive baselines. More details are available at our project page: https://karaozgur.com/BudgetPix
Figures & tables
Figure 1: Test-time compute adjustment. BudgetPix dynamically scales inference costs, achieving a strong quality-efficiency tradeoff across JiT (top), MiniT2I (bottom left), PixelDiT (bottom right), 3 popular architectures for pixel-space image diffusion. Percentages show the fraction of tokens used.
Figure 2: Uniform vs. content-adaptive tokenization. For a 512×512 image, uniform tokenization requires 1,024 tokens with 16×16 patches (a) and 256 tokens with 32×32 patches (c). BudgetPix adapts to visual content, requiring only 307 (b) and 127 (d) tokens. Baselines in (a) and (c) represent full-budget MiniT2I-L and JiT-L; BudgetPix in (b) and (d) operates at 30% and 50% budgets, respectively.
Figure 3: Overview. BudgetPix enables compute-adaptive pixel-space diffusion with three components: an encoder, a decoder, and flexible training and inference across token counts. The Encoder and Decoder (purple) are trained from scratch; the Transformer (blue) is initialized from existing baselines.
Figure 4: Non-uniform layout construction. Given an image (the original image x during training, or the denoising image x^ during inference), the objective is to partition it into cells across multiple scales: p×p , 2p×2p , …, np×np , where n=2j for j≥1 . The process begins by dividing the image into np×np cells, which is the coarsest scale. For each cell, we construct a pixel intensity histogram and compute its entropy H . If H<τnp+δ (where τnp is the scale-specific threshold and δ is a global offset), the cell remains a single token. Otherwise, it is subdivided into four cells of the next finer scale ( n←n/2 ). This recursive subdivision terminates at n=1 , corresponding to the minimum cell size of p×p .
Figure 5: BudgetPix Architecture. The encoder receives the noisy input zt and its corresponding spatial layout L , converts the image into a sequence of multi-scale patches. The finest p×p tokens route directly through the fine path, while coarser patches navigate a dedicated coarse path in both the encoder and decoder. In the encoder, each coarse patch undergoes simultaneous downscaling (to p×p ) and splitting via a Patch Aggregator (into n2 sub-patches). These streams are processed through MLP layers and a scale mixer before being fused into a single coarse token. In the decoder, each denoised coarse patch is upscaled back to its target resolution ( np×np ). Concurrently, a Patch Refiner splits the patch into 4×4 sub-patches, processes them through an MLP and a lightweight transformer, and fuses them with the upscaled patch to restore local texture. The central diffusion transformer is directly initialized from existing pixel-space models, modified only to operate across variable, dynamically determined token counts rather than a fixed count.
Figure 6: Qualitative comparison across budgets, MiniT2I-L/16 at 5122 . Same prompt and noise were used for all methods. BudgetPix generates image with high quality even at 10%.
Figure 7: Text-to-image results, MiniT2I-L/16 at 5122 . (a) Each metric plotted against the wall-clock speed-up over the released model (DPG: DPG-Bench; IR: ImageReward denote metrics). (b) Faithfulness to the method’s own full-budget image plotted against the token budget (Sim.: CLIP or DINOv2 feature similarity).
Metric
Method
100%
90%
80%
70%
60%
50%
25%
GenEval ↑
Base model
0.721
0.722
0.724
0.725
0.708
0.674
0.323
BudgetPix
0.741
0.740
0.741
0.738
0.746
0.743
0.725
DPG-Bench ↑
Base model
84.8
84.6
84.2
84.0
83.3
81.8
55.1
BudgetPix
85.2
85.2
85.0
85.0
84.6
84.6
83.7
Table 1: Text-to-image generation with PixelDiT-T2I at 10242 . BudgetPix consistently outperform the base model across different compute budgets.
Figure 8: FID and Inception Score (IS) on class-conditional generation across token budgets on ImageNet, JiT-L/32 at 5122 . FID is cut at 40.
Figure 9: User study (top) and VLM-as-Judge (bottom), MiniT2I-L/16 at 5122 : (a) How many percentage from the compute budget can be cut without a noticeable difference for human and VLM judges. (b) Head-to-head comparison of BudgetPix against baselines by human and VLM judges, showing our method is consistently preferred across all token budgets.
Model
100%
90%
80%
70%
60%
50%
25%
5122
Base Model
4.12
5.00
6.86
12.24
24.79
60.61
190.71
+ Patch Aggregator
3.96
4.52
5.26
6.67
8.56
11.58
25.07
+ Patch Refiner
3.94
4.35
4.80
5.57
6.59
8.30
16.71
2562
Base Model
3.61
4.24
5.84
10.67
23.46
63.94
145.46
+ Patch Aggregator
3.57
3.92
4.34
5.14
6.17
7.84
16.90
+ Patch Refiner
3.60
3.89
4.25
4.89
5.73
7.01
13.05
Table 2: Ablation on JiT-B, FID-50k. BudgetPix with both Patch Aggregator (sec 3.2 ) and Refiner (sec 3.3 ) works best across different compute budgets.
Figure 10: BudgetPix on one-step pixel MeanFlow pMF-L/16. BudgetPix provides a better quality-efficiency trade-off compared with the base model and other token merging baselines.
Appendix figures & tables60 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11: Every method across the budget: an umbrella.
Figure 12: Every method across the budget: a herd of cows.
Figure 13: Every method across the budget: a seamless vector pattern.
Figure 14: Every method across the budget: a pencil case and binoculars.
Figure 15: Every method across the budget: an ornate treasure chest.
Figure 16: Every method across the budget: three birds on a branch.
Figure 17: BudgetPix on JiT-L/32 at 5122 , one ImageNet class per row.
Figure 18: BudgetPix on JiT-L/32 at 5122 , one ImageNet class per row (continued).
Figure 19: BudgetPix on MiniT2I-L/16 at 5122 , one prompt per row.
Figure 20: BudgetPix on MiniT2I-L/16 at 5122 , one prompt per row (continued).
Figure 21: BudgetPix on PixelDiT-T2I at 10242 , one prompt per row.
Figure 22: BudgetPix on PixelDiT-T2I at 10242 , one prompt per row (continued).
Figure 23: One-step samples, pMF-L/16 at 2562 , one ImageNet class per row.
Figure 24: BudgetPix on pMF-L/16 at 2562 , one step, one ImageNet class per row.
Figure 25: BudgetPix on pMF-L/16 at 2562 , one step, one ImageNet class per row (continued).
Figure 26: BudgetPix on JiT-L/32 at 5122 : samples and their layouts, one ImageNet class per row.
Figure 27: BudgetPix on JiT-L/32 at 5122 : samples and their layouts, one ImageNet class per row (continued).
Figure 28: BudgetPix on MiniT2I-L/16 at 5122 : samples and their layouts, one prompt per row.
Figure 29: BudgetPix on MiniT2I-L/16 at 5122 : samples and their layouts, one prompt per row (continued).
Figure 30: BudgetPix on PixelDiT-T2I at 10242 : samples and their layouts, one prompt per row.
Figure 31: BudgetPix on PixelDiT-T2I at 10242 : samples and their layouts, one prompt per row (continued).
FID-50k ↓
Inception score ↑
parent
method
100%
90%
80%
70%
60%
50%
25%
100%
90%
80%
70%
60%
50%
25%
JiT-L/32, 5122
Base model
2.66
3.07
4.34
8.09
15.51
31.57
107.51
337
327
305
259
203
127
13
ToMe
2.66
3.25
6.82
14.83
28.40
49.16
129.69
337
307
257
197
134
78
13
FeatSim
2.66
3.20
6.04
16.62
47.19
110.60
172.31
337
312
264
179
77
16
5
BudgetPix
2.71
2.94
3.20
3.69
4.41
5.57
11.54
341
336
328
316
302
281
206
JiT-L/16, 2562
Base model
2.71
3.01
4.44
9.07
18.09
35.95
116.16
328
318
294
246
188
115
12
Appendix
Table 3: Class-conditional generation on ImageNet, JiT-L and JiT-B: FID-50k and Inception score.
Figure 32: Class-conditional generation on ImageNet, JiT-L and JiT-B.
Figure 33: Paired metrics against the token budget, JiT-L and JiT-B.
JiT-B/32, 5122
JiT-B/16, 2562
metric
method
90%
80%
70%
60%
50%
25%
90%
80%
70%
60%
50%
25%
PSNR ↑
Base model
29.4
25.8
23.1
21.4
19.7
16.0
34.1
29.4
25.8
23.4
21.2
18.3
BudgetPix
27.8
25.8
24.6
23.8
23.2
22.1
34.5
31.7
29.6
28.2
26.9
24.4
PSNR, merged pixels ↑
Base model
28.5
25.7
23.4
21.7
20.0
16.1
30.5
27.4
24.8
22.9
21.2
18.4
BudgetPix
28.6
26.8
25.5
24.6
23.9
22.3
32.0
30.1
28.6
27.6
26.7
24.5
PSNR, kept pixels ↑
Base model
30.1
26.4
23.5
21.5
19.6
15.2
35.9
31.2
27.4
24.6
21.7
17.3
Appendix
Table 4: Faithfulness to the dense image, JiT-B: paired metrics.
MiniT2I-B/16
MiniT2I-L/16
metric
method
100%
90%
80%
70%
60%
50%
25%
100%
90%
80%
70%
60%
50%
25%
GenEval ↑
Base model
0.876
0.884
0.872
0.865
0.506
0.040
0.006
0.882
0.877
0.868
0.854
0.773
0.183
0.017
ToMe
0.877
0.878
0.867
0.861
0.841
0.779
0.283
0.886
0.891
0.886
0.873
0.845
0.741
0.229
FeatSim
0.877
0.882
0.871
0.865
0.830
0.610
0.106
0.886
0.880
0.850
0.798
0.676
0.292
0.007
RTI
0.850
0.864
0.870
0.869
0.870
0.868
0.856
0.863
0.873
0.867
0.870
0.868
0.871
0.859
BudgetPix
0.877
0.884
0.875
0.879
0.875
0.871
0.866
0.880
0.880
0.875
0.875
0.874
0.880
0.874
Appendix
Table 5: Text-to-image generation, MiniT2I-B/16 and MiniT2I-L/16 at 5122 .
Figure 34: Five metrics against the measured wall-clock speed-up, MiniT2I-B/16 at 5122 .
Figure 35: Wall-clock speed-up against the token budget, four parents.
parent
method
100%
90%
80%
70%
60%
50%
25%
MiniT2I-L/16, 5122
Base model
1.01 ×
1.02 ×
1.18 ×
1.59 ×
1.88 ×
2.30 ×
3.59 ×
ToMe
0.91 ×
0.80 ×
0.90 ×
0.92 ×
1.21 ×
1.13 ×
1.37 ×
FeatSim
0.91 ×
0.82 ×
0.91 ×
1.36 ×
1.52 ×
1.70 ×
2.20 ×
RTI
0.93 ×
0.83 ×
0.92 ×
1.36 ×
1.52 ×
1.69 ×
2.16 ×
BudgetPix
1.01 ×
1.01 ×
1.18 ×
1.59 ×
1.88 ×
2.29 ×
3.59 ×
MiniT2I-B/16, 5122
Base model
1.02 ×
1.02 ×
1.17 ×
1.57 ×
1.84 ×
2.19 ×
3.35 ×
Appendix
Table 6: Wall-clock speed-up over the released model, every parent.
Figure 36: Quality against measured speed-up, JiT-L/32 and MiniT2I-L/16 at 5122 .
measure
method
100%
90%
80%
70%
60%
50%
25%
seconds per image
Base model
4.912
4.870
4.205
3.105
2.634
2.151
1.376
ToMe
5.413
6.171
5.493
5.353
4.082
4.382
3.600
FeatSim
5.419
6.023
5.452
3.643
3.249
2.903
2.252
RTI
5.305
5.947
5.402
3.636
3.257
2.926
2.294
BudgetPix
4.913
4.872
4.206
3.113
2.626
2.164
1.376
speed-up
Base model
1.01 ×
1.02 ×
1.18 ×
1.59 ×
1.88 ×
2.30 ×
3.59 ×
Appendix
Table 7: Wall-clock per image and speed-up, MiniT2I-L/16 at 5122 .
measure
method
100%
90%
80%
70%
60%
50%
25%
seconds per image
Base model
1.782
1.784
1.549
1.161
0.988
0.828
0.543
ToMe
1.986
2.340
2.061
2.029
1.564
1.695
1.255
FeatSim
1.988
2.222
2.057
1.472
1.329
1.191
0.983
RTI
1.970
2.223
2.067
1.471
1.353
1.218
1.018
BudgetPix
1.785
1.776
1.547
1.161
0.988
0.829
0.545
speed-up
Base model
1.02 ×
1.02 ×
1.17 ×
1.57 ×
1.84 ×
2.19 ×
3.35 ×
Appendix
Table 8: Wall-clock per image and speed-up, MiniT2I-B/16 at 5122 .
measure
method
100%
90%
80%
70%
60%
50%
25%
seconds per image
Base model
2.992
3.760
3.383
2.983
2.652
2.306
1.527
BudgetPix
2.974
3.754
3.378
2.973
2.651
2.307
1.526
speed-up
Base model
0.99 ×
0.79 ×
0.88 ×
0.99 ×
1.12 ×
1.29 ×
1.94 ×
BudgetPix
1.00 ×
0.79 ×
0.88 ×
1.00 ×
1.12 ×
1.29 ×
1.95 ×
Appendix
Table 9: Wall-clock per image and speed-up, PixelDiT-T2I at 10242 .
measure
method
100%
90%
80%
70%
60%
50%
25%
seconds per image
Base model
0.394
0.376
0.346
0.323
0.300
0.267
0.195
ToMe
0.377
0.381
0.361
0.350
0.339
0.325
0.287
FeatSim
0.378
0.381
0.359
0.341
0.324
0.301
0.256
BudgetPix
0.396
0.377
0.348
0.322
0.297
0.265
0.194
speed-up
Base model
0.96 ×
1.00 ×
1.09 ×
1.17 ×
1.26 ×
1.41 ×
1.94 ×
ToMe
1.00 ×
0.99 ×
1.05 ×
1.08 ×
1.12 ×
1.16 ×
1.32 ×
Appendix
Table 10: Wall-clock per image and speed-up, JiT-L/32 at 5122 .
measure
method
100%
90%
80%
70%
60%
50%
25%
seconds per image
Base model
0.131
0.134
0.129
0.122
0.115
0.108
0.089
ToMe
0.124
0.134
0.129
0.125
0.122
0.116
0.103
FeatSim
0.122
0.127
0.123
0.118
0.113
0.108
0.097
BudgetPix
0.132
0.135
0.128
0.123
0.116
0.109
0.091
speed-up
Base model
0.95 ×
0.92 ×
0.96 ×
1.02 ×
1.08 ×
1.15 ×
1.39 ×
ToMe
1.00 ×
0.92 ×
0.96 ×
0.99 ×
1.02 ×
1.07 ×
1.20 ×
Appendix
Table 11: Wall-clock per image and speed-up, JiT-B/32 at 5122 .
measure
method
100%
90%
80%
70%
60%
50%
25%
seconds per image
Base model
0.383
0.365
0.339
0.314
0.291
0.259
0.197
ToMe
0.371
0.374
0.354
0.343
0.332
0.319
0.279
FeatSim
0.372
0.377
0.350
0.333
0.314
0.293
0.249
BudgetPix
0.384
0.368
0.342
0.315
0.292
0.261
0.195
speed-up
Base model
0.97 ×
1.02 ×
1.09 ×
1.18 ×
1.27 ×
1.43 ×
1.88 ×
ToMe
1.00 ×
0.99 ×
1.05 ×
1.08 ×
1.12 ×
1.17 ×
1.33 ×
Appendix
Table 12: Wall-clock per image and speed-up, JiT-L/16 at 2562 .
measure
method
100%
90%
80%
70%
60%
50%
25%
seconds per image
Base model
0.121
0.119
0.113
0.106
0.098
0.089
0.072
ToMe
0.118
0.130
0.122
0.119
0.113
0.111
0.097
FeatSim
0.117
0.122
0.116
0.112
0.107
0.101
0.091
BudgetPix
0.123
0.122
0.115
0.108
0.099
0.091
0.074
speed-up
Base model
0.97 ×
0.98 ×
1.03 ×
1.11 ×
1.19 ×
1.32 ×
1.63 ×
ToMe
0.99 ×
0.90 ×
0.96 ×
0.98 ×
1.04 ×
1.06 ×
1.21 ×
Appendix
Table 13: Wall-clock per image and speed-up, JiT-B/16 at 2562 .
Figure 37: Budget mode against threshold mode, FID-50k on JiT-L/32 at 5122 .
Figure 38: The layout as the budget falls, JiT-L/32 at 5122 , budget mode.
budget mode, budget
100%
90%
80%
70%
60%
50%
40%
30%
25%
20%
10%
tokens (% of grid)
100.0
89.5
80.1
69.5
60.2
49.6
39.1
29.7
25.0
19.1
9.8
FID-50k ↓
2.71
2.94
3.20
3.69
4.41
5.57
7.33
9.73
11.54
18.12
44.62
Inception score ↑
341
336
328
316
302
281
256
228
206
164
76
threshold mode, δ
− 1
− 0.5
0
+ 0.5
+ 1
+ 1.5
+ 1.9
+ 2
+ 2.7
+ 3
+ 4
tokens (% of grid)
91.6
86.3
78.8
68.7
56.0
41.4
30.3
28.0
18.8
17.1
9.2
s.d. across images (points)
10.9
13.7
16.2
18.0
18.2
15.6
11.3
10.0
4.7
4.9
3.5
Appendix
Table 14: Budget mode against threshold mode, JiT-L/32 at 5122 : every point of Figure 37 .
metric
method
100%
90%
80%
70%
60%
50%
40%
30%
25%
tokens
256
230
205
179
154
128
102
77
64
FID-50k ↓
Base model
2.54
6.91
118.76
242.62
257.20
244.95
236.82
253.49
233.64
ToMe
2.54
3.06
5.93
14.66
36.04
70.62
109.40
122.59
149.41
FeatSim
2.54
5.71
20.44
57.37
100.26
134.35
180.26
194.43
192.94
BudgetPix
2.65
2.65
2.77
3.06
3.46
3.94
4.50
4.96
5.44
Inception score ↑
Base model
263
226
20
2
2
2
3
3
3
Appendix
Table 15: One-step generation, pMF-L/16 at 2562 : FID-50k and Inception score.
JiT
pMF-L/16
MiniT2I
PixelDiT-T2I
2562 and 5122
2562
5122
10242
width D
768 (B), 1024 (L)
1024
768 (B), 1248 (L)
1536
cell embedding (released)
p×p convolution to 128 channels, then 1×1 convolution to D
linear map of the flattened 16×16×3 cell
downscale path
released embedding of the patch resized to p×p
none
split path: scale mixer
n2 cell embeddings + learned positions; 1×1 conv to 256, GELU, depthwise n×n conv, 1×1 conv, GELU, linear to D
n2 cell embeddings + learned positions; 1×1 conv to 256, depthwise n×n conv, linear to D ; no activation
coarse token
downscale path + scale mixer + scale embedding
mean of the n2 cell embeddings + scale mixer
Appendix
Table 16: Architecture of the added modules, every family.
quantity
JiT-B/16
JiT-L/16
JiT-B/32
JiT-L/32
MiniT2I-B/16
MiniT2I-L/16
PixelDiT-T2I
2562
2562
5122
5122
5122
5122
10242
max ∣Δ∣ at initialisation
1.9×10−6
1.6×10−6
2.5×10−6
1.9×10−6
2.0×10−6
4.5×10−6
9.6×10−4
FID, released model
3.61
2.71
4.12
2.66
24.6
24.8
43.7
FID, adapted model at step 0
3.61
2.71
4.12
2.66
24.6
24.8
43.7
FID, adapted model after training
3.60
2.56
3.94
2.71
17.1
22.9
36.2
Appendix
Table 17: Zero-initialisation check, every parent.
setting
JiT-B/16
JiT-B/32
JiT-L/16
JiT-L/32
2562
5122
2562
5122
patch / token scales
16 / 16, 32, 64
32 / 32, 64, 128
16 / 16, 32, 64
32 / 32, 64, 128
parameters, base + added
131.3M + 2.39M
133.4M + 2.40M
459.1M + 3.06M
461.8M + 3.06M
steps / epochs / images seen
142,987 / 100 / 128.1M
71,494 / 50 / 64.1M
global batch (per GPU × GPUs × accum.)
896 ( 112×4×2 )
896 ( 56×4×4 )
learning rate, backbone / added
3.5×10−5 / 2.8×10−4
Appendix
Table 18: Fine-tuning and sampling settings, JiT parents (ImageNet-1k).
setting
MiniT2I-B/16
MiniT2I-L/16
PixelDiT-T2I
5122
10242
patch / token scales
16 / 16, 32, 64
parameters, base + added
258.1M + 2.26M
911.8M + 3.50M
1302.5M + 4.37M
training data
CC12M, LLaVA recaptions, 1M images
CC12M (1M), then the 120K mix
PD12M, 200K images
steps / images seen
100k / 25.6M
40k (24k CC12M, then 16k mix) / 10.2M
30k / 0.96M
global batch (per GPU × GPUs × accum.)
256 ( 16×4×4 )
256 ( 4×4×16 )
32 ( 4×4×2 )
Appendix
Table 19: Fine-tuning and sampling settings, text-to-image parents.
Figure 39: User-study form on where a difference is first seen: its start and first question.
Figure 40: User-study form on which image is more faithful: its start and first question.
cut before a difference is seen (%)
Base model
ToMe
FeatSim
RTI
BudgetPix
people, 20 prompts
30.3 [29.2, 31.4]
32.5 [30.1, 35.1]
23.3 [21.5, 25.3]
38.0 [31.8, 44.0]
51.8 [48.0, 55.7]
judge, the same 20 prompts
18.5 [14.5, 22.5]
20.0 [16.5, 24.0]
12.5 [11.0, 14.5]
24.0 [14.5, 35.0]
43.5 [35.0, 51.0]
judge, all 100 prompts
15.4 [14.1, 16.9]
15.9 [14.3, 17.5]
11.7 [11.0, 12.4]
21.1 [16.9, 25.8]
36.2 [31.9, 40.4]
Appendix
Table 20: User study and VLM-as-Judge, MiniT2I-L/16 at 5122 : the numbers of Figure 9 .
setting
JiT-B/32, L/32 ( 5122 )
JiT-B/16, L/16 ( 2562 )
MiniT2I-B/16
MiniT2I-L/16
PixelDiT-T2I
resolution
512×512
256×256
512×512
1024×1024
dense grid
256 tokens (32 px)
256 tokens (16 px)
1024 tokens (16 px)
4096 tokens (16 px)
token scales of the layout
32 / 64 / 128 px
16 / 32 / 64 px
budgets 100, 90, 80, 70, 60, 50, 25 %
256, 230, 205, 179, 154, 128, 64 tokens
1024, 922, 819, 717, 614, 512, 256 tokens
4096, 3686, 3277, 2867, 2458, 2048, 1024 tokens
conditioning
class label, random
text, pre-encoded
batch (images per sampler call)
50
32
20
8
Appendix
Table 21: Conditions of the wall-clock measurements, every parent.
parent
method
100%
90%
80%
70%
60%
50%
40%
30%
tokens
256
230
205
179
154
128
102
77
JiT-B/32, 5122
Base model
6.53
7.44
9.45
14.96
27.61
63.91
118.04
156.99
BudgetPix, 50 epochs
6.47
6.93
7.36
8.33
9.36
11.18
13.89
17.33
JiT-B/16, 2562
Base model
6.02
6.67
8.30
13.26
26.34
67.38
116.44
152.40
Appendix
Table 22: Reference runs, JiT-B: FID-10k.
parent
training
100%
90%
80%
70%
60%
50%
40%
30%
tokens
256
230
205
179
154
128
102
77
JiT-B/32, 5122
10 epochs
7.48
8.16
9.33
11.30
13.87
18.19
23.95
29.94
20 epochs
7.45
8.15
9.10
10.91
13.08
16.27
21.80
28.29
30 epochs
7.40
8.05
9.05
10.72
12.43
15.95
20.50
26.78
50 epochs
7.33
8.01
8.94
10.36
12.15
15.17
19.41
24.96
Appendix
Table 23: Pixel-stream decoder across training length, JiT-B/32 at 5122 : FID-10k.
parent
schedule
10 steps
15 steps
25 steps
50 steps
JiT-B/32, 5122
cosine
10.02
9.15
9.01
8.82
shift-3
11.71
11.21
10.61
10.10
uniform
12.14
11.88
10.42
9.10
Appendix
Table 24: Timestep schedule against step count, JiT-B/32 at 5122 : FID-10k.
parent
sampler
90%
80%
70%
60%
40%
tokens
230
205
179
154
102
JiT-B/32, 5122
cosine, guidance 3.0
7.19
7.75
8.75
10.17
14.50
cosine, guidance 4.0
7.21
7.31
7.62
8.37
11.14
cosine, guidance 5.0
8.03
7.87
7.92
8.17
9.96
cosine, guidance 6.0
8.93
8.55
8.46
8.45
9.66
uniform, guidance 3.0 (default)
6.93
7.36
8.33
9.36
13.89
Appendix
Table 25: Guidance scale along the budget, JiT-B/32 at 5122 : FID-10k.
parent
arm
80%
50%
30%
tokens
205
128
77
guidance
3.0
6.0
6.0
JiT-B/32, 5122
reference
7.36
9.10
11.69
downsample-lanczos
7.42
8.92
11.14
oracle-fine
7.38
9.34
11.56
Appendix
Table 26: Upper bound on a better decoder, JiT-B/32 at 5122 : FID-10k.
50%
25%
setting
value
128
64
layout signal
entropy quadtree (default)
10.99
19.79
random placement
19.51
23.41
oracle, 50-step draft
11.30
19.58
dense warm-up steps
0
12.44
20.41
2
11.20
20.33
Appendix
Table 27: Layout and sampler settings, JiT-B/32 at 5122 : FID-10k.
50%
25%
setting
value
128
64
guidance
2.0
19.09
31.73
2.5
13.77
24.27
3.0 (default)
10.99
19.79
4.0
8.96
15.13
5.0
8.67
13.50
Appendix
Table 28: Guidance scale and timestep schedule, JiT-B/32 at 5122 : FID-10k.
skew κ
−1.5
−1
−0.5
0 (default)
+0.5
+1
+1.5
FID-10k ↓
15.79
13.02
11.42
10.99
10.95
10.95
10.97
Inception score ↑
152
165
173
175
176
177
177
Appendix
Table 29: Entropy-threshold skew, JiT-B/32 at 5122 : FID-10k and Inception score at 128 tokens.
Figure 41: Layout evolution during sampling, JiT-B/32 at 5122 .
70%
50%
25%
metric
arm
179
128
64
FID-10k ↓
fixed (default)
8.20
10.99
19.79
pooled
7.79
10.82
20.39
PSNR ↑
fixed (default)
24.57 (20.86 / 28.40)
23.22 (19.83 / 26.84)
22.14 (18.96 / 25.59)
pooled
26.35 (22.05 / 29.54)
23.49 (20.47 / 26.56)
22.08 (19.18 / 25.27)
base model
20.58 (17.44 / 24.17)
18.68 (15.95 / 21.59)
16.02 (13.75 / 18.52)
Appendix
Table 30: Pooled against fixed budget, JiT-B/32 at 5122 : FID-10k and paired fidelity.
Figure 42: Pooled budget, JiT-B/32 at 5122 , a mean of 128 tokens.
Pixel-space diffusion models avoid the reconstruction ceiling of latent diffusion models by generating directly in image space. However, their substantially higher token count makes generation expensive due to the quadratic complexity of self-attention. Several existing efficiency methods reduce this cost by using larger patches at selected denoising steps, thereby representing the image with fewer tokens. Yet, each step still uses a single patch size uniformly across the entire image, overlooking that different regions suffer different fidelity losses when coarsened. We introduce MOSAIK, a damage-guided framework that varies patch size across regions and denoising steps. MOSAIK adapts the PixelDiT backbone to generate arbitrary heterogeneous patch layouts, and a lightweight predictor uses intermediate denoising features to estimate the fidelity loss caused by coarsening each region. Given a token budget, our damage-guided layout predictor assigns fine patches to sensitive regions and coarse patches elsewhere. Remarkably, while reducing FLOPs by 70% and token count by 83%, MOSAIK matches the full-compute PixelDiT on GenEval and its DPG-Bench score drops by only 1.0 point. Compared to diverse efficiency paradigms, including temporal patch scheduling and feature caching, our approach delivers highly competitive performance at moderate budgets and consistently outperforms these baselines in highly constrained compute regimes.
Image tokenizers, from 2D grids to recent 1D sequences, typically encode every image with the same fixed number of tokens. Yet visual complexity is highly heterogeneous, so a uniform budget overspends on simple inputs and underserves complex ones. Existing elastic tokenizers expose variable-length reconstructions, but often leave token length as a deployment-time operating point, a search target, or an external prediction rather than an output of the tokenizer itself. In this work, we ask whether a discrete visual tokenizer can budget itself in one pass. Our central finding is that actionable elasticity requires a representation--allocation co-design: prefixes must remain decodable across budgets, and the tokenizer must learn which prefix each image needs. We propose AdaTok, a self-budgeting discrete 1D tokenizer. AdaTok combines Prioritized Representation Learning, which orders tokens with nested tail masking and resolves budget-dependent semantic shift through Multi-Head LoRA decoder heads, with Adaptive Token Allocation, which trains a lightweight deterministic-group GRPO policy over candidate budgets. Dynamic Pareto Weighting balances fidelity and efficiency during policy training without manual trade-off sweeps. On ImageNet-1K, AdaTok-Full reaches rFID 1.31 at 256 tokens, while AdaTok-Adaptive attains rFID 1.50 using only ~118 tokens on average, outperforming discrete 1D baselines at comparable budgets. In autoregressive image generation, the shorter adaptive representation yields ~2.1x throughput over a fixed 256-token decode, suggesting that visual token count can be learned as a content-conditioned output rather than set as a fixed hyperparameter.
Xiaocheng Lu, Yuxi Chen, Jie Zhang +5
The Hong Kong University of Science and Technology · The Hong Kong Polytechnic University
Image and video diffusion models allocate equal computation to every region, even when the intended scene calls for varying levels of detail. The spatial distribution of detail can often be anticipated before generation, indicating where computation can be reduced. We introduce Level-of-Token (LoT) Diffusion, a framework that turns this knowledge into an explicit multiresolution token layout (Level-of-Token layout) for adaptive and efficient generation. Tokens represent rectangular patches of varying sizes and shapes, allocating finer tokens where detail is needed and coarser tokens elsewhere. We adapt pretrained diffusion transformers to LoT layouts through a patch-wise asymmetric flow parametrization and embeddings for multiresolution tokens, preserving full-resolution flow prediction at every denoising step while processing only a reduced token sequence. LoT Diffusion enables layout-adaptive generation while preserving pretrained generative priors. We demonstrate LoT with layouts derived from semantic masks, bounding boxes, texture variance, and depth-of-field cues, as well as agentic plans. Across image and video generation, LoT offers favorable quality-efficiency tradeoffs, with significant speedups determined by the layout's token budget. Our project website is at https://georgenakayama.github.io/lotdiffusion/.