Image and video diffusion models allocate equal computation to every region, even when the intended scene calls for varying levels of detail. The spatial distribution of detail can often be anticipated before generation, indicating where computation can be reduced. We introduce Level-of-Token (LoT) Diffusion, a framework that turns this knowledge into an explicit multiresolution token layout (Level-of-Token layout) for adaptive and efficient generation. Tokens represent rectangular patches of varying sizes and shapes, allocating finer tokens where detail is needed and coarser tokens elsewhere. We adapt pretrained diffusion transformers to LoT layouts through a patch-wise asymmetric flow parametrization and embeddings for multiresolution tokens, preserving full-resolution flow prediction at every denoising step while processing only a reduced token sequence. LoT Diffusion enables layout-adaptive generation while preserving pretrained generative priors. We demonstrate LoT with layouts derived from semantic masks, bounding boxes, texture variance, and depth-of-field cues, as well as agentic plans. Across image and video generation, LoT offers favorable quality-efficiency tradeoffs, with significant speedups determined by the layout's token budget. Our project website is at https://georgenakayama.github.io/lotdiffusion/.
Figures & tables
Figure 1: Level-of-Token (LoT) Diffusion allocates compute using a spatial token layout that specifies the spatially varying level of detail for efficient image and video generation. Derived from diverse spatial priors, token layouts guide detail allocation with speedups tied to the token budget.
Figure 2: Token layout notation.
Figure 3: Level-of-Token Diffusion pipeline. LoT Diffusion maps a full-resolution token grid xt to a sequence of reduced tokens specified by the LoT layout. The DiT blocks process these tokens conditioned on token shapes s , timestep t , and text c . Extent-dependent heads hei restore the asymmetric velocity u^ , which is converted into the full-rank velocity for flow-matching training.
Figure 4: Image generation with different layouts. We generate images using arbitrary LoT layouts derived from signals that indicate spatial importance, including bounding boxes, semantic masks, texture variance, or depth maps, allocating finer tokens to regions that require detail.
Figure 5: Video generation with different layouts. Our model generates video under arbitrary LoT layouts derived from bounding box, semantic mask, texture variance, or depth input, allocating tokens where an application needs detail and reducing them elsewhere, such as sky and road surface in autonomous driving and flat untextured surfaces in real-time gaming.
Figure 6: Layout-adaptive image generation.
Method
Compression
Speed
IR ↑
HPSv2.1 ↑
HPSv3 ↑
FID ↓
pFID ↓
TOPIQ ↑
MUSIQ ↑
Full Res. FLUX2 4B
1.000 ×
1.001 ×
0.9518
0.2808
10.4299
16.748
20.884
0.5852
69.783
ToMe-SD
1.893 ×
1.538 ×
0.1757
0.2182
5.3569
25.757
30.502
0.4888
58.991
DDiT
1.340 ×
0.8850
0.2613
9.0052
19.671
23.739
0.5593
66.233
Foveated Diffusion
1.040 ×
0.8280
0.2615
9.0092
14.944
18.595
0.5822
67.494
Ours (LoT-Flux2 4B)
1.560 ×
0.8833
0.2728
10.1211
13.361
16.393
0.6080
69.691
ToMe-SD
2.406 ×
1.793 ×
-0.2547
0.1946
2.9150
34.321
41.187
0.4488
53.558
Table 1: Image generation baseline comparisons. Our method achieves superior performance on most generation metrics across different token budgets, while remaining faster than the baselines.
Method
Compression
Speed ↑
Aesthetic ↑
Imaging ↑
Dynamic ↑
Background ↑
Subject ↑
Motion ↑
Full Res. Wan2.1 14B
1 ×
1.000 ×
0.5518
0.6465
0.8900
0.9332
0.9254
0.9899
ToMe-SD
2 ×
2.223 ×
0.3497
0.4925
0.9000
0.9017
0.8235
0.9594
Foveated Diffusion
2.153 ×
0.5258
0.5598
0.9650
0.9252
0.8800
0.9844
Ours (LoT-Wan2.1 14B)
2.262 ×
0.5324
0.6120
0.9650
0.9242
0.8775
0.9836
ToMe-SD
2.5 ×
2.802 ×
0.2930
0.4477
0.9300
0.9122
0.8071
0.9555
Foveated Diffusion
2.746 ×
0.5153
0.5263
0.9550
0.9236
0.8776
0.9839
Table 2: Video generation quantitative comparisons. Our model obtains the best Aesthetic and Imaging scores compared with baselines and achieves the most speedup under the same budgets.
Figure 7: Qualitative baseline comparisons for (a) image and (b) video generation with the same token budget. ToMe-SD and DDiT lead to noisy low-res. regions. Foveated Diffusion results in distortions near mixed-res. token boundaries. We preserve visual detail while accelerating generation.
Figure 10: Mask-guided Image Generation with Varying Token Budget.
Figure 11: Mask-guided Image Generation with Varying Token Budget.
Figure 12: VRS-guided Image Generation with Varying Token Budget – Part 1.
Figure 13: VRS-guided Image Generation with Varying Token Budget – Part 2
Figure 14: Image Quality at Matched Token Budgets. (a) Quantitative comparison. (b) Ten qualitative pairs, with token-matched outputs shown at their relative generation resolution. Hatched areas are display padding, not generated content.
Table 4: Layout Controllability Compared with Dense Generation. All values are percentages (higher is better).
Figure 19: Layout Controllability: LoT-FLUX.2 9B and Dense FLUX.2 9B. Selected comparisons (1/2): original grid (A), ours (A), relocated grid (B), ours (B), and pretrained FLUX.2 9B. Each row shares the prompt and initial noise; A and B also share the token budget. Yellow boxes mark the 1×1 -token regions in the layouts. Dense FLUX.2 receives no layout.
Figure 20: Layout Controllability: LoT-FLUX.2 9B and Dense FLUX.2 9B. Selected comparisons (2/2), following the same arrangement and protocol as Figure 19 .
Figure 21: Ablation Study Qualitative Results.
Method
Compression
Time (s) ↓
Speed ↑
IR ↑
HPSv2.1 ↑
HPSv3 ↑
FID ↓
pFID ↓
TOPIQ ↑
MUSIQ ↑
Ours w/o size embedding
1.523 ×
4.372
1.322 ×
0.9294
0.2717
10.0284
14.584
17.472
0.6030
69.433
Ours w/ direct HR prediction
1.523 ×
4.264
1.355 ×
0.6301
0.2601
7.0595
26.487
36.072
0.5134
63.346
Ours w/ binary grid
1.518 ×
4.167
1.387 ×
0.7472
0.2520
8.1955
20.603
25.073
0.5517
66.443
Ours
1.523 ×
4.366
1.323 ×
0.9105
0.2764
10.3477
13.800
17.019
0.6170
70.410
Ours w/o size embedding
1.893 ×
3.718
1.554 ×
0.9008
0.2665
9.7299
13.934
17.320
0.5888
68.563
Ours w/ direct HR prediction
1.893 ×
3.609
1.601 ×
0.4335
0.2503
5.3532
37.414
55.443
0.4581
59.576
Appendix
Table 5: Ablation Study Across Different Token Budgets.
Figure 22: Limitation: Fine-Structure Degradation at Lower Token Budgets. The lower-budget result shows more broken staff lines.