Image and video diffusion models allocate equal computation to every region, even when the intended scene calls for varying levels of detail. The spatial distribution of detail can often be anticipated before generation, indicating where computation can be reduced. We introduce Level-of-Token (LoT) Diffusion, a framework that turns this knowledge into an explicit multiresolution token layout (Level-of-Token layout) for adaptive and efficient generation. Tokens represent rectangular patches of varying sizes and shapes, allocating finer tokens where detail is needed and coarser tokens elsewhere. We adapt pretrained diffusion transformers to LoT layouts through a patch-wise asymmetric flow parametrization and embeddings for multiresolution tokens, preserving full-resolution flow prediction at every denoising step while processing only a reduced token sequence. LoT Diffusion enables layout-adaptive generation while preserving pretrained generative priors. We demonstrate LoT with layouts derived from semantic masks, bounding boxes, texture variance, and depth-of-field cues, as well as agentic plans. Across image and video generation, LoT offers favorable quality-efficiency tradeoffs, with significant speedups determined by the layout's token budget. Our project website is at https://georgenakayama.github.io/lotdiffusion/.
Figures & tables
Figure 1: Level-of-Token (LoT) Diffusion allocates compute using a spatial token layout that specifies the spatially varying level of detail for efficient image and video generation. Derived from diverse spatial priors, token layouts guide detail allocation with speedups tied to the token budget.
Figure 2: Token layout notation.
Figure 3: Level-of-Token Diffusion pipeline. LoT Diffusion maps a full-resolution token grid xt to a sequence of reduced tokens specified by the LoT layout. The DiT blocks process these tokens conditioned on token shapes s , timestep t , and text c . Extent-dependent heads hei restore the asymmetric velocity u^ , which is converted into the full-rank velocity for flow-matching training.
Figure 4: Image generation with different layouts. We generate images using arbitrary LoT layouts derived from signals that indicate spatial importance, including bounding boxes, semantic masks, texture variance, or depth maps, allocating finer tokens to regions that require detail.
Figure 5: Video generation with different layouts. Our model generates video under arbitrary LoT layouts derived from bounding box, semantic mask, texture variance, or depth input, allocating tokens where an application needs detail and reducing them elsewhere, such as sky and road surface in autonomous driving and flat untextured surfaces in real-time gaming.
Figure 6: Layout-adaptive image generation.
Method
Compression
Speed
IR ↑
HPSv2.1 ↑
HPSv3 ↑
FID ↓
pFID ↓
TOPIQ ↑
MUSIQ ↑
Full Res. FLUX2 4B
1.000 ×
1.001 ×
0.9518
0.2808
10.4299
16.748
20.884
0.5852
69.783
ToMe-SD
1.893 ×
1.538 ×
0.1757
0.2182
5.3569
25.757
30.502
0.4888
58.991
DDiT
1.340 ×
0.8850
0.2613
9.0052
19.671
23.739
0.5593
66.233
Foveated Diffusion
1.040 ×
0.8280
0.2615
9.0092
14.944
18.595
0.5822
67.494
Ours (LoT-Flux2 4B)
1.560 ×
0.8833
0.2728
10.1211
13.361
16.393
0.6080
69.691
ToMe-SD
2.406 ×
1.793 ×
-0.2547
0.1946
2.9150
34.321
41.187
0.4488
53.558
Table 1: Image generation baseline comparisons. Our method achieves superior performance on most generation metrics across different token budgets, while remaining faster than the baselines.
Method
Compression
Speed ↑
Aesthetic ↑
Imaging ↑
Dynamic ↑
Background ↑
Subject ↑
Motion ↑
Full Res. Wan2.1 14B
1 ×
1.000 ×
0.5518
0.6465
0.8900
0.9332
0.9254
0.9899
ToMe-SD
2 ×
2.223 ×
0.3497
0.4925
0.9000
0.9017
0.8235
0.9594
Foveated Diffusion
2.153 ×
0.5258
0.5598
0.9650
0.9252
0.8800
0.9844
Ours (LoT-Wan2.1 14B)
2.262 ×
0.5324
0.6120
0.9650
0.9242
0.8775
0.9836
ToMe-SD
2.5 ×
2.802 ×
0.2930
0.4477
0.9300
0.9122
0.8071
0.9555
Foveated Diffusion
2.746 ×
0.5153
0.5263
0.9550
0.9236
0.8776
0.9839
Table 2: Video generation quantitative comparisons. Our model obtains the best Aesthetic and Imaging scores compared with baselines and achieves the most speedup under the same budgets.
Figure 7: Qualitative baseline comparisons for (a) image and (b) video generation with the same token budget. ToMe-SD and DDiT lead to noisy low-res. regions. Foveated Diffusion results in distortions near mixed-res. token boundaries. We preserve visual detail while accelerating generation.
Figure 10: Mask-guided Image Generation with Varying Token Budget.
Figure 11: Mask-guided Image Generation with Varying Token Budget.
Figure 12: VRS-guided Image Generation with Varying Token Budget – Part 1.
Figure 13: VRS-guided Image Generation with Varying Token Budget – Part 2
Figure 14: Image Quality at Matched Token Budgets. (a) Quantitative comparison. (b) Ten qualitative pairs, with token-matched outputs shown at their relative generation resolution. Hatched areas are display padding, not generated content.
Table 4: Layout Controllability Compared with Dense Generation. All values are percentages (higher is better).
Figure 19: Layout Controllability: LoT-FLUX.2 9B and Dense FLUX.2 9B. Selected comparisons (1/2): original grid (A), ours (A), relocated grid (B), ours (B), and pretrained FLUX.2 9B. Each row shares the prompt and initial noise; A and B also share the token budget. Yellow boxes mark the 1×1 -token regions in the layouts. Dense FLUX.2 receives no layout.
Figure 20: Layout Controllability: LoT-FLUX.2 9B and Dense FLUX.2 9B. Selected comparisons (2/2), following the same arrangement and protocol as Figure 19 .
Figure 21: Ablation Study Qualitative Results.
Method
Compression
Time (s) ↓
Speed ↑
IR ↑
HPSv2.1 ↑
HPSv3 ↑
FID ↓
pFID ↓
TOPIQ ↑
MUSIQ ↑
Ours w/o size embedding
1.523 ×
4.372
1.322 ×
0.9294
0.2717
10.0284
14.584
17.472
0.6030
69.433
Ours w/ direct HR prediction
1.523 ×
4.264
1.355 ×
0.6301
0.2601
7.0595
26.487
36.072
0.5134
63.346
Ours w/ binary grid
1.518 ×
4.167
1.387 ×
0.7472
0.2520
8.1955
20.603
25.073
0.5517
66.443
Ours
1.523 ×
4.366
1.323 ×
0.9105
0.2764
10.3477
13.800
17.019
0.6170
70.410
Ours w/o size embedding
1.893 ×
3.718
1.554 ×
0.9008
0.2665
9.7299
13.934
17.320
0.5888
68.563
Ours w/ direct HR prediction
1.893 ×
3.609
1.601 ×
0.4335
0.2503
5.3532
37.414
55.443
0.4581
59.576
Appendix
Table 5: Ablation Study Across Different Token Budgets.
Figure 22: Limitation: Fine-Structure Degradation at Lower Token Budgets. The lower-budget result shows more broken staff lines.
Pixel-space diffusion models avoid the reconstruction ceiling of latent diffusion models by generating directly in image space. However, their substantially higher token count makes generation expensive due to the quadratic complexity of self-attention. Several existing efficiency methods reduce this cost by using larger patches at selected denoising steps, thereby representing the image with fewer tokens. Yet, each step still uses a single patch size uniformly across the entire image, overlooking that different regions suffer different fidelity losses when coarsened. We introduce MOSAIK, a damage-guided framework that varies patch size across regions and denoising steps. MOSAIK adapts the PixelDiT backbone to generate arbitrary heterogeneous patch layouts, and a lightweight predictor uses intermediate denoising features to estimate the fidelity loss caused by coarsening each region. Given a token budget, our damage-guided layout predictor assigns fine patches to sensitive regions and coarse patches elsewhere. Remarkably, while reducing FLOPs by 70% and token count by 83%, MOSAIK matches the full-compute PixelDiT on GenEval and its DPG-Bench score drops by only 1.0 point. Compared to diverse efficiency paradigms, including temporal patch scheduling and feature caching, our approach delivers highly competitive performance at moderate budgets and consistently outperforms these baselines in highly constrained compute regimes.
Latent Diffusion Models (LDMs) have become dominant in visual synthesis, but their quality-compute trade-off is largely constrained by the tokenizer's fixed compression ratio. Variable-length tokenizers (VLTs) promise adaptive compression by varying token counts, allowing diffusion models to flexibly balance quality and compute. However, conventional VLTs modulate length by truncating ordered token sequences, which makes token semantics depend on token position and breaks representational alignment across lengths. This leads to a cross-length shift in the latent distribution that hinders a single variable-length diffusion model from operating effectively. To address this, we propose a novel variable-length tokenizer that modulates length by merging tokens. We show that encouraging similar tokens to merge enables direct cross-length representation alignment when the diffusion transformer operates according to the merging pattern. Since conventional merging methods are data-dependent, making the merging pattern inaccessible during generation, we introduce learnable global merging, which is data-independent, to ensure compatibility with diffusion transformers. On ImageNet 256×256 generation, our merging-based variable-length tokenizer integrated with a diffusion transformer achieves a superior gFID-compute trade-off compared to prior VLT methods. Code is available at this https URL
Dong Hoon Lee, Seunghoon Hong
Kim Jaechul Graduate School of AI, KAIST, Daejeon, South Korea · School of Computing, KAIST, Daejeon, South Korea
Diffusion Transformers (DiTs) have achieved state-of-the-art video generation quality, but they incur immense computational cost because standard inference applies the same number of denoising steps uniformly to every token in the sequence. It is well known that human vision ignores vast amounts of redundant motion. Why, then, do our densest models treat every spatiotemporal token with equal priority? In this paper, we introduce Heterogeneous Step Allocation (HSA), a training-free inference algorithm that assigns varying step budgets to different spatiotemporal tokens based on their velocity dynamics. To resolve the resulting sequence-length mismatch without sacrificing global context, HSA introduces a KV-cache synchronization mechanism that allows active tokens to attend to the full sequence while entirely bypassing inactive tokens. Furthermore, we derive a cached Euler update that advances the latent states of skipped tokens in a single operation without additional model evaluations. We evaluate HSA on the Wan-2 and LTX-2 models for both text-to-video (T2V) and image-to-video (I2V) generation. Our results demonstrate that HSA significantly outperforms previous state-of-the-art caching methods and the vanilla Flow Matching baseline, especially at aggressive acceleration regimes (e.g., 50% and 25% runtimes). Crucially, HSA achieves a superior quality-runtime Pareto frontier without the need for expensive offline profiling, robustly preserving structural integrity and generation quality even under tight computational budgets. Project page: https://ernestchu.github.io/hsa