Organizations: The Hong Kong University of Science and Technology (Guangzhou) · Griffith University · Data61, CSIRO · Artificial Intelligence Lab, Institute of Deep Perception Technology, JITRI
Diffusion transformers incur substantial inference cost through repeated model evaluations along a sampling trajectory. We introduce GeoShrink, a training-free acceleration method that retains the original solver grid while evaluating the model only at a prescribed set of anchors. At skipped stages, GeoShrink predicts the solver-facing output by adding a geometrically retained fraction of the latest observed innovation to the most recent exact output. We derive this rule from chordal tangent transport and round-trip line projection, and establish a geometric anchor-spacing principle that minimizes the largest adjacent gap expansion under fixed coverage and first span. The analysis characterizes the geometric closure and propagation of prediction errors without assuming access to future model outputs. Experiments cover image, video, motion, and audio generation, together with adapted 3D backends. At approximately 5× acceleration, GeoShrink improves FLUX PSNR by 3.10 dB over the strongest listed baseline. On HunyuanVideo, it achieves a reported 4.99× speedup and improves ChronoMagic-Bench-150 PSNR by 5.44 dB over the strongest listed fidelity baseline. Comparisons at fixed evaluation budgets further show substantial gains on motion, audio, music, and 3D generation.
Figures & tables
Figure 1 : Qualitative comparison of GeoShrink. Direct NFE reduction (left) vs. GeoShrink (right) under the same model-evaluation budget. GeoShrink better preserves visual structure and temporal evolution under aggressive acceleration.
Figure 2 : Motivation for adaptive innovation retention. From left to right, we show the full-NFE reference, fixed-coefficient predictions with ω∈{0,0.5,1.5,2} , and GeoShrink. Small ω underuses the latest innovation, while aggressive reuse ( ω>1 ) progressively introduces color, contrast, and appearance drift. GeoShrink instead adapts ωt from observable trajectory geometry and more closely preserves the full-NFE reference.
Figure 3 : Overview of GeoShrink. (a) Previous methods rely on internal feature caching, reusing or forecasting intermediate DiT representations across denoising steps. (b) GeoShrink operates at the solver interface, allocating 10 exact model queries over a 50-step grid through geometric anchoring. At skipped stages, it predicts the solver-facing field from the three latest exact outputs while preserving the original solver and its grid. (c) The retention coefficient ωt=(1+ρt2sin2θ)−1 combines the acute angle θ between successive innovations with the normalized midpoint forecast horizon ρt . Stronger turning or a longer horizon reduces the retained innovation in vt=vt(1)+ωtΔv .
vt←Vθ(xt,σt,c),t(3)←t(2),t(2)←t(1),t(1)←t
Algorithm 1 GeoShrink
Figure 4 : Qualitative comparison on text-to-image generation. GeoShrink better preserves the full-sampling reference than competing acceleration methods at comparable speedups.
Figure 5 : Qualitative comparison on text-to-video generation. GeoShrink preserves spatial details and temporal consistency more faithfully under aggressive acceleration.
Model
Benchmark
Method
Setting
Speedup
PSNR ↑
SSIM ↑
LPIPS ↓
Task Score ↑
SD3.5 Medium
Pick-a-Pic
SpeCa
τ0=7,β=0.3
2.40 ×
20.759
0.7743
0.2704
0.7873
TaylorSeer
N=3,O=1
2.49 ×
18.493
0.7454
0.2918
0.7815
TeaCache
l=0.4
2.66 ×
20.872
0.7616
0.2942
0.7671
ToCa
N=10,R=90%
2.09 ×
20.431
0.7057
0.3111
0.7662
ResilPhase
N=3,O=1
2.76 ×
18.483
0.7451
0.2921
0.7813
GeoShrink
B=10
2.80 ×
24.744
0.8450
0.2078
0.7924
Table 1 : Quantitative comparison of acceleration methods across image and video generation models. For the task-specific metric, we report ImageReward on Pick-a-Pic, MTScore on ChronoMagic-Bench-150, and VBench Score on VBench.
NFE
Method
Visual
Audio
CLIP ↑
FVD ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Mel Cos ↑
LogMel L1 ↓
SI-SDR ↑
10
Direct
33.7164
166.0071
17.4910
0.6190
0.3632
0.7730
1.1660
-0.2793
GeoShrink
33.8711
164.5004
24.2122
0.7883
0.1678
0.8528
0.6998
1.5133
15
Direct
33.6598
166.8756
19.2673
0.6732
0.2914
0.8537
0.8240
3.2408
GeoShrink
33.7885
166.1741
29.1320
0.8810
0.0818
0.9595
0.3645
7.6285
20
Direct
33.6435
166.9863
21.2167
0.7250
0.2317
0.9011
0.5850
6.3888
Table 2 : Quantitative comparison between GeoShrink and direct step reduction on MiniMax-H3.
NFE
Method
RA-MPJPE ↓
Global MPJPE ↓
RootErr ↓
VelErr ↓
AccErr ↓
uTMR R@3 ↑
Matching ↓
10
Direct
120.618
501.158
465.109
7.325
2.484
0.6260
49.8151
GeoShrink
5.303
26.009
25.024
0.694
0.301
0.7280
46.8595
20
Direct
81.194
360.800
335.515
5.822
1.948
0.7105
47.9181
GeoShrink
1.252
6.004
5.786
0.199
0.102
0.7320
46.7981
30
Direct
51.350
194.661
177.461
4.035
1.363
0.7175
47.2669
GeoShrink
0.394
2.337
2.279
0.093
0.063
0.7300
46.8036
Table 3 : Comparison between GeoShrink and direct step reduction on the HY-Motion.
Diffusion Transformers (DiTs) have achieved state-of-the-art video generation quality, but they incur immense computational cost because standard inference applies the same number of denoising steps uniformly to every token in the sequence. It is well known that human vision ignores vast amounts of redundant motion. Why, then, do our densest models treat every spatiotemporal token with equal priority? In this paper, we introduce Heterogeneous Step Allocation (HSA), a training-free inference algorithm that assigns varying step budgets to different spatiotemporal tokens based on their velocity dynamics. To resolve the resulting sequence-length mismatch without sacrificing global context, HSA introduces a KV-cache synchronization mechanism that allows active tokens to attend to the full sequence while entirely bypassing inactive tokens. Furthermore, we derive a cached Euler update that advances the latent states of skipped tokens in a single operation without additional model evaluations. We evaluate HSA on the Wan-2 and LTX-2 models for both text-to-video (T2V) and image-to-video (I2V) generation. Our results demonstrate that HSA significantly outperforms previous state-of-the-art caching methods and the vanilla Flow Matching baseline, especially at aggressive acceleration regimes (e.g., 50% and 25% runtimes). Crucially, HSA achieves a superior quality-runtime Pareto frontier without the need for expensive offline profiling, robustly preserving structural integrity and generation quality even under tight computational budgets. Project page: https://ernestchu.github.io/hsa
Diffusion models are widely used to generate high-quality images and videos, but their iterative denoising process remains computationally intensive. A growing class of training-free accelerators reduces this cost by reusing cached intermediate features or forecasting future ones. To control draft drift, these methods sometimes compute an exact block feature for verification. Yet the resulting exact feature is typically used only to measure discrepancy or guide a later decision and is then discarded. We find that this previously computed feature can instead be reused for correction. Forwarding it at the verification site resets the local draft residual and reduces downstream feature error. Based on this observation, we introduce FeatFix, a local exact-feature correction method for cached diffusion inference. FeatFix operates at a fixed sparse set of layer--timestep sites. At each selected site, it replaces the complete draft block output with the exact output computed from the same incoming state, avoiding token- or channel-level partial replacement and full-timestep recomputation. Experiments across four image and video backbones show that FeatFix consistently accelerates generation, achieving a speedup of up to 6.70× over Vanilla while maintaining competitive output quality.
Hanshuai Cui, Zhiqing Tang, Zhi Yao +3
School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China · Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai 519087, China
Diffusion Transformers (DiTs) are a dominant backbone for high-fidelity text-to-image generation due to strong scalability and alignment at high resolutions. However, quadratic self-attention over dense spatial tokens leads to high inference latency and limits deployment. We observe that denoising is spatially non-uniform with respect to aesthetic descriptors in the prompt. Regions associated with aesthetic tokens receive concentrated cross-attention and show larger temporal variation, while low-affinity regions evolve smoothly with redundant computation. Based on this insight, we propose AccelAes, a training-free framework that accelerates DiTs through aesthetics-aware spatio-temporal reduction while improving perceptual aesthetics. AccelAes builds AesMask, a one-shot aesthetic focus mask derived from prompt semantics and cross-attention signals. When localized computation is feasible, SkipSparse reallocates computation and guidance to masked regions. We further reduce temporal redundancy using a lightweight step-level prediction cache that periodically replaces full Transformer evaluations. Experiments on representative DiT families show consistent acceleration and improved aesthetics-oriented quality. On Lumina-Next, AccelAes achieves a 2.11× speedup and improves ImageReward by +11.9% over the dense baseline. Code is available at https://github.com/xuanhuayin/AccelAes.
Xuanhua Yin, Chuanzhi Xu, Haoxian Zhou +2
School of Computer Science, The University of Sydney, NSW 2006, Australia