Organizations: The Hong Kong University of Science and Technology (Guangzhou) · Griffith University · Data61, CSIRO · Artificial Intelligence Lab, Institute of Deep Perception Technology, JITRI
Diffusion transformers incur substantial inference cost through repeated model evaluations along a sampling trajectory. We introduce GeoShrink, a training-free acceleration method that retains the original solver grid while evaluating the model only at a prescribed set of anchors. At skipped stages, GeoShrink predicts the solver-facing output by adding a geometrically retained fraction of the latest observed innovation to the most recent exact output. We derive this rule from chordal tangent transport and round-trip line projection, and establish a geometric anchor-spacing principle that minimizes the largest adjacent gap expansion under fixed coverage and first span. The analysis characterizes the geometric closure and propagation of prediction errors without assuming access to future model outputs. Experiments cover image, video, motion, and audio generation, together with adapted 3D backends. At approximately 5× acceleration, GeoShrink improves FLUX PSNR by 3.10 dB over the strongest listed baseline. On HunyuanVideo, it achieves a reported 4.99× speedup and improves ChronoMagic-Bench-150 PSNR by 5.44 dB over the strongest listed fidelity baseline. Comparisons at fixed evaluation budgets further show substantial gains on motion, audio, music, and 3D generation.
Figures & tables
Figure 1 : Qualitative comparison of GeoShrink. Direct NFE reduction (left) vs. GeoShrink (right) under the same model-evaluation budget. GeoShrink better preserves visual structure and temporal evolution under aggressive acceleration.
Figure 2 : Motivation for adaptive innovation retention. From left to right, we show the full-NFE reference, fixed-coefficient predictions with ω∈{0,0.5,1.5,2} , and GeoShrink. Small ω underuses the latest innovation, while aggressive reuse ( ω>1 ) progressively introduces color, contrast, and appearance drift. GeoShrink instead adapts ωt from observable trajectory geometry and more closely preserves the full-NFE reference.
Figure 3 : Overview of GeoShrink. (a) Previous methods rely on internal feature caching, reusing or forecasting intermediate DiT representations across denoising steps. (b) GeoShrink operates at the solver interface, allocating 10 exact model queries over a 50-step grid through geometric anchoring. At skipped stages, it predicts the solver-facing field from the three latest exact outputs while preserving the original solver and its grid. (c) The retention coefficient ωt=(1+ρt2sin2θ)−1 combines the acute angle θ between successive innovations with the normalized midpoint forecast horizon ρt . Stronger turning or a longer horizon reduces the retained innovation in vt=vt(1)+ωtΔv .
vt←Vθ(xt,σt,c),t(3)←t(2),t(2)←t(1),t(1)←t
Algorithm 1 GeoShrink
Figure 4 : Qualitative comparison on text-to-image generation. GeoShrink better preserves the full-sampling reference than competing acceleration methods at comparable speedups.
Figure 5 : Qualitative comparison on text-to-video generation. GeoShrink preserves spatial details and temporal consistency more faithfully under aggressive acceleration.
Model
Benchmark
Method
Setting
Speedup
PSNR ↑
SSIM ↑
LPIPS ↓
Task Score ↑
SD3.5 Medium
Pick-a-Pic
SpeCa
τ0=7,β=0.3
2.40 ×
20.759
0.7743
0.2704
0.7873
TaylorSeer
N=3,O=1
2.49 ×
18.493
0.7454
0.2918
0.7815
TeaCache
l=0.4
2.66 ×
20.872
0.7616
0.2942
0.7671
ToCa
N=10,R=90%
2.09 ×
20.431
0.7057
0.3111
0.7662
ResilPhase
N=3,O=1
2.76 ×
18.483
0.7451
0.2921
0.7813
GeoShrink
B=10
2.80 ×
24.744
0.8450
0.2078
0.7924
Table 1 : Quantitative comparison of acceleration methods across image and video generation models. For the task-specific metric, we report ImageReward on Pick-a-Pic, MTScore on ChronoMagic-Bench-150, and VBench Score on VBench.
NFE
Method
Visual
Audio
CLIP ↑
FVD ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Mel Cos ↑
LogMel L1 ↓
SI-SDR ↑
10
Direct
33.7164
166.0071
17.4910
0.6190
0.3632
0.7730
1.1660
-0.2793
GeoShrink
33.8711
164.5004
24.2122
0.7883
0.1678
0.8528
0.6998
1.5133
15
Direct
33.6598
166.8756
19.2673
0.6732
0.2914
0.8537
0.8240
3.2408
GeoShrink
33.7885
166.1741
29.1320
0.8810
0.0818
0.9595
0.3645
7.6285
20
Direct
33.6435
166.9863
21.2167
0.7250
0.2317
0.9011
0.5850
6.3888
Table 2 : Quantitative comparison between GeoShrink and direct step reduction on MiniMax-H3.
NFE
Method
RA-MPJPE ↓
Global MPJPE ↓
RootErr ↓
VelErr ↓
AccErr ↓
uTMR R@3 ↑
Matching ↓
10
Direct
120.618
501.158
465.109
7.325
2.484
0.6260
49.8151
GeoShrink
5.303
26.009
25.024
0.694
0.301
0.7280
46.8595
20
Direct
81.194
360.800
335.515
5.822
1.948
0.7105
47.9181
GeoShrink
1.252
6.004
5.786
0.199
0.102
0.7320
46.7981
30
Direct
51.350
194.661
177.461
4.035
1.363
0.7175
47.2669
GeoShrink
0.394
2.337
2.279
0.093
0.063
0.7300
46.8036
Table 3 : Comparison between GeoShrink and direct step reduction on the HY-Motion.
Jul 30, 2026·Hanshuai Cui, Zhiqing Tang, Zhi Yao +3DrafterCorrection
School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China · Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai 519087, China