Highly compressed video autoencoders offer an effective way to accelerate video diffusion models, as the Diffusion Transformer (DiT) operates on far fewer tokens. However, such autoencoders are challenging to train, since a higher compression ratio degrades reconstruction quality and recovering it requires more channels, which is known to slow the convergence of the DiT. The compressed latent also differs from the one the DiT was trained on, so the pretrained DiT must be either retrained from scratch or adapted at considerable cost. Compressing the autoencoder the DiT was trained with appears to preserve compatibility, yet optimizing it for reconstruction alone still shifts the latent away from the distribution the DiT has learned. To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT. Specifically, we keep a frozen base latent from the pretrained encoder and learn a residual latent for the information lost under stronger compression, while aligning the compressed latent with the pretrained latent in the feature space of the frozen DiT so that the autoencoder is optimized for generation. We then adapt the DiT with lightweight fine-tuning and asymmetric denoising, where the base is denoised ahead of the residual. GRACE reduces the token count of Wan2.1-I2V-14B by 8x and its latency by 11.1x at 480x832x81, while matching the generation quality of the pretrained pipeline before compression on VBench.
Figures & tables
Figure 1: Teaser. We present GRACE , a novel latent compression technique that fine-tunes pretrained video generation models, such as Wan2.1-14B ( Wan et al., 2025 ) , to generate from substantially fewer latent tokens (e.g., nearly 8× fewer). (Top) Text-to-video (T2V) samples at 736p from Wan2.1-14B before compression and from GRACE, using the same prompt. (Bottom left) The same comparison at 480p. (Bottom right) On VBench-T2V ( Huang et al., 2023 ; Huang et al., 2024 ) , GRACE preserves the generation quality of Wan2.1-14B before compression and achieves higher scores than existing high-compression autoencoders ( HaCohen et al., 2024 ; Zheng et al., 2026 ; He et al., 2026 ) , while generating 11.1× faster at 480p and 15.5× faster at 736p. All models in the plot are evaluated at matched resolutions with 50 sampling steps. The plot reports text-to-video, and the speedups above are measured on image-to-video.
Figure 2: Overall architecture. Stage 1 trains the autoencoder. The frozen encoder E maps the downsampled input to zbase , the residual encoder Eres maps the full-resolution input to zres , and D^ decodes both, while Lalign matches the compressed latent to the pretrained latent inside the frozen DiT. The dashed box on the right of Stage 1 shows a simplified view of the blocks added to Eres and D^ . Stage 2 adapts the DiT to the compressed latent with LoRA, denoising zbase at a lower noise level than zres at every step, so the base is denoised first.
Figure 3: Effect of the dual latent and Lalign . (a) single latent, (b) dual-latent representation without Lalign , and (c) ours, all generated at 480×832×81 . (Left) latent distribution, projected with t-SNE and colored by kernel density, with its uniformity metrics ( Yao et al., 2025 ) and the total VBench ( Huang et al., 2023 ; Huang et al., 2024 ) score below, followed by T2V samples generated after Stage 2. (Right) spatio-temporal structure of the latent, shown as its principal components mapped to RGB under the input frames. More samples are in Figs. I.11 and I.12 .
Figure 4: Denoising order. Each row shows the sampling steps of zbase and zres ( Left ) and the generated frames ( Right ), with τ=1 pure noise and τ=0 clean. (Top) Both parts are denoised at the same noise level. (Bottom) zbase stays δ ahead of zres in schedule time u (Eq. 11 ) at every step.
Autoencoder
Config
AE Training
Latent Tokens
Reconstruction
VBench-I2V
PSNR ↑
SSIM ↑
LPIPS ↓
rFVD ↓
Total ↑
Wan2.1-VAE ( Wan et al., 2025 )
f8t4c16p2
scratch †
32.8k
35.15
0.958
0.016
1.13
87.92
Step-Video-VAE ( Ma et al., 2025 )
f16t8c64p1
scratch ‡
17.2k
33.88
0.950
0.029
3.16
84.05
Video DC-AE ( Zheng et al., 2026 )
f32t4c128p1
scratch
8.2k
34.61
0.956
0.024
3.70
84.94
LTX-VAE ( HaCohen et al., 2024 )
f32t8c128p1
scratch
4.3k
31.97
0.914
0.051
19.53
87.06
Single-latent Baseline
f16t8c32p2
fine-tuned
4.3k
33.76
0.956
0.031
13.11
86.44
Table 1: Video autoencoder comparison at 256×256×81 reconstruction and 480×832×81 generation. We report the VBench-I2V ( Huang et al., 2023 ; Huang et al., 2024 ) total after adapting the same pretrained Wan2.1-I2V-14B ( Wan et al., 2025 ) to every latent under the same budget. Config lists f , t , channel count c , and patch size p , which set the latent token count. Single-latent Baseline is our baseline without the dual latent or the alignment loss. The first row is the pretrained autoencoder before compression, and our method is shaded . Bold marks the best value in each column among the last three rows (4.3k tokens). † Initialized by inflating its own 2D image autoencoder. ‡ Trained first at f8t4 and then extended to f16t8 with additional modules.
Figure 6Figure 7
Figure 7: Convergence during DiT adaptation. With and without Lalign ; solid lines show EMA.
Method
Components
Reconstruction
VBench-T2V
VBench-I2V
[zbase;zres]
Lalign
δ>0
PSNR ↑
LPIPS ↓
rFVD ↓
Quality ↑
Semantic ↑
Total ↑
I2V ↑
Quality ↑
Total ↑
(I)
Wan2.1-14B (pretrained)
–
–
–
35.15
0.016
1.13
85.24
78.70
83.93
95.82
80.01
87.92
(II)
Single-latent Baseline
×
×
×
33.76
0.031
13.11
83.21
77.75
82.12
93.67
79.21
86.44
(III)
(II) + DiT from scratch
×
×
×
33.76
0.031
13.11
74.20
58.40
71.04
53.06
73.94
63.50
(IV)
(II) + dual-latent
✓
×
×
33.10
0.031
13.89
84.77
78.65
83.55
93.96
79.32
86.64
(V)
(IV) + Lalign
✓
✓
×
32.63
0.032
13.53
85.22
82.55
84.68
95.23
79.78
87.51
Table 4: Ablation on design components. Row (III) keeps the autoencoder of (II) but trains the DiT from scratch for the same number of steps. Bold marks the best value in each column among rows (II) and (IV)–(VI).
Alignment Target
Reconstruction
VBench-T2V
PSNR ↑
rFVD ↓
Quality ↑
Semantic ↑
Total ↑
None
33.10
13.89
84.77
78.65
83.55
V-JEPA 2.1 features
32.61
13.93
84.98
80.04
83.99
Pretrained DiT features ( Ours )
32.63
13.53
85.22
82.55
84.68
Table 5: Alignment target. Asymmetric denoising is disabled in all rows.
Denoising Order
Offset
VBench-I2V
I2V ↑
Quality ↑
Total ↑
τbase=τres
δ=0
95.23
79.78
87.51
τbase>τres
δ<0
94.61
79.72
87.17
τbase<τres (Ours)
δ>0
95.48
80.31
87.90
Table 6: Ablation on denoising order.
Figure A.1: Detailed autoencoder architecture. Eres adds a downsampling stage between its middle blocks and head, consisting of residual blocks and a strided causal convolution, paired with a parameter-free shortcut that folds space and time into channels.
Hyperparameter
Phase 1
Phase 2
Architecture
pretrained autoencoder
Wan2.1-VAE
Wan2.1-VAE
(C,C′)
(16,16)
(16,16)
(rs,rt)
(2,2)
(2,2)
Training setup
input shape
256×256×81
512×512×81
trained modules
Eres , D^
D^
optimizer
AdamW
AdamW
Table A.1: Stage 1 hyperparameters. All settings follow the training recipe of the pretrained autoencoder.
Hyperparameter
I2V
T2V
Architecture
pretrained DiT
Wan2.1-I2V-14B
Wan2.1-T2V-14B
input dim
72
32
hidden dim
5120
5120
blocks
40
40
num. heads
40
40
patch size
2
2
Table B.1: Stage 2 hyperparameters. Offset and shift are sampled from a range in training, fixed at inference.
Table D.1: Per-dimension VBench ( Huang et al., 2023 ; Huang et al., 2024 ) scores at 480×832×81 . Single-latent Baseline is our baseline without the dual latent or the alignment loss. The best score in each row is in bold. † Reported for reference only, as the official VBench-I2V quality score does not include this dimension.
Table D.2: Per-dimension VBench ( Huang et al., 2023 ; Huang et al., 2024 ) scores at 736×1280×81 , following Tab. D.1 . The best score in each row is in bold. ‡ Measured with spatial tiling in the VAE encoder to avoid running out of memory when encoding the conditioning image at this resolution.
Task
Aspect
vs. Wan2.1-14B ( Wan et al., 2025 ) (%)
vs. DC-Gen ( He et al., 2026 ) (%)
Ours
Equal
Wan2.1-14B
Ours
Equal
DC-Gen
T2V
Visual quality
49.4
15.4
35.3
71.2
11.5
17.3
Temporal consistency
46.8
21.2
32.1
64.7
21.2
14.1
Text alignment
52.6
21.8
25.6
62.2
23.1
14.7
I2V
Visual quality
23.7
32.7
43.6
60.3
16.7
23.1
Temporal consistency
29.5
25.6
44.9
64.1
11.5
24.4
Table D.3: Human evaluation. Participants compared each pair of videos without knowing which model produced which. Each cell gives the percentage of votes preferring our model (Ours), rating both videos about equal (Equal), or preferring the baseline named in the column.
Figure D.1: User study interface for T2V samples.
Figure D.2: User study interface for I2V samples.
First frame
PSNR ↑
SSIM ↑
LPIPS ↓
×
31.97
0.928
0.038
✓
32.63
0.930
0.032
Table E.1: Ablation on first-frame conditioning.
Depth
Reconstruction
VBench-I2V
PSNR ↑
SSIM ↑
LPIPS ↓
I2V ↑
Quality ↑
Total ↑
10
32.63
0.930
0.032
95.48
80.31
87.90
40
32.21
0.929
0.038
95.70
80.19
87.94
Table E.2: Ablation on alignment depth.
Figure F.1: Reconstruction under latent noise. We noise zbase or zres at level τ , leave the other unchanged, and reconstruct. (a) Reconstruction quality on Panda-70M ( Chen et al., 2024 ) at 480×832×81 as τ grows, measured in PSNR and LPIPS. (b) Reconstructed frames at two levels of τ , shown below the input and the noise-free reconstruction, with the first block noising zbase and the second noising zres .
Figure F.2: PCA visualization of the base and residual latents. Principal Component Analysis (PCA) of each part at f16t8p2, encoded from a 480×832×81 video, for (a) zbase , (b) zres without Lalign , and (c) zres with Lalign . Since E is frozen, zbase is the same in both settings.
Figure F.3: Convergence of the base and the residual. Flow matching loss on the base channels (left) and the residual channels (right) during DiT adaptation, for the dual latent with and without Lalign . Each part is standardized with its own statistics. Faint curves show the raw loss and solid curves its EMA.
T2V
I2V
Wan2.1-14B
Ours
Wan2.1-14B
Ours
Training
Stage 1, full autoencoder
–
6.9
–
6.9
Stage 1, decoder-only (+EMA)
–
1.6
–
1.6
Stage 2
–
30
–
30
Total (H200 GPU days)
–
38.5
–
38.5
Inference latency (s) ↓
851.5
75.8
863.2
77.7
Table G.1: Computational cost. Training in H200 GPU days, and inference at 480×832×81 on a single A100 with 50 sampling steps.
Figure I.1: Additional image-to-video results. Samples from VBench-I2V ( Huang et al., 2023 ; Huang et al., 2024 ) at 480×832×81 , generated by the pretrained Wan2.1-14B ( Wan et al., 2025 ) at f8t4p2, ours at f16t8p2, and DC-Gen ( He et al., 2026 ) at f32t4p1, from the same conditioning frame and prompt.
Figure I.2: Additional image-to-video results (continued). Samples from VBench-I2V ( Huang et al., 2023 ; Huang et al., 2024 ) at 480×832×81 , generated by the pretrained Wan2.1-14B ( Wan et al., 2025 ) at f8t4p2, ours at f16t8p2, and DC-Gen ( He et al., 2026 ) at f32t4p1, from the same conditioning frame and prompt.
Figure I.3: Additional text-to-video results. Samples from VBench-T2V ( Huang et al., 2023 ; Huang et al., 2024 ) at 480×832×81 , generated by the pretrained Wan2.1-14B ( Wan et al., 2025 ) at f8t4p2, ours at f16t8p2, and DC-Gen ( He et al., 2026 ) at f32t4p1, from the same prompt.
Figure I.4: Additional text-to-video results (continued). Samples from VBench-T2V ( Huang et al., 2023 ; Huang et al., 2024 ) at 480×832×81 , generated by the pretrained Wan2.1-14B ( Wan et al., 2025 ) at f8t4p2, ours at f16t8p2, and DC-Gen ( He et al., 2026 ) at f32t4p1, from the same prompt.
Figure I.5: Additional high-resolution text-to-video results. Samples at 736×1280×81 , generated by the pretrained Wan2.1-14B ( Wan et al., 2025 ) at f8t4p2, ours at f16t8p2, and DC-Gen ( He et al., 2026 ) at f32t4p1, from the same prompt.
Figure I.6: Additional high-resolution text-to-video results (continued). Samples at 736×1280×81 , generated by the pretrained Wan2.1-14B ( Wan et al., 2025 ) at f8t4p2, ours at f16t8p2, and DC-Gen ( He et al., 2026 ) at f32t4p1, from the same prompt.
Figure I.7: Additional high-resolution image-to-video results. Samples at 736×1280×81 , generated by the pretrained Wan2.1-14B ( Wan et al., 2025 ) at f8t4p2, ours at f16t8p2, and LTX-Video ( HaCohen et al., 2024 ) 0.9.7 at f32t8p1, from the same conditioning frame and prompt. The motion of LTX-Video tends to come from a global zoom or a slow camera movement over the conditioning frame, while the scene stays static.
Figure I.8: Additional high-resolution image-to-video results (continued). Samples at 736×1280×81 , generated by the pretrained Wan2.1-14B ( Wan et al., 2025 ) at f8t4p2, ours at f16t8p2, and LTX-Video ( HaCohen et al., 2024 ) 0.9.7 at f32t8p1, from the same conditioning frame and prompt. The motion of LTX-Video tends to come from a global zoom or a slow camera movement over the conditioning frame, while the scene stays static.
Figure I.9: Additional high-resolution image-to-video results (continued). Samples at 736×1280×81 , generated by the pretrained Wan2.1-14B ( Wan et al., 2025 ) at f8t4p2, ours at f16t8p2, and LTX-Video ( HaCohen et al., 2024 ) 0.9.7 at f32t8p1, from the same conditioning frame and prompt. The motion of LTX-Video tends to come from a global zoom or a slow camera movement over the conditioning frame, while the scene stays static.
Figure I.10: Generation in various styles. Text-to-video samples from ours at 736×1280×81 , generated from a single prompt with a different style suffix in each row.
Figure I.11: Additional latent PCA visualizations. Principal components of the full C+C′ latent at f16t8p2, encoded from 480×832×81 videos, for (a) a single encoder, (b) ours without Lalign , and (c) ours, with the input frames on top. The three leading components are computed per clip and mapped to RGB, so colors are not comparable across panels.
Figure I.12: Additional latent PCA visualizations (continued). Principal components of the full C+C′ latent at f16t8p2, encoded from 480×832×81 videos, for (a) a single encoder, (b) ours without Lalign , and (c) ours, with the input frames on top. The three leading components are computed per clip and mapped to RGB, so colors are not comparable across panels.
Diffusion models provide strong generative capabilities for video compression at ultra-low bitrates. Existing diffusion-based video codecs adapt base models originally developed for text-conditioned generation, whereas diffusion models designed and trained specifically for compression remain unexplored. To fill this gap, we introduce our Generative Video Codec (GenVC), built on a video diffusion model trained from scratch for compression. To our knowledge, this is the first compression-oriented video diffusion model. We realize this model directly in pixel space with a global-to-local hierarchy that recovers fine spatio-temporal details, enabling high-quality generative reconstruction from compressed representations. To accelerate inference, we distill the multi-step model into one step using distribution matching distillation (DMD). Applying DMD directly, however, drives the student toward motion-stalled reconstructions. We trace this to a teacher-side guidance failure: once student-induced perturbations leave the frozen teacher's training region, its guidance can become misleading, causing DMD updates to reinforce rather than correct the student drift. To break the resulting feedback loop, we propose Adaptive Score Distillation, which gates DMD updates according to their alignment with the ground-truth direction, enabling high-quality reconstruction with coherent motion. Experimental results show that GenVC achieves state-of-the-art perceptual quality at ultra-low bitrates, with average bitrate savings of 62.5% at matched LPIPS and 71.3% at matched FID over GLVC. Unlike prior codecs that inherit billion-scale pretrained backbones, our diffusion model has only 478.0M parameters and decodes 1080p video in a single step at 15.1 fps on an A100 GPU.
Naifu Xue, Zhaoyang Jia, Haosen Li +7
Communication University of China, Beijing, China · Microsoft Research Asia · University of Science and Technology of China, Hefei, China +1
Video Diffusion Transformers (DiTs) generate high-quality videos but demand substantial compute due to wide blocks, deep architectures, and iterative sampling. Recent methods reduce cost by compressing width, depth, or sampling steps, but typically commit to a fixed architecture that cannot adapt to individual inputs or denoising stages. We propose PARE (Pruning and Adaptive Routing for Efficient video generation), which jointly compresses width and depth with structure-aware pruning and input-adaptive routing. For width, we observe that attention heads specialize into spatial and temporal roles, and design importance scoring that accounts for this distinction to prevent motion-critical temporal heads from being pruned prematurely. For depth, we train a lightweight router conditioned on denoising timestep and visual content to dynamically select which blocks to execute at each step, enabling per-input compute adaptation rather than static block removal. A progressive pipeline first recovers width-pruned quality via distillation, then jointly optimizes the student and router to decouple the two learning objectives. Experiments on Wan2.1-14B for both image-to-video and text-to-video generation show that PARE substantially reduces per-step computation while preserving quality across VBench dimensions, and composes with step distillation for further acceleration.
Yutong Wang, Yunke Wang, Tianfan Xue +4
The University of Sydney · The Chinese University of Hong Kong · Shanghai AI Laboratory
Video generation, while capable of generating realistic videos, is computationally expensive and slow, prohibiting real-time applications. In this paper, we observe that video latents encoded via an autoencoder under the Latent Diffusion Model (LDM) framework contain redundancy along the temporal axis. Analogous to how traditional video compression algorithms avoid transmitting redundant frame data, we propose the Latent Inter-frame Pruning framework to prune (skip the re-computation of) duplicated latent patches, thereby reducing computational burden and increasing throughput. However, direct pruning results in visual artifacts due to the discrepancy between full-sequence training and pruned inference. To resolve these artifacts, we propose an Attention Recovery mechanism to bridge the train-inference gap. With our proposed method, we increase video editing throughput by 1.44×, achieving 12.44 FPS on an NVIDIA RTX 6000 while maintaining video quality. We hope our work inspires further research into integrating traditional video compression methods with modern video generation pipelines. This work is a preliminary work on Training-free Latent Inter-Frame Pruning with Attention Recovery.
Dennis Menn, Chih-Hsien Chou
The University of Texas at Austin · Futurewei Technologies, Inc.