Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, diluting learning signals for informative content and disrupting pretrained priors. Existing sparse attention methods either target training-free acceleration or overlook the unique structure of joint video-audio data, where cross-modal interactions are inherently concentrated around sound-producing regions. To address this, we propose Prism, a dynamic sparse attention framework for natively training joint video-audio generation models at 2K. In particular, Prism organizes the token sequence into spatiotemporal macro-zones, enabling the attention structure to adapt to local content. For each zone, it estimates local information structure via video feature variance along the channel and feature norms from the audio-to-video cross-attention, jointly capturing how visual content varies directionally and how strongly audio influences each visual region. Based on these signals, Prism dynamically assigns a tailored block shape to each zone, applying finer partitioning along axes of rapid visual content variation and strong audio-visual coupling. This encourages tokens within each block to remain semantically coherent, allowing block-level features to capture both visual content and joint video-audio interaction patterns. Prism further adopts a hybrid block selection strategy to dynamically determine per-query sparsity. Experiments show that Prism achieves 2.5× training speedup compared to full attention, while surpassing it in generation quality.
Figures & tables
Figure 1 : Videos generated by Prism, showing its power to natively synthesize 2K video-audio. LTX-2.3 [ 12 ] first generates at 720p and then upsamples (SR) to 2K, and MOVA [ 59 ] performs native 2K generation by training with full attention on 2K videos.
Figure 2 : Architecture of Prism. Prism partitions the token sequence into spatiotemporal macro-zones, dynamically assigns each zone a tailored block shape, and applies hybrid Top- k /Top- p block-sparse attention to focus computation on informative interactions.
Figure 3 : Qualitative Comparisons with previous open-source methods. Please refer to the demo video for audio. More results are in the Appx. A.7 .
Model
AQ ↑
DD ↑
TF ↑
ID ↑
PQ ↑
CU ↑
DeSync ↓
Sync-D ↓
Sync-C ↑
cpCER ↓
MUSIQ ↑
MANIQA ↑
MotionQ ↑
Ovi [ 39 ]
0.38
0.35
0.891
0.85
6.68
5.92
1.12
8.28
5.21
0.468
51.43
0.326
0.42
LTX-2.3 [ 12 ]
0.48
0.41
0.943
0.91
7.05
6.83
0.95
7.62
5.92
0.382
55.60
0.403
0.66
MagiHuman [ 54 ]
0.40
0.37
0.912
0.87
6.72
6.14
1.08
8.05
5.45
0.420
52.68
0.337
0.58
MOVA [ 59 ]
0.42
0.39
0.903
0.88
6.80
6.26
1.05
7.96
5.62
0.374
53.21
0.363
0.53
Ours
0.61
0.52
0.982
0.94
7.69
7.34
0.63
6.74
7.27
0.187
62.25
0.438
0.89
Table 1: Quantitative comparisons with previous open-source methods on 2K-Bench.
Category
Model
AQ ↑
TF ↑
PQ ↑
CU ↑
DeSync ↓
cpCER ↓
MANIQA ↑
MotionQ ↑
Base Sparsity
Training-Free
SVG-2 [ 83 ]
0.35
0.861
6.23
5.71
1.28
0.462
0.318
0.38
71%
Sol-Attn [ 29 ]
0.40
0.892
6.64
6.08
1.11
0.398
0.349
0.47
85%
Trainable
Full Attn
0.42
0.903
6.80
6.26
1.05
0.374
0.363
0.53
0%
Block Sparse Attn (BSA) [ 58 ]
0.33
0.842
6.12
5.58
1.36
0.487
0.306
0.34
90%
VSA [ 93 ]
0.39
0.878
6.53
6.04
1.08
0.391
0.347
0.46
90%
SSTA [ 77 ]
0.43
0.911
6.82
6.31
0.97
0.356
0.368
0.54
85%
Table 2: Ablation study on different sparse attention for native 2K joint video-audio training. Training-free methods are applied at inference to the Full Attn. Trainable methods replace original video self-attention during native 2K training. All competitors use their optimal sparsity settings.
Appendix figures & tables40 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Architecture of the dual-branch DiT of Prism.
Model
Audio-Aes ↑
T-V Align ↑
T-A Align ↑
A-V Align ↑
DeSync ↓
Visual Realism ↑
Audio Realism ↑
Audio QA ↑
Video QA ↑
Ovi [ 39 ]
3.494
0.226
0.307
0.168
0.983
4.963
4.593
0.679
0.725
LTX-2.3 [ 12 ]
3.683
0.225
0.277
0.224
0.871
4.958
4.570
0.751
0.743
MagiHuman [ 54 ]
3.452
0.196
0.290
0.192
0.694
4.134
4.248
0.610
0.686
MOVA [ 59 ]
3.318
0.219
0.386
0.243
0.942
4.969
4.537
0.783
0.716
Ours
3.691
0.228
0.394
0.262
0.691
4.973
4.608
0.786
0.772
Appendix
Table 3: Quantitative comparisons with previous open-source methods on VABench. We use their original evaluation metrics for fair comparison.
Model
IS ↑
DNSMOS ↑
DeSync ↓
IB-Score ↑
LSE-D ↓
LSE-C ↑
cpCER ↓
Ovi [ 39 ]
3.680
3.516
0.515
0.190
7.468
6.378
0.436
LTX-2.3 [ 12 ]
3.326
3.708
0.342
0.312
8.063
6.447
0.197
MagiHuman [ 54 ]
3.075
3.420
0.561
0.238
11.637
2.566
0.487
MOVA [ 59 ]
3.814
3.751
0.370
0.297
7.094
7.452
0.218
Ours
4.103
3.866
0.314
0.335
6.827
7.689
0.128
Appendix
Table 4: Quantitative comparisons with previous open-source methods on MOVA-Bench. We use their original evaluation metrics for a fair comparison.
Category
Model
AQ ↑
TF ↑
PQ ↑
CU ↑
DeSync ↓
cpCER ↓
MANIQA ↑
MotionQ ↑
Training Time ↓
Infer GPU Mem ↓
Training-Free
MOVA-720p → 2K [ 59 ]
0.36
0.871
6.37
5.79
1.21
0.443
0.291
0.37
-
72.8G
UltraGen [ 17 ]
0.43
0.908
6.71
6.15
1.06
0.386
0.347
0.46
-
48.0G
Natively Training
Full Attn
0.42
0.903
6.80
6.26
1.05
0.374
0.363
0.53
26.5min
72.8G
SpargeAttn2 [ 92 ]
0.47
0.928
7.10
6.55
0.90
0.319
0.384
0.61
8.4min
36.7G
LUVE [ 95 ]
0.50
0.926
7.19
6.68
0.85
0.312
0.402
0.62
18.7min
62.0G
PyramidFlow [ 25 ]
0.51
0.944
7.08
6.59
0.84
0.295
0.391
0.70
18.3min
65.8G
Appendix
Table 5: Comparison with native 2K training/inference methods. Training Time is the per-step training time. GPU Mem is per-GPU memory during 4-GPU parallel inference, as the token count at 2K (10s, FPS=24) reaches approximately 870K, making single-GPU inference prohibitively expensive. Each competitor is trained until loss convergence.
Model
AQ ↑
TF ↑
PQ ↑
CU ↑
DeSync ↓
cpCER ↓
MANIQA ↑
MotionQ ↑
LTX-2.3 [ 12 ]
0.48
0.943
7.05
6.83
0.95
0.382
0.403
0.66
LTX-2.3 + Prism
0.63
0.986
7.74
7.41
0.66
0.208
0.452
0.86
Ovi [ 39 ]
0.38
0.891
6.68
5.92
1.12
0.468
0.326
0.42
Ovi + Prism
0.53
0.952
7.38
6.68
0.79
0.251
0.387
0.71
Appendix
Table 6: Ablation on different DiT backbones.
Setting
AQ ↑
TF ↑
PQ ↑
CU ↑
DeSync ↓
cpCER ↓
MANIQA ↑
MotionQ ↑
Training Time ↓
Only Top- k (75%)
0.50
0.937
7.02
6.58
0.87
0.294
0.386
0.64
14.3min
Only Top- k (85%)
0.56
0.961
7.41
6.97
0.74
0.243
0.417
0.76
11.8min
Only Top- k (95%)
0.48
0.926
6.89
6.43
0.92
0.312
0.379
0.61
9.8min
Only Top- p (0.1)
0.29
0.831
5.87
5.32
1.48
0.513
0.291
0.26
9.2min
Only Top- p (0.2)
0.36
0.862
6.28
5.74
1.24
0.441
0.321
0.37
9.9min
Only Top- p (0.3)
0.43
0.894
6.71
6.18
1.02
0.378
0.356
0.49
11.1min
Appendix
Table 7: Ablation study on sparsity strategies. Top- k ( x %) applies x sparsity per query. Top- p ( y ) selects the smallest set of key blocks whose cumulative attention weight exceeds y .
Model
AQ ↑
TF ↑
PQ ↑
CU ↑
DeSync ↓
cpCER ↓
MANIQA ↑
MotionQ ↑
Kling3.0 [ 28 ]
0.65
0.986
7.93
7.61
0.52
0.143
0.456
0.91
MiniMax-H3 [ 43 ]
0.66
0.983
8.01
7.72
0.54
0.148
0.449
0.90
Seedance2.5 [ 3 ]
0.71
0.989
8.27
7.98
0.44
0.112
0.471
0.96
Wan3.0 [ 1 ]
0.74
0.991
8.38
8.11
0.40
0.103
0.482
0.96
Ours (16B)
0.61
0.982
7.69
7.34
0.63
0.187
0.438
0.89
Appendix
Table 8: Comparison with industry-leading models on 2K-Bench. All commercial models are evaluated using their native 2K/1080p generation APIs.
Category
Setting
AQ ↑
TF ↑
PQ ↑
CU ↑
DeSync ↓
cpCER ↓
MANIQA ↑
MotionQ ↑
Training Time ↓
Infer GPU Mem ↓
Block Shape
Random Shape
0.40
0.894
6.73
6.18
1.03
0.372
0.339
0.46
8.7min
36.2G
Fixed 43
0.55
0.961
7.43
6.99
0.72
0.234
0.413
0.77
12.2min
42.3G
Fixed 83
0.43
0.906
6.86
6.32
1.02
0.367
0.352
0.51
6.8min
33.4G
Only C64
0.57
0.965
7.49
7.08
0.70
0.221
0.419
0.80
13.5min
44.1G
Only C128
0.54
0.953
7.36
6.91
0.76
0.251
0.406
0.74
10.9min
39.3G
Only C256
0.49
0.934
7.17
6.64
0.86
0.298
0.388
0.64
8.4min
35.2G
Appendix
Table 9: Ablation study on dynamic block shape. rv,d is the video channel-wise variance guidance. aˉ is the audio coupling gate. v^a,d is the audio directional variance. Global σv,d2 ( w/o ch.) replaces the per-channel variance in Eq. 1 with global feature variance computed across all channels jointly. σq,d2 / σk,d2 replace value features with query/key features for variance computation.
Category
Setting
AQ ↑
TF ↑
PQ ↑
CU ↑
DeSync ↓
cpCER ↓
MANIQA ↑
MotionQ ↑
Training Time ↓
Infer GPU Mem ↓
Shape Mapping
α=1/3
0.56
0.958
7.41
6.97
0.72
0.236
0.411
0.76
10.6min
38.0G
α=1
0.58
0.965
7.53
7.14
0.68
0.214
0.421
0.82
10.6min
38.0G
α=2
0.54
0.949
7.31
6.86
0.76
0.248
0.402
0.71
10.6min
38.0G
Softmax
0.57
0.964
7.46
7.03
0.69
0.227
0.414
0.77
10.6min
38.0G
Rank
0.55
0.963
7.43
6.98
0.73
0.239
0.417
0.75
10.6min
38.0G
Entropy
0.58
0.971
7.49
7.08
0.71
0.223
0.423
0.79
10.6min
38.0G
Appendix
Table 10: Ablation study on dynamic block shape. Shape Mapping ablates the quantization in Eq. 7 , where α controls bd∗∝gd−α . Heads/Layers ablates whether block shape is computed independently per attention head and per layer. τ128/τ256 ablates the thresholds for block shape assignment.
C64 (%)
C128 (%)
C256 (%)
Index
All
444
824
842
248
284
428
482
All
288
828
882
448
484
844
All
488
848
884
Layer (HN)
10
40.9
36.8
2.3
0.6
0.2
0.4
0.4
0.2
59.1
0.4
0.2
0.2
16.9
15.0
26.4
0.0
–
–
–
20
27.0
20.9
3.3
1.5
0.3
0.4
0.2
0.4
65.2
0.3
0.6
0.3
18.6
14.1
31.3
7.8
0.5
4.3
3.0
24
25.2
18.1
4.2
1.5
0.4
0.4
0.4
0.2
28.8
0.2
1.4
0.4
7.1
3.7
16.0
46.0
11.3
21.6
13.1
32
21.4
12.6
4.6
1.0
0.2
0.2
2.5
0.3
22.1
0.4
6.3
0.7
2.4
1.8
10.5
56.5
8.1
36.6
11.8
40
16.6
8.3
3.6
0.5
0.3
0.3
3.2
0.4
20.5
0.4
10.3
0.4
1.5
0.4
7.5
62.9
5.7
49.7
7.5
Appendix
Table 11: Block shape proportion across DiT layers, denoised steps, and attention heads for the same 2K video-audio clip (a two-speaker dialogue with a sounding hand action). Our MOVA [ 59 ] backbone follows the Wan2.2-style MoE design [ 71 ] , with a high-noise (HN) and a low-noise (LN) expert of 40 layers and 40 heads each. Shapes are denoted bTbHbW (e.g., 824 means bT=8,bH=2,bW=4 )
Res.
Method
AQ ↑
TF ↑
PQ ↑
CU ↑
DeSync ↓
cpCER ↓
MANIQA ↑
MotionQ ↑
Training Time ↓
Infer GPU Mem ↓
480P
Full Attn
0.56
0.968
7.42
6.95
0.71
0.234
0.348
0.59
0.6min
10.4G
VMoBA
0.52
0.961
7.26
6.78
0.77
0.258
0.336
0.54
0.5min
7.6G
SpargeAttn2
0.49
0.951
7.16
6.68
0.84
0.276
0.326
0.49
0.3min
5.1G
PyramidFlow
0.50
0.964
7.11
6.63
0.75
0.249
0.323
0.47
0.7min
11.4G
w/o Video Guidance
0.48
0.953
7.19
6.71
0.81
0.244
0.328
0.50
0.3min
4.7G
w/o Audio Guidance
0.53
0.963
7.32
6.85
0.76
0.268
0.339
0.55
0.5min
5.3G
Appendix
Table 12: Ablation study on different training resolutions. GPU Mem is per-GPU memory during 4-GPU parallel inference. Each model generates at its natively training resolution (480p/720p/1080p rows), with all outputs bicubically resized to 1080p for comparable resolution-sensitive AQ/MANIQA across rows. The 2K row is additionally evaluated at its native 2K resolution.
Model
AQ ↑
TF ↑
PQ ↑
CU ↑
DeSync ↓
cpCER ↓
MANIQA ↑
MotionQ ↑
Inference Speed ↓
Infer GPU Mem ↓
LTX-2.3 [ 12 ]
0.48
0.943
7.05
6.83
0.95
0.382
0.403
0.66
347s
78.4G
MagiHuman [ 54 ]
0.40
0.912
6.72
6.14
1.08
0.420
0.337
0.58
386s
74.2G
Ovi [ 39 ]
0.38
0.891
6.68
5.92
1.12
0.468
0.326
0.42
724s
68.7G
MOVA (Full Attn) [ 59 ]
0.42
0.903
6.80
6.26
1.05
0.374
0.363
0.53
1091s
72.8G
VMoBA [ 78 ]
0.46
0.924
7.04
6.52
0.91
0.328
0.381
0.59
483s
43.6G
SpargeAttn2 [ 92 ]
0.47
0.928
7.10
6.55
0.90
0.319
0.384
0.61
248s
36.7G
Appendix
Table 13: Inference speed and GPU memory comparison. GPU Mem is per-GPU memory during 4-GPU parallel inference, as the token count at 2K (10s, FPS=24) reaches approximately 870K, making single-GPU inference prohibitively expensive.
Prism vs
V-A ↑
A-A ↑
V-Q ↑
A-Q ↑
M-Q ↑
V-A-S ↑
Ovi [ 39 ]
93.4%
92.6%
96.9%
96.2%
98.1%
91.7%
LTX-2.3 [ 12 ]
91.2%
88.7%
94.3%
93.5%
95.7%
87.4%
MagiHuman [ 54 ]
92.8%
90.3%
95.7%
94.1%
97.6%
89.5%
MOVA [ 59 ]
92.0%
91.3%
96.5%
95.8%
97.6%
91.3%
Appendix
Table 14: User preference of Prism compared to other competitors. A higher score indicates users prefer more to our model.
Figure 5 : The system prompt used for Gemini-based motion quality evaluation.
Figure 6 : Examples of 2K-Bench.
Figure 7 : More comparison results (1/4). Please refer to the demo video for audio.
Figure 8 : More comparison results (2/4). Please refer to the demo video for audio.
Figure 9 : More comparison results (3/4). Please refer to the demo video for audio.
Figure 10 : More comparison results (4/4). Please refer to the demo video for audio.
Figure 11 : Qualitative Comparisons with commercial models (1/2).
Figure 12 : Qualitative Comparisons with commercial models (2/2)
Figure 13 : Ablation study on different sparse attention methods.
Figure 14 : Ablation study on different native high-resolution training methods.
Figure 15 : Ablation study on block shape.
Figure 16 : Ablation study on guidance.
Figure 17 : Ablation study on different dynamic block shape mapping functions.
Figure 18 : Ablation study on different DiT layers, denoised steps, and attention heads.
Figure 19 : Ablation study on different training resolutions. Please refer to the demo video for a clear comparison. VG and AG refer to video channel-wise variance guidance and audio-to-video cross-attention norm guidance.
Figure 20 : Visualization of video training loss.
Figure 21 : Visualization of audio training loss.
Figure 22 : Visualization of our attention maps.
Figure 23 : Visualization of dynamic block shapes.
Figure 24 : Synthesized video-audio content involving multiple speakers. Please refer to the demo video for audio.
Figure 25 : Complex scene results (1/5). Please refer to the demo video for audio.
Figure 26 : Complex scene results (2/5). Please refer to the demo video for audio.
Figure 27 : Complex scene results (3/5). Please refer to the demo video for audio.
Figure 28 : Complex scene results (4/5). Please refer to the demo video for audio.
Figure 29 : Complex scene results (5/5). Please refer to the demo video for audio.