Audio-video generation using heterogeneous multimodal references has emerged as a new challenge, requiring both compositional control over generation and grounded understanding of multimodal context. In this paper, we introduce ORAV Bench for Omni Reference Audio-Video Generation, comprising 380 task instances with 2-10 references, 9 semantic roles, and 30 role compositions. Instructions specify the relationships among references; the media supply the identities, dynamics, and audio characteristics to be realized. To evaluate these open-ended outputs, we develop a reference-aware pairwise protocol that prepares visual and auditory evidence, compares the intended contribution of each reference, and checks the overall verdict in both presentation orders. On held-out instances, it achieves 86.08% effective agreement with human judgments. Across 5 frontier systems, overall rankings conceal distinct strengths across reference compositions. A recurring failure is to reproduce unintended source content in place of the requested result, despite closely resembling a reference. Reproducible pointwise diagnostics of quality, reference affinity, and speech reveal distinct dimensions of model behavior. ORAV thus offers a benchmark for tracking progress toward controllable, compositional, and reference-faithful audio-video generation.
Figures & tables
Fig. 1: Omni-reference audio-video generation requires selectively composing information from heterogeneous references. (a) Representative reference-conditioned task types include identity-conditioned generation, motion transfer, and audio-driven animation. (b) An omni-reference task instance combines subject and scene images, a motion video, and a speech reference under a textual instruction. The target subject speaks the specified utterance using the referenced vocal timbre, then performs the referenced motion in the designated scene. (c) Successful generation preserves assigned properties, prevents attribute leakage, and coherently combines speech and motion.
Benchmark
Generation
Reference-derived factors
Joint composition
Output
References
Content
Dynamics
Audio
T2AV-Compass ( 2026 )
AV
–
–
–
–
–
VABench ( 2026 )
AV
I
✓
–
–
–
UI2V-Bench ( 2025 )
V
I
✓
–
–
–
LongAV-Compass ( 2026 )
AV
I, V
✓
–
–
–
OpenS2V-Eval ( 2025 )
V
I
✓
–
–
C
Tab. 1: Comparison with related video and audio–video generation benchmarks. Reference modalities and semantic coverage are pooled across task instances; joint composition reports the largest set of semantic families required together from distinct references in a single instance.
Fig. 2: Overview of the benchmark. Top: Nine semantic roles describe the intended contributions of image, video, and audio references. Middle: Representative task instances combine heterogeneous references with a global textual instruction, which specifies what to extract from each reference and how to bind and compose the selected information in the generated audio-video output. Bottom: Benchmark statistics summarize dataset scale, semantic-role frequencies, composition complexity, and video/audio reference-duration distributions.
Fig. 3: From reference evidence to a joint judgment. With subject likeness tied, Y’s closer reproduction of forward leg extension with a lowered ball outweighs X’s better fixed-camera compliance. The visual case contrasts Gemini-3.1-Pro-Preview on native video with GPT-6-Astra on panels; the speech branch illustrates the integration of auditory evidence.
#
System
Overall
Image
Video
Audio
Strength
95% CI
Subj.
Prop
Scene
Style
Motion
Cam.
VFX
Speech
Music
1
Seedance-2.5
+0.70
[+0.56,+0.85]
68.92
64.07
68.61
65.48
71.01
65.38
82.63
68.04
76.56
2
Seedance-2.0
+0.48
[+0.36,+0.59]
62.82
61.01
61.85
63.54
65.70
64.90
71.36
59.95
50.98
3
MiniMax-H3
+0.15
[+0.02,+0.28]
53.39
56.40
56.87
57.51
42.37
65.87
58.28
73.58
39.85
4
Wan3.0-Video
−0.56
[−0.70,−0.42]
35.50
38.90
34.71
36.47
30.06
25.00
30.00
40.90
40.42
5
Kling-v3-Omni
−0.77
[−0.93,−0.63]
29.37
29.63
27.96
27.00
40.86
28.85
7.73
7.53
42.19
Tab. 2: Overall and role-conditioned performance on ORAV. Systems are ordered by Davidson–Bradley–Terry strength on 350 shared-delivery instances, with 95% bootstrap intervals over instances. Role columns report expected win rates (%) for overlapping subsets containing each role, including their co-occurring requirements. Bold marks the highest estimate in each column.
System
Visual quality ↑
Ref. affinity ↑
Audio ↑
Speech
ORAV ↑
Comp. z
TQ
AP
Subject
Scene
PQ
Voice ↑
CER ↓
WER ↓
Win rate
Seedance-2.5
+0.10
−1.78
4.95
40.61
77.87
6.99
0.45
5.51
5.58
68.38
Seedance-2.0
+0.17
−0.67
4.99
39.29
79.01
7.11
0.48
19.17
16.23
62.69
MiniMax-H3
+0.13
−1.30
4.60
40.15
75.49
6.82
0.59
14.76
18.09
54.00
Wan3.0-Video
−0.16
−1.91
5.05
38.94
77.09
7.12
0.39
10.27
112.48
35.04
Kling-v3-Omni
−0.19
−3.97
4.46
40.45
77.64
6.72
1.00
187.84
178.14
29.90
Tab. 3: Diagnostic measurements and ORAV performance. Diagnostic means weight instances measurable for all five systems equally. Comp. z : VBench-style quality; TQ: raw DOVER++ technical score ×100 ; AP: Aesthetic Predictor V2.5. Subject/Scene are full-frame DINO/CLIP similarities ×100 ; PQ is Audiobox production quality; Voice is ECAPA cosine. CER/WER are character/word error rates (%); insertions can yield errors above 100%. ORAV reports Davidson–Bradley–Terry expected win rates (%) on shared-delivery instances. Kling-v3-Omni receives audio through a retained-soundtrack adapter because its tested interface lacks an audio-reference channel.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
System
Delivered
Rate
Usable comparisons
Rate
Seedance-2.5
362/380
95.3%
1344/1436
93.6%
Seedance-2.0
359/380
94.5%
1340/1427
93.9%
MiniMax-H3
375/380
98.7%
1373/1459
94.1%
Wan3.0-Video
376/380
98.9%
1368/1464
93.4%
Kling-v3-Omni
380/380
100.0%
1375/1472
93.4%
Appendix
Tab. 4: Delivery and evaluation coverage. Delivered videos are counted over all 380 task instances. Usable comparisons are the system’s resolved or tied comparisons divided by attempted comparisons between delivered candidates, across all instances. These denominators differ from the 350-instance ranking subset. Neither missing videos nor unavailable judgments are counted as task defeats.
Fig. 5: Exact pairwise outcomes on shared-delivery instances. Each cell gives the row system’s empirical win rate, (W+T/2)/(W+L+T) , above its wins–losses–ties record. Ties include order conflicts; unavailable judgments are excluded. Column numbers identify the systems listed in the rows. Color encodes the win rate, with blue above and red below 50%.
#
System
Strength
95% CI
Win rate (%)
1
Seedance-2.5
+0.67
[+0.51,+0.84]
67.93
2
Seedance-2.0
+0.53
[+0.39,+0.68]
64.24
3
MiniMax-H3
−0.03
[−0.17,+0.10]
49.18
4
Kling-v3-Omni
−0.53
[−0.69,−0.37]
35.76
5
Wan3.0-Video
−0.64
[−0.84,−0.47]
32.89
Appendix
Tab. 5: System ranking without audio references. Davidson–Bradley–Terry estimates use 2,406 comparisons on 259 shared-delivery instances with no speech or music references. Intervals use 2,000 instance-bootstrap replicates. Strengths sum to zero; win rates (%) are expected against a uniformly chosen other system. Bold marks the highest estimate.
Measure
Comp. z
TQ
AP
Subject
Scene
PQ
Voice
CER
WER
ORAV
Shared instances ( n )
350
350
350
332
166
91
72
52
22
350
Appendix
Tab. 6: Shared instance counts for the main comparison. Diagnostic counts require valid values from all five systems for each metric; Subject uses DINO and Scene uses CLIP. ORAV uses instances with all five outputs delivered, retaining 3,280 available pairwise outcomes for the strength fit.
Measurement
Component
Source
Visual affinity
CLIP ViT-B/32
Radford et al. (2021)
Visual affinity
DINOv1 ViT-B/16
Caron et al. (2021)
Voice similarity
ECAPA-TDNN
Desplanques et al. (2020)
Speech text
Whisper-small
Radford et al. (2023)
Motion direction
CoTracker3 scaled_offline
Karaev et al. (2025)
Appendix
Tab. 7: Pretrained components for affinity, speech, and motion diagnostics. Component names identify the evaluated variants; references identify their published methods.
System
Subj.
Bkgd.
Smooth.
Aesth.
Imag.
AQ
Shared instances ( n )
350
350
350
350
350
350
Seedance-2.5
91.6
93.9
99.3
56.3
67.3
+0.035
Seedance-2.0
92.4
93.3
99.2
57.0
71.0
+0.040
MiniMax-H3
90.6
92.2
99.3
59.4
70.7
+0.036
Wan3.0-Video
89.8
91.2
99.1
54.7
70.0
+0.038
Kling-v3-Omni
90.9
93.0
99.1
53.1
64.2
+0.028
Appendix
Tab. 8: Additional visual-quality dimensions. VBench-style consistency, smoothness, and aesthetic scores are multiplied by 100; imaging quality retains its 0–100 scale. DOVER++ AQ reports its raw aesthetic score. Means use the shared instances in each column. Higher values indicate better quality within each measure; scales differ across columns.
System
Face
ALADIN
CLIP-Style
DINO-Style
Motion dir.
Shared instances ( n )
252
43
43
43
73
Seedance-2.5
0.360
0.324
0.627
0.272
0.290
Seedance-2.0
0.347
0.349
0.621
0.286
0.301
MiniMax-H3
0.336
0.246
0.610
0.244
0.311
Wan3.0-Video
0.332
0.371
0.643
0.342
0.572
Kling-v3-Omni
0.234
0.287
0.635
0.309
0.606
Appendix
Tab. 9: Complementary reference diagnostics. Face cosine uses ArcFace on detected faces; style cosines use ALADIN, CLIP, and DINO; motion compares visible foreground track directions. Each column uses its own five-system common subset. These affinities describe the measured feature and do not establish the identity of the actor performing an action or the correctness of the complete task.
System
NISQA
DNSMOS
CLAP
MuQ
Chroma
Waveform
Shared instances ( n )
74
72
17
17
17
17
Seedance-2.5
2.83
2.80
0.404
0.668
0.884
0.229
Seedance-2.0
3.11
2.97
0.509
0.763
0.890
0.216
MiniMax-H3
3.20
2.90
0.562
0.783
0.940
0.800
Wan3.0-Video
2.89
2.76
0.560
0.792
0.925
0.508
Kling-v3-Omni
3.01
2.68
0.661
0.929
0.988
1.000
Appendix
Tab. 10: Speech quality and music-reference affinity. NISQA MOS and DNSMOS OVRL measure speech quality. CLAP, MuQ, chroma-DTW, and waveform correlation compare music references with the output. Music affinity has no universal quality direction: a task may request a new composition rather than the source recording. Kling-v3-Omni uses retained reference soundtracks.
System
Ruled unusable
Substitution cues
Reference voice
All
Video ref.
No video ref.
Share of vetoes
Replayed
Seedance-2.5
2.3%
3.2%
0.0%
95%
0/75
Seedance-2.0
4.7%
6.0%
1.4%
91%
0/75
MiniMax-H3
7.7%
10.7%
0.0%
94%
0/78
Wan3.0-Video
43.8%
55.9%
12.0%
94%
9/78
Kling-v3-Omni
32.0%
31.4%
33.9%
96%
79/79
Appendix
Tab. 11: Reference-use failures. Unusability rates count ordered candidate judgments; substitution cues are a share of vetoes. Replay counts use unique outputs for speech instances.
Fig. 9: Reference-specific motion beyond overall appearance. In subject-motion-7f681a , forward leg extension with a lowered ball distinguishes the demonstrated dunk from a conventional rendition. Frames and verbatim quotations come from the recorded evaluation.
Fig. 10: Selective transfer of motion and performers. In subject-motion-be4824 , reproducing the competition hall and its original dancers replaces the requested performers and scene. Frames and verbatim quotations reveal reference substitution rather than task fulfillment.
Component
Reasoning
Temperature
Output budget
GPT-6-Astra judge
medium
—
8,192
Gemini-3.1-Pro listener
medium
0.3
32,768
Gemini-3.1-Pro vanilla
high
0.3
16,384
Appendix
Tab. 12: Judge and listener settings. Output budgets are in tokens.
Comparison
Agreement (%)
Δ (pp)
95% CI ( Δ )
ORAV Judge vs. Simplified Astra
86.97/79.52
+7.45
[+1.08,+15.83]
Simplified Astra vs. Vanilla Gemini
82.09/71.89
+10.20
[+4.48,+17.74]
Appendix
Tab. 13: Paired human agreement on the 38 held-out instances. Each row uses its own shared eligible comparisons. Scores follow evaluator order; Δ is their paired difference. Instances receive equal weight; 95% intervals for Δ use 2,000 instance-bootstrap replicates.
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these settings, MR2AV requires models to jointly reason over multiple references while generating synchronized visual and audio content. Models must not only preserve each reference faithfully but also correctly bind and compose multiple referenced entities into coherent audio-visual events. To address this gap, we introduce MultiRef-Compass, a unified benchmark for MR2AV generation. It comprises 350 carefully curated samples constructed through a scalable and controllable asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. To provide interpretable assessment, MultiRef-Compass defines an evaluation protocol with four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following, using 14 sub-metrics. MultiRef-Compass integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition. Extensive experiments on eight representative MR2AV systems reveal substantial room for improvement across multiple evaluation dimensions, underscoring the need for a comprehensive benchmark and positioning MultiRef-Compass as a foundation for future MR2AV research.
Xiaohan Zhang, Yuqing Wen, Junlin Chen +9
Nanjing University · Kling Team · National University of Singapore +3
Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios involving speech and interactions. Yet evaluation for AV generation remains at an early stage, with only a few coarse-grained benchmarks for human-related scenarios and relying on limited preset evaluations with generic multimodal LLMs, leading to inaccurate assessments of model capabilities. To address these issues, we introduce AVBench, a fully automated benchmark tailored for human-centric AV generation. AVBench is built on two key designs for comprehensive and accurate evaluation: (i) Human-centric and fine-grained metrics. AVBench integrates ten evaluation dimensions designed for human-centered real-world scenarios, covering visual quality, audio quality, and multi-level consistency across modalities. These practical metrics capture human-related details that existing benchmarks often overlook. (ii) Specialized evaluators via preference learning. To address the lack of specialized training data, we construct large-scale supervision by transforming real-world videos into diverse training pairs with controlled perturbations. After fine-tuning on this high-quality dataset, the evaluators learn to reliably detect subtle cross-modal inconsistencies. Crucially, instead of producing discrete textual judgment, AVBench derives continuous evaluation scores from the model's prediction confidence on binary decisions. This probabilistic scoring mechanism enables a more reliable assessment than traditional VQA-style evaluation and aligns closely with human judgment. Taken together, AVBench offers automated evaluation for AV generation, demonstrates strong potential for data filtering, and serves as a differentiable reward signal for Reinforcement Learning from Human Feedback (RLHF).
Jialiang Yang, Bin Xia, Ruihang Chu +6
Tsinghua University · The Chinese University of Hong Kong
Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio modality. Existing benchmarks either treat audio as an auxiliary component of video quality or assess it in isolation from audiovisual grounding, making it difficult to diagnose where current systems truly succeed or fail in audio generation. We present PRISM-Bench, the first audio-centric diagnostic benchmark for T2AV generation. Built from a rigorously curated dataset of 900 human-verified samples, PRISM-Bench factorizes audio evaluation along two orthogonal axes: audio type (Speech, Music, and Sound) and sound-source visibility (On-screen vs. Off-screen). It evaluates generated content across four perceptual dimensions (Audio-Visual Coherence, Audio Quality, Audio Expressiveness, and Prompt Following) with 35 fine-grained criteria. To ensure reliable assessment, we adopt an enhanced MLLM-as-a-Judge protocol based on blind, side-by-side comparison against ground-truth references, demonstrating strong alignment (over 70% mean agreement) with human raters. Our evaluation of recent T2AV systems highlights a significant performance gap between frontier and open-source models. Furthermore, we demonstrate that current generation paradigms overfit to perceptual fidelity while struggling with complex grounding and control tasks, particularly in generating music and synchronized On-screen audio.