Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide coverage without localization, local probes provide localization without coverage, and unconstrained LLM judges introduce instability. We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic framework. OmniCapBench shifts the prediction target from free-form text to sets of atomic, verifiable evaluation units across three tracks: entity references, visual shots, and audio events, enabling reliable scoring with deterministic constraint checks and localized LLM-based semantic comparisons. With 786 densely annotated videos, OmniCapBench effectively distinguishes MLLM perception errors, including temporal grounding failures, identity drift, cross-modal misalignment, and hallucinated descriptions. Evaluating frontier MLLMs reveals strong local perception but weak long-horizon audio--visual reasoning, particularly in identity drift and cross-modal misalignment, providing a fine-grained roadmap for omnimodal development.
Figures & tables
Benchmark
Evaluated Unit
Scoring Operator
Identity Tracking (Visual)
Temporal Grounding (Audio/Visual)
Audio-Visual Association (Audio-Visual)
Diagnostic Traceability
Whole-Caption Paradigm
AuroraCap [ 5 ]
Caption
Global LLM
✗
✗
✗
✗
VCapsBench [ 62 ]
Caption
Global LLM
✗
✗
✗
✗
UGC-VideoCap [ 43 ]
Caption
Global LLM
✗
✗
✗
✗
video-SALMONN 2 [ 34 ]
Caption
Global LLM
✗
✗
✗
✗
Probe-Based Paradigm
Table 1: Evaluation paradigm comparison. We compare whether existing protocols natively support the structural dimensions of omnimodal diagnosis. Identity Tracking explicitly evaluates entity permanence across discontinuous shots (Visual). Temporal Grounding enforces precise continuous time boundaries for actions and sounds (Audio/Visual). Audio-Visual Association evaluates whether audio events are correctly attached to concurrent visual shots (Audio-Visual). Diagnostic Traceability indicates whether the evaluation isolates explicit structural breakdowns from descriptive hallucinations. ✓ : natively supported as a verifiable unit; △ : implicitly judged; ✗ : not supported.
Figure 2: OmniCapBench data construction pipeline. The construction proceeds in three stages. Stage I filters raw videos to isolate inputs with rich audio–visual complexity. Stage II generates atomic units (References → Events → Shots) through an iterative multi-model loop, enforcing structural integrity before semantic refinement. Stage III applies traceable human audits to finalize the rigorously verified audio–visual reference system.
Figure 3
Model
SGC
Visual
Audio
Audio-Visual
RefUse
CCC
Ref Subject F1
Ref Scene F1
Shot F1
Shot tIoU
Subshot F1
Subshot tIoU
Event F1
Event tIoU
Speaker F1
EVSA F1
Proprietary Models
Gemini 3.1-Pro [ 15 ]
97.36
70.39
34.43
78.60
82.68
76.24
70.74
65.27
48.76
63.32
71.92
91.18
51.46
Gemini 2.5-Pro [ 10 ]
96.73
70.75
37.81
80.70
83.65
70.93
72.58
66.69
49.18
58.52
69.35
89.72
43.74
Qwen3.5-Omni-Plus [ 32 ]
96.09
66.92
29.55
73.57
79.97
71.25
72.35
65.69
48.29
53.92
67.77
91.20
40.20
Qwen3.5-Omni-Flash [ 32 ]
95.58
63.10
24.31
72.27
66.07
65.65
68.99
58.94
46.52
50.68
69.34
89.91
35.58
Table 3: Rule-based structural evaluation. Metrics quantify structural compliance across Visual, Audio, and Audio-Visual dimensions. SGC is an independent schema prerequisite. Scores are macro-averaged across videos. Best and second-best results are highlighted.
Model
Visual
Audio
Shot
Subject
Scene
Subshot
Dialogue
Non-dialogue
Recall
Precision
Recall
Precision
Recall
Precision
Proprietary Models
Gemini 3.1-Pro [ 15 ]
59.05
43.78
67.07
51.78
75.62
53.29
70.92
91.18
41.49
Gemini 2.5-Pro [ 10 ]
55.58
51.53
57.08
56.80
65.12
53.05
71.21
89.72
37.38
Qwen3.5-Omni-Plus [ 32 ]
54.17
45.41
62.84
50.75
69.80
50.22
70.66
91.20
36.82
Table 4: Localized semantic evaluation. Semantic fidelity is assessed exclusively on structurally aligned units. Bidirectional scoring isolates distinct failure modes: Recall penalizes factual omissions, while Precision penalizes hallucinations. Best and second-best results are highlighted.
Figure 4: Evaluation paradigms and structural failures. (a) Global text metrics make distinct models appear similar, whereas structured constraints better match human audit. (b) Error composition shows different bottlenecks: frontier models concentrate failures in fine-grained hallucination and audio-visual alignment, while open-source models exhibit broader degradation.
Judge
Holistic
QA
OmniCapBench
Weak
71.24
61.17
56.82
Mid
83.08
53.49
58.24
Strong
93.92
56.57
57.60
Table 5: Judge sensitivity. Stability across judges.
Figure 5: Qualitative case study of deep-structured diagnostics. Traditional holistic evaluations often over-reward fluent prose, missing critical errors like hallucinated “protective eyewear” or temporally unsupported actions. By performing bounded semantic checks on structurally valid units and assigning localized Recall and Precision, OmniCapBench decouples genuine descriptive fidelity from structural noise, providing a high-resolution diagnostic trace.
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Input
Output
Diagnostic Target
Generation
Raw Video V
Scored structured units
Native recovery and precision
Probing
Predicted JSON S^(V)
Probe pass/fail
Internal structural validity
Verification
V + Candidate JSON unit
Accept/reject boolean
Local contradiction detection
Appendix
Table A.1: OmniCapBench Task Suite. All diagnostic tasks derive from the same frozen system of atomic units S⋆(V) , but each exposes a distinct model failure mode.
Bucket
Metric
Required Range
All
shots/min
[5,30]
All
events/min
[2,15]
<1 min
subjects/min
[3,12]
<1 min
scenes/min
[3,15]
<1 min
total shots
≥4
<1 min
total events
≥2
Appendix
Table A.2 : Candidate curation thresholds. Video-level filters applied before reference-unit construction. Per-minute intervals are inclusive and follow [center−std,3×center] ; absolute thresholds apply only to the listed duration buckets.
Component
Fields
Description
references
ref_id , type , semantic_description , appearance_anchor
Persistent entities or scenes with detailed visual descriptions under id_features
Verifiable evidence fields utilized for local semantic diagnostics
Appendix
Table A.3 : Schema Field Definitions. Descriptions and typing rules for every structural component required from the evaluated model.
Metric
Value
Global Counts and Averages
Total Videos
786
Mean Duration (seconds)
58.82
Median Duration (seconds)
28.32
Avg. References / Video
7.40
Avg. Events / Video
8.32
Appendix
Table A.4 : Full Dataset Statistics. Comprehensive macro averages and distribution metrics supplementing Table 2 . These statistics explicitly define the coverage and complexity of the OmniCapBench evaluation scope.
Figure A.1 : Data Source and Category Distributions. The left panel illustrates the distribution of source datasets, ensuring domain diversity. The right panel shows the composition of video durations across the top 10 semantic categories.
Figure A.2 : Temporal Footprint and Structural Complexity. The left panel provides an exhaustive view of the video duration distribution. The right panel contextualizes the structural complexity by plotting visual cuts density against action density for different categories.
Figure A.3 : Structured Unit Counts per Video. Individual histograms of References, Events, and Shots per video, illustrating the dense tracking requirements placed on the models.
Figure A.4 : Temporal Scale and Per-Minute Structural Density. The left panel relates video duration to event count; the right panel summarizes shots-per-minute and events-per-minute distributions. These two views separate absolute temporal scale from normalized structural density.
Layer
Matching Metric
Default threshold
Tie-breaking
Reference
LLM bipartite matching over detail_description
None (LLM output)
Defensive ID/type/one-to-one check
Dialogue event
1−WER over content.line
≥0.50
Max line score, one-to-one
Non-dialogue event
Temporal IoU
≥0.20
Max tIoU, one-to-one
Shot
Temporal IoU
≥0.30
Max tIoU, one-to-one
Appendix
Table A.5 : Matching Rules Summary. Parameters defining structural alignment across the reference, event, and shot layers.
Prediction type
Matched
Valid structure
Supported
Credit / penalty
Denominator
Matched unit
yes
yes
n/a
Full credit; proceeds to structural/semantic checks.
Matched metric denominators
Unmatched reference
no
yes/no
no
No credit; retained in diagnostic logs.
Appendix audit
Unmatched event
no
yes/no
no
No credit; retained in diagnostic logs.
Appendix audit
Unmatched sub-shot
no
yes/no
no
No credit; retained in diagnostic logs.
Appendix audit
Invalid or dangling link
no
no
no
Zero credit; penalized as a structural error.
Appendix audit
Valid but out-of-scope fact
no
yes
unknown
No manual credit, but no strict penalty; logged for auditing.
Reported diagnostics
Appendix
Table A.6 : Open-World Scoring Policy. This table outlines exactly how we credit, penalize, or ignore different types of predictions within our bounded evaluation scope. Crucially, while OmniCapBench rewards models for recovering annotated truths, it does not unfairly penalize models for correctly identifying true facts that merely fell outside our annotation budget. Unsupported extras are logged for diagnostic analysis but do not directly dock the main precision metrics.
Model
Checkpoint / Endpoint
Params
Backend
Audio
Frames / FPS
Max Tokens
Duration
Long-Video Strategy
Proprietary Models
Gemini 3.1-Pro
gemini-3.1-pro-preview
–
API
✓
no cap / 2 fps
65536
All
Single-pass
Gemini 2.5-Pro
gemini-2.5-pro
–
API
✓
no cap / 2 fps
65536
All
Single-pass
Qwen3.5-Omni-Plus
qwen3.5-omni-plus
–
API
✓
no cap / 2 fps
32768
All
Single-pass
Qwen3.5-Omni-Flash
qwen3.5-omni-flash
–
API
✓
no cap / 2 fps
32768
All
Single-pass
Open-source Omnimodal Models
Appendix
Table A.7 : Model inference and configuration settings. One row per evaluated model entry. All models use greedy decoding (temperature = 0, top_p = 1) to minimize structural hallucination. Closed-source models are accessed via API with no frame cap; local Qwen3-Omni models (30B-A3B) are served via vLLM with TP = 4 and max 192 frames at 2 fps; videos exceeding 1 min use a 2-pass strategy. Qwen2.5-Omni-based models (7B and fine-tuned variants) only evaluate videos ≤ 60 s due to context limits.
Error Family (Category)
Primary Metric
What triggers this error?
Count
Rate
Schema Failure (Gate)
SGC
Output cannot be parsed into references, events, and shots, or is missing required JSON fields
31
3.8%
Reference Recovery (Visual)
Ref F1
Required ground-truth reference is missed, or a hallucinated predicted reference cannot be matched
3778
31.7%
Reference-Use (Visual)
RefUse
A matched shot or subshot cites the wrong persistent reference ID
8649
34.1%
Identity Drift (Visual)
CCC
A single persistent reference is split, merged, or inconsistently tracked across multiple shots
2656
65.2%
Shot Recovery/Alignment (Visual)
Shot F1/tIoU
A visual micro-action or boundary is missed, hallucinated, or severely misaligned in time
5156
26.2%
Event Recovery/Alignment (Audio)
Event F1/tIoU
A required audio event is missed, hallucinated, or aligned to a completely wrong time segment
4218
36.2%
Appendix
Table A.8 : Error analysis summary. This table defines exactly how we classify failures in structurally valid outputs. It explicitly links each failure mode to the corresponding rule-based metric. Counts and rates are derived from our evaluation logs over applicable samples for the Gemini 3.1-Pro model. This detailed breakdown ensures that every capability deficit is independently traceable rather than being obscured by a single aggregate score.
Failure type
Category
Metric
What it looks like in practice
Rate
Broken identity persistence
Visual
CCC
The model correctly describes a character but assigns them a new ID, such as PERSON_2 , when they reappear in a later shot, losing track of their persistent identity.
65.2%
Wrong event reference
Visual
RefUse
An off-screen voice or a background action is attributed to the wrong character ID in the structured output.
34.1%
Event alignment error
Audio
Event tIoU
The model correctly identifies an event, like a dog barking, but aligns it to a completely wrong time segment where the dog is silent.
42.0%
Missing structural field
Gate
SGC
The JSON output is missing required fields like time_range or detail_description , rendering the unit untestable.
3.8%
Unsupported extra event
Audio
Event (Precision)
The model hallucinates an action or event that never actually occurred in the source video.
36.2%
Appendix
Table A.9 : Failure taxonomy for outputs exhibiting high text-level scores but low structural scores. This table dissects exactly why a seemingly fluent holistic prose caption can fail rigorous deep-structured evaluation. We map specific failure modes to the exact metric designed to penalize them. All rates are derived from our validation logs for the Gemini 3.1-Pro model.
Probe type
Formal Predicate
Failure condition
Count / Prop.
Reference-use validity
ValidRefUse(x,r)
Missing aligned region, dangling ID, or reference mismatch after alignment
8649 / 34.1%
Temporal compatibility
TemporalCompatible(e,h)
Predicted interval is disjoint from the required shot/event support or violates ordering tolerance
12240 / 42.0%
Reference-support validity
SupportFieldPresent(r)
Missing support field, malformed support field, or support field incompatible with the matched reference
3778 / 31.7%
Cross-reference integrity
ConsistentCoreference(r)
Referenced ID is undefined, inconsistent across shots, or incompatible with the matched persistent reference
2656 / 65.2%
Appendix
Table B.1 : Structural Consistency Probe Types. Each row defines a deterministic probe family, its formal predicate, and the exact condition counted as a failure. Proportions are computed over the subset of structurally applicable predictions for the Gemini 3.1-Pro model to isolate structural reasoning from raw detection capability.
Figure B.1 : Instance-level correlation among diagnostic metrics. Pearson correlation coefficients computed across 786 videos (Gemini 3.1-Pro). Left: Structural correlations confirm that complex capabilities like identity tracking ( CCC ) are decoupled from basic temporal grounding ( Shot F1 , r=0.17 ). Right: Semantic LLM-as-judge scores exhibit near-zero correlation between Precision and Recall for Subject descriptions ( r=−0.03 ), proving that hallucination tendency and factual coverage are independent failure modes.
Model
RefUse-shot
RefUse-sub
CCC
Ref R
Ref P
Shot R
Shot P
Shot tIoU
Subshot R
Subshot P
Subshot tIoU
Proprietary Models
Gemini 3.1-Pro [ 15 ]
68.49
68.77
34.43
77.02
70.56
71.73
88.77
70.74
75.40
60.29
48.76
Gemini 2.5-Pro [ 10 ]
67.01
71.42
37.81
76.96
71.39
66.29
86.27
72.58
70.26
66.89
49.18
Qwen3.5-Omni-Plus [ 32 ]
63.99
65.22
29.55
65.02
70.99
63.90
89.62
72.35
69.00
66.66
48.29
Qwen3.5-Omni-Flash [ 32 ]
59.39
61.07
24.31
59.50
69.26
55.77
89.71
68.99
61.93
60.45
46.52
Open-source Omnimodal Models
Appendix
Table B.2 : Expanded Visual rule-based submetrics. Ref R/P are reference recall and precision. RefUse-shot and RefUse-sub report WHO_shot and WHO_sub . CCC is the macro average of cross-shot reference-trajectory IoU. Shot/Subshot R/P report alignment recall and precision, and tIoU reports temporal overlap ( WHEN_H , WHEN_S ). Within each column, bold is best and underline is second-best among models (ties allowed).
Model
Audio
Audio-Visual
Event R
Event P
Event tIoU
Speaker
EVSA-P
EVSA-R
EVSA
Proprietary Models
Gemini 3.1-Pro [ 15 ]
63.96
67.20
71.92
91.18
78.35
42.14
51.46
Gemini 2.5-Pro [ 10 ]
58.09
64.23
69.35
89.72
74.05
35.04
43.74
Qwen3.5-Omni-Plus [ 32 ]
51.22
62.53
67.77
91.20
71.11
31.82
40.20
Qwen3.5-Omni-Flash [ 32 ]
48.01
58.91
69.34
89.91
66.28
27.69
35.58
Appendix
Table B.3 : Expanded Audio and Audio-Visual rule-based submetrics. Within each column, bold is best and underline is second-best among models (ties allowed).
Model
Visual
Audio
Subject Ref
Scene Ref
Shot Fact
Dialogue
Non-dialogue
R
P
Mean
R
P
Mean
Mean
Mean
Mean
Proprietary Models
Gemini 3.1-Pro [ 15 ]
43.78
67.07
55.43
51.78
75.62
63.70
59.05
91.18
41.49
Gemini 2.5-Pro [ 10 ]
51.53
57.08
54.30
56.80
65.12
60.96
55.58
89.72
37.38
Qwen3.5-Omni-Plus [ 32 ]
45.41
62.84
54.13
50.75
69.80
60.27
54.17
91.20
36.82
Appendix
Table B.4 : Expanded LLM-as-judge submetrics. Dialogue is the rule-based dialogue line structural score. Within each column, bold is best and underline is second-best among models (ties allowed).
Model
SGC
Visual
Audio
Audio-Visual
RefUse
CCC
Ref Subj. F1
Ref S. F1
Shot F1
Shot tIoU
Sub F1
Sub tIoU
Evt F1
Evt tIoU
Spk F1
EVSA F1
Duration: < 1 min
Gemini 3.1-Pro
97.46
74.13
38.41
83.09
86.54
77.90
69.76
67.55
47.73
63.09
70.65
91.15
53.30
Gemini 2.5-Pro
97.57
73.40
40.60
84.80
87.39
76.61
71.72
70.94
48.36
60.46
67.46
89.08
49.31
Qwen3.5-Omni-Plus
96.90
70.94
32.71
80.28
84.99
77.60
73.57
69.65
48.73
57.12
64.57
90.77
46.87
Qwen3.5-Omni-Flash
96.26
66.83
26.81
77.37
66.70
72.86
69.94
64.57
47.07
52.72
66.90
89.34
41.65
Appendix
Table B.5 : Per-duration rule-based structural results. Stratification follows the three duration buckets ( <1 min, 1 – 3 min, and 3 – 5 min); Ref Subj./Ref S. are recomputed per video in each bucket as in the main table. Within each column, bold is best and underline is second-best among models (ties allowed).
Model
Visual
Audio
Shot
Subject
Scene
Subshot
Dialogue
Non-dialogue
Recall
Precision
Recall
Precision
Recall
Precision
Duration: < 1 min
Gemini 3.1-Pro
62.84
46.74
68.18
52.62
78.01
53.90
71.75
91.15
43.39
Gemini 2.5-Pro
59.70
54.97
57.11
57.66
67.97
54.19
72.40
89.08
40.03
Qwen3.5-Omni-Plus
58.92
48.27
62.41
51.83
71.96
51.88
70.08
90.77
38.23
Appendix
Table B.6 : Per-duration local semantic-equivalence score results. Stratification follows the three duration buckets ( <1 min, 1 – 3 min, and 3 – 5 min). Within each column, bold is best and underline is second-best among models.
Figure B.2 : Case study of identity persistence measured by CCC. The timeline compares identity assignments across models over the same video interval. Colored spans show ground-truth and predicted identity coverage. Higher CCC indicates more consistent cross-shot identity reuse; lower CCC indicates identity fragmentation. In this case, Gemini 3.1 Pro, Gemini 2.5 Pro, and Qwen3.5-Omni-Plus score 76.92, 78.10, and 80.00, while Qwen3-Omni-Instruct scores 30.77, revealing a clear failure in preserving identity continuity.
Figure B.3 : Instruction-following failures induced by caption-specialized SFT. The two models either reject the requested JSON contract outright or satisfy only the generic request for machine-readable captions while missing the required Reference–Event–Shot structure.
Figure B.4 : Structural-output failures after partial schema compliance. Even when caption-specialized models emit OmniCapBench-like keys, the generated objects do not form usable evaluation graphs: events are missing, references can be undefined, and local descriptions collapse into repeated zero-duration fragments.
Model
Text Metric (Prose)
Native Constraints
Human Audit
Pair 1: Qwen3.5-Omni-Flash vs. Qwen3-Omni
Weaker (Qwen3-Omni)
52.2
37.5
21.6
Stronger (Qwen3.5-Omni-Flash)
47.8
62.5
78.4
Pair 2: Qwen3.5-Omni-Plus vs. Qwen3.5-Omni-Flash
Weaker (Qwen3.5-Omni-Flash)
41.7
21.4
29.9
Stronger (Qwen3.5-Omni-Plus)
58.3
78.6
70.1
Appendix
Table B.7 : The illusion of global text scores. We construct two hard-to-distinguish subsets where a weaker baseline and a stronger upgraded model exhibit near-identical text-centric scores. Across these pairs, we track performance through progressive strictness: (1) traditional text-centric metrics on dense captions, (2) strict structural verification using our native atomic units, and (3) human audit. Global text metrics falsely penalize the stronger models and overestimate the weaker ones, creating an illusion of parity. However, evaluating structural constraints directly resolves this gap, perfectly aligning with human judgment.
Model
Track
SGC
RefUse
CCC
Ref Subj
Shot F1
Shot tIoU
Evt F1
Evt tIoU
Speaker
EVSA
Gemini 3.1-Pro [ 15 ]
Native
97.36
70.39
34.43
78.60
76.24
70.74
63.32
71.92
91.18
51.46
Post-hoc
99.19
46.91
5.20
57.11
97.39
95.62
40.73
56.65
87.36
38.29
Qwen3-Omni-Instruct [ 49 ]
Native
89.02
40.96
9.85
54.84
49.98
67.92
41.70
40.99
88.35
13.33
Post-hoc
96.61
40.93
9.62
49.21
96.44
95.59
12.82
54.40
81.65
12.23
ASID-Captioner-7B [ 25 ]
Post-hoc
96.17
39.58
9.65
41.89
96.51
95.73
11.09
54.83
87.85
10.87
Appendix
Table B.8 : Native generation vs. the post-hoc parsing track (rule-based metrics). Native asks the model for OmniCapBench units directly; Post-hoc lets the model write a free-form caption that a fixed, video-blind LLM parser converts into units. Post-hoc raises structural well-formedness (SGC, Shot F1, Shot tIoU) while every grounding-dependent metric (CCC, Event F1, EVSA) degrades, showing that format repair does not recover information the caption never contained.
Model
Track
Shot
Subj-R
Subj-P
Scene-R
Scene-P
Subsh-R
Subsh-P
Dialogue
Non-dial
Gemini 3.1-Pro [ 15 ]
Native
59.05
43.78
67.07
51.78
75.62
53.29
70.92
91.18
41.49
Post-hoc
35.84
24.75
75.90
24.72
82.59
36.15
62.94
87.36
21.42
Qwen3-Omni-Instruct [ 49 ]
Native
34.63
40.31
62.02
36.07
70.73
29.42
57.84
88.35
20.51
Post-hoc
28.67
28.49
71.82
29.32
77.37
33.77
50.75
81.65
13.15
ASID-Captioner-7B [ 25 ]
Post-hoc
22.87
25.24
65.66
32.96
64.35
33.25
40.17
87.85
10.37
Appendix
Table B.9 : Native generation vs. the post-hoc parsing track (localized LLM-judge metrics). Every local semantic score drops under post-hoc conversion, including for the caption-specialized model that the track is meant to accommodate. Precision rises on some fields only because the parser emits fewer, more conservative units.
Sweep
Metric
Setting
Gemini 3.1-Pro
Gemini 2.5-Pro
Qwen3.5- Omni-Plus
Qwen3.5- Omni-Flash
Qwen3-Omni- Instruct
Qwen3-Omni- Captioner
Non-dialogue tIoU
Event F1
0.10
64.02
59.54
54.63
51.43
42.29
41.88
0.20
63.32
58.52
53.92
50.68
41.70
41.31
0.40
61.86
56.34
52.38
48.95
40.71
40.07
Shot tIoU
Shot F1
0.20
77.84
72.32
72.96
67.81
54.22
55.34
0.30
76.24
70.93
71.25
65.65
49.98
52.29
0.50
64.07
60.87
60.37
52.79
37.41
38.82
Appendix
Table B.10 : Matching-threshold sweep. Each block varies one cutoff while predictions, references, and the one-to-one matching rule stay fixed; no inference is re-run. Default settings are shown in bold . Absolute scores shift with the cutoff, but the induced ranking does not: Spearman’s ρ against the default ranges from 0.914 to 1.000 (Appendix B.7 ).
Table B.11 : License Information for Scientific Artifacts. This table details the licenses for the underlying video data sources and software models used in our evaluation.
Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language. However, their audio-visual intelligence (AVI) remains insufficiently evaluated due to the lack of systematic and comprehensive benchmarks. We introduce AVI-Bench, a cognitively inspired benchmark that evaluates Omni-MLLMs across three stages, perception, understanding, and reasoning, through cross-modal tasks requiring joint audio-visual interpretation. This design enables fine-grained diagnosis of model capabilities and failure modes. To further assess robustness beyond familiar domains, we propose AVI-Bench-PriSe, an extension that probes models' primitive audio-visual sensation using unfamiliar, low-semantic stimuli, testing generalization beyond common training distributions. Extensive experiments on both open-source and closed-source models reveal substantial limitations in current Omni-MLLMs. Based on these findings, we present a four-level AVI taxonomy. Overall, AVI-Bench provides a principled evaluation framework to guide the development of more robust and generalizable AVI. Project website: https://fudancvl.github.io/AVI-Bench/
Yaoting Wang, Ziyi Zhang, Wenming Tu +10
Institute of Big Data, College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China · Huazhong University of Science and Technology · Shanghai Jiaotong University +5
Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams, whereas existing benchmarks largely fail to evaluate this capability. They restrict videos to short clips, isolate modalities, or reduce questions to one-hop perception. We introduce TraceAV-Bench, the first benchmark to jointly evaluate multi-hop reasoning over long audio-visual trajectories and multimodal hallucination robustness. TraceAV-Bench comprises 2,200 rigorously validated multiple-choice questions over 578 long videos, totaling 339.5 hours, spanning 4 evaluation dimensions and 15 sub-tasks. Each question is grounded in an explicit reasoning chain that averages 3.68 hops across a 15.1-minute temporal span. The dataset is built by a three-step semi-automated pipeline followed by a strict quality assurance process. Evaluation of multiple representative OmniLLMs on TraceAV-Bench reveals that the benchmark poses a persistent challenge across all models, with the strongest closed-source model (Gemini 3.1 Pro) reaching only 68.29% on general tasks, and the best open-source model (Ming-Flash-Omni-2.0) reaching 51.70%, leaving substantial headroom. Moreover, we find that robustness to multimodal hallucination is largely decoupled from general multimodal reasoning performance. We anticipate that TraceAV-Bench will stimulate further research toward OmniLLMs that can reason coherently and faithfully over long-form audio-visual content.
Hengyi Feng, Hao Liang, Mingrui Chen +6
University of Electronic Science and Technology of China · Peking University · Zhongguancun Academy +1
Multimodal Large Language Models have demonstrated impressive video understanding, yet their ability to reason over long-form narratives is often masked by visual-centric evaluations and inefficient context processing. Existing benchmarks over-rely on visual heuristics while marginalizing auditory cues, effectively reducing models to "silent observers" that bypass genuine cross-modal reasoning. Moreover, standard dense sampling creates an evidence-context trade-off: increasing frames to capture evidence inevitably leads to attention distraction and token explosion. To bridge these gaps, we present Video-HolmesV2, a novel benchmark designed for Deep Audio-Visual Coupling. Unlike previous works, it enforces an Evidence-Based Evaluation, requiring models to justify answers with precise spatio-temporal audio-visual evidence, thereby reducing confounding effects of guessing and hallucinated evidence. To support this, we introduce: (1) a Multi-Model Cross-Verification pipeline to ensure task rigor; (2) a Spatio-temporal Evidence-Aware Metric for fine-grained calibration. Furthermore, we propose an Audio-Text Guided Token Compression framework. By fusing task intent with auditory anchors, our method distills high-value reasoning cues to mitigate long-context noise. In our evaluation, even strong proprietary models achieve below 60% accuracy, while our approach outperforms comparable open-source omni-models.
Zhaoyang Wei, Zipeng Wang, Yushe Cao +10
University of Chinese Academy of Sciences, China · Tsinghua University, China · Tencent, China