OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
Organizations: Hunyuan, Tencent · Nanyang Technological University · Northumbria University
Abstract
Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide coverage without localization, local probes provide localization without coverage, and unconstrained LLM judges introduce instability. We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic framework. OmniCapBench shifts the prediction target from free-form text to sets of atomic, verifiable evaluation units across three tracks: entity references, visual shots, and audio events, enabling reliable scoring with deterministic constraint checks and localized LLM-based semantic comparisons. With 786 densely annotated videos, OmniCapBench effectively distinguishes MLLM perception errors, including temporal grounding failures, identity drift, cross-modal misalignment, and hallucinated descriptions. Evaluating frontier MLLMs reveals strong local perception but weak long-horizon audio--visual reasoning, particularly in identity drift and cross-modal misalignment, providing a fine-grained roadmap for omnimodal development.
Figures & tables
| Benchmark | Evaluated Unit | Scoring Operator | Identity Tracking (Visual) | Temporal Grounding (Audio/Visual) | Audio-Visual Association (Audio-Visual) | Diagnostic Traceability |
|---|---|---|---|---|---|---|
| Whole-Caption Paradigm | ||||||
| AuroraCap [ 5 ] | Caption | Global LLM | ✗ | ✗ | ✗ | ✗ |
| VCapsBench [ 62 ] | Caption | Global LLM | ✗ | ✗ | ✗ | ✗ |
| UGC-VideoCap [ 43 ] | Caption | Global LLM | ✗ | ✗ | ✗ | ✗ |
| video-SALMONN 2 [ 34 ] | Caption | Global LLM | ✗ | ✗ | ✗ | ✗ |
| Probe-Based Paradigm | ||||||
| Model | SGC | Visual | Audio | Audio-Visual | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RefUse | CCC | Ref Subject F1 | Ref Scene F1 | Shot F1 | Shot tIoU | Subshot F1 | Subshot tIoU | Event F1 | Event tIoU | Speaker F1 | EVSA F1 | ||
| Proprietary Models | |||||||||||||
| Gemini 3.1-Pro [ 15 ] | 97.36 | 70.39 | 34.43 | 78.60 | 82.68 | 76.24 | 70.74 | 65.27 | 48.76 | 63.32 | 71.92 | 91.18 | 51.46 |
| Gemini 2.5-Pro [ 10 ] | 96.73 | 70.75 | 37.81 | 80.70 | 83.65 | 70.93 | 72.58 | 66.69 | 49.18 | 58.52 | 69.35 | 89.72 | 43.74 |
| Qwen3.5-Omni-Plus [ 32 ] | 96.09 | 66.92 | 29.55 | 73.57 | 79.97 | 71.25 | 72.35 | 65.69 | 48.29 | 53.92 | 67.77 | 91.20 | 40.20 |
| Qwen3.5-Omni-Flash [ 32 ] | 95.58 | 63.10 | 24.31 | 72.27 | 66.07 | 65.65 | 68.99 | 58.94 | 46.52 | 50.68 | 69.34 | 89.91 | 35.58 |
| Model | Visual | Audio | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Shot | Subject | Scene | Subshot | Dialogue | Non-dialogue | ||||
| Recall | Precision | Recall | Precision | Recall | Precision | ||||
| Proprietary Models | |||||||||
| Gemini 3.1-Pro [ 15 ] | 59.05 | 43.78 | 67.07 | 51.78 | 75.62 | 53.29 | 70.92 | 91.18 | 41.49 |
| Gemini 2.5-Pro [ 10 ] | 55.58 | 51.53 | 57.08 | 56.80 | 65.12 | 53.05 | 71.21 | 89.72 | 37.38 |
| Qwen3.5-Omni-Plus [ 32 ] | 54.17 | 45.41 | 62.84 | 50.75 | 69.80 | 50.22 | 70.66 | 91.20 | 36.82 |
| Judge | Holistic | QA | OmniCapBench |
|---|---|---|---|
| Weak | 71.24 | 61.17 | 56.82 |
| Mid | 83.08 | 53.49 | 58.24 |
| Strong | 93.92 | 56.57 | 57.60 |
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Input | Output | Diagnostic Target |
|---|---|---|---|
| Generation | Raw Video | Scored structured units | Native recovery and precision |
| Probing | Predicted JSON | Probe pass/fail | Internal structural validity |
| Verification | + Candidate JSON unit | Accept/reject boolean | Local contradiction detection |
| Bucket | Metric | Required Range |
|---|---|---|
| All | shots/min | |
| All | events/min | |
| min | subjects/min | |
| min | scenes/min | |
| min | total shots | |
| min | total events |
| Component | Fields | Description |
|---|---|---|
| references | ref_id , type , semantic_description , appearance_anchor | Persistent entities or scenes with detailed visual descriptions under id_features |
| events | event_id , type , time_range , content | Timed audio-centered units; spoken dialogue requires exact content.line |
| shots | shot_id , time_range , shot_transfer , visual_description , camera | Visual timeline segments containing fine-grained sub-shot units |
| Link fields | active_events , references_in_shot , reference IDs in content | Checkable structural links governing event-shot membership and reference context |
| Support fields | appearance_anchor , visual_description , time_range | Verifiable evidence fields utilized for local semantic diagnostics |
| Metric | Value |
|---|---|
| Global Counts and Averages | |
| Total Videos | 786 |
| Mean Duration (seconds) | 58.82 |
| Median Duration (seconds) | 28.32 |
| Avg. References / Video | 7.40 |
| Avg. Events / Video | 8.32 |
| Layer | Matching Metric | Default threshold | Tie-breaking |
|---|---|---|---|
| Reference | LLM bipartite matching over detail_description | None (LLM output) | Defensive ID/type/one-to-one check |
| Dialogue event | over content.line | Max line score, one-to-one | |
| Non-dialogue event | Temporal IoU | Max tIoU, one-to-one | |
| Shot | Temporal IoU | Max tIoU, one-to-one |
| Prediction type | Matched | Valid structure | Supported | Credit / penalty | Denominator |
|---|---|---|---|---|---|
| Matched unit | yes | yes | n/a | Full credit; proceeds to structural/semantic checks. | Matched metric denominators |
| Unmatched reference | no | yes/no | no | No credit; retained in diagnostic logs. | Appendix audit |
| Unmatched event | no | yes/no | no | No credit; retained in diagnostic logs. | Appendix audit |
| Unmatched sub-shot | no | yes/no | no | No credit; retained in diagnostic logs. | Appendix audit |
| Invalid or dangling link | no | no | no | Zero credit; penalized as a structural error. | Appendix audit |
| Valid but out-of-scope fact | no | yes | unknown | No manual credit, but no strict penalty; logged for auditing. | Reported diagnostics |
| Model | Checkpoint / Endpoint | Params | Backend | Audio | Frames / FPS | Max Tokens | Duration | Long-Video Strategy |
| Proprietary Models | ||||||||
| Gemini 3.1-Pro | gemini-3.1-pro-preview | – | API | ✓ | no cap / 2 fps | 65536 | All | Single-pass |
| Gemini 2.5-Pro | gemini-2.5-pro | – | API | ✓ | no cap / 2 fps | 65536 | All | Single-pass |
| Qwen3.5-Omni-Plus | qwen3.5-omni-plus | – | API | ✓ | no cap / 2 fps | 32768 | All | Single-pass |
| Qwen3.5-Omni-Flash | qwen3.5-omni-flash | – | API | ✓ | no cap / 2 fps | 32768 | All | Single-pass |
| Open-source Omnimodal Models | ||||||||
| Error Family (Category) | Primary Metric | What triggers this error? | Count | Rate |
|---|---|---|---|---|
| Schema Failure (Gate) | SGC | Output cannot be parsed into references, events, and shots, or is missing required JSON fields | 31 | 3.8% |
| Reference Recovery (Visual) | Ref F1 | Required ground-truth reference is missed, or a hallucinated predicted reference cannot be matched | 3778 | 31.7% |
| Reference-Use (Visual) | RefUse | A matched shot or subshot cites the wrong persistent reference ID | 8649 | 34.1% |
| Identity Drift (Visual) | CCC | A single persistent reference is split, merged, or inconsistently tracked across multiple shots | 2656 | 65.2% |
| Shot Recovery/Alignment (Visual) | Shot F1/tIoU | A visual micro-action or boundary is missed, hallucinated, or severely misaligned in time | 5156 | 26.2% |
| Event Recovery/Alignment (Audio) | Event F1/tIoU | A required audio event is missed, hallucinated, or aligned to a completely wrong time segment | 4218 | 36.2% |
| Failure type | Category | Metric | What it looks like in practice | Rate |
|---|---|---|---|---|
| Broken identity persistence | Visual | CCC | The model correctly describes a character but assigns them a new ID, such as PERSON_2 , when they reappear in a later shot, losing track of their persistent identity. | 65.2% |
| Wrong event reference | Visual | RefUse | An off-screen voice or a background action is attributed to the wrong character ID in the structured output. | 34.1% |
| Event alignment error | Audio | Event tIoU | The model correctly identifies an event, like a dog barking, but aligns it to a completely wrong time segment where the dog is silent. | 42.0% |
| Missing structural field | Gate | SGC | The JSON output is missing required fields like time_range or detail_description , rendering the unit untestable. | 3.8% |
| Unsupported extra event | Audio | Event (Precision) | The model hallucinates an action or event that never actually occurred in the source video. | 36.2% |
| Probe type | Formal Predicate | Failure condition | Count / Prop. |
|---|---|---|---|
| Reference-use validity | Missing aligned region, dangling ID, or reference mismatch after alignment | 8649 / 34.1% | |
| Temporal compatibility | Predicted interval is disjoint from the required shot/event support or violates ordering tolerance | 12240 / 42.0% | |
| Reference-support validity | Missing support field, malformed support field, or support field incompatible with the matched reference | 3778 / 31.7% | |
| Cross-reference integrity | Referenced ID is undefined, inconsistent across shots, or incompatible with the matched persistent reference | 2656 / 65.2% |
| Model | RefUse-shot | RefUse-sub | CCC | Ref R | Ref P | Shot R | Shot P | Shot tIoU | Subshot R | Subshot P | Subshot tIoU |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Proprietary Models | |||||||||||
| Gemini 3.1-Pro [ 15 ] | 68.49 | 68.77 | 34.43 | 77.02 | 70.56 | 71.73 | 88.77 | 70.74 | 75.40 | 60.29 | 48.76 |
| Gemini 2.5-Pro [ 10 ] | 67.01 | 71.42 | 37.81 | 76.96 | 71.39 | 66.29 | 86.27 | 72.58 | 70.26 | 66.89 | 49.18 |
| Qwen3.5-Omni-Plus [ 32 ] | 63.99 | 65.22 | 29.55 | 65.02 | 70.99 | 63.90 | 89.62 | 72.35 | 69.00 | 66.66 | 48.29 |
| Qwen3.5-Omni-Flash [ 32 ] | 59.39 | 61.07 | 24.31 | 59.50 | 69.26 | 55.77 | 89.71 | 68.99 | 61.93 | 60.45 | 46.52 |
| Open-source Omnimodal Models | |||||||||||
| Model | Audio | Audio-Visual | |||||
|---|---|---|---|---|---|---|---|
| Event R | Event P | Event tIoU | Speaker | EVSA-P | EVSA-R | EVSA | |
| Proprietary Models | |||||||
| Gemini 3.1-Pro [ 15 ] | 63.96 | 67.20 | 71.92 | 91.18 | 78.35 | 42.14 | 51.46 |
| Gemini 2.5-Pro [ 10 ] | 58.09 | 64.23 | 69.35 | 89.72 | 74.05 | 35.04 | 43.74 |
| Qwen3.5-Omni-Plus [ 32 ] | 51.22 | 62.53 | 67.77 | 91.20 | 71.11 | 31.82 | 40.20 |
| Qwen3.5-Omni-Flash [ 32 ] | 48.01 | 58.91 | 69.34 | 89.91 | 66.28 | 27.69 | 35.58 |
| Model | Visual | Audio | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Subject Ref | Scene Ref | Shot Fact | Dialogue | Non-dialogue | |||||
| R | P | Mean | R | P | Mean | Mean | Mean | Mean | |
| Proprietary Models | |||||||||
| Gemini 3.1-Pro [ 15 ] | 43.78 | 67.07 | 55.43 | 51.78 | 75.62 | 63.70 | 59.05 | 91.18 | 41.49 |
| Gemini 2.5-Pro [ 10 ] | 51.53 | 57.08 | 54.30 | 56.80 | 65.12 | 60.96 | 55.58 | 89.72 | 37.38 |
| Qwen3.5-Omni-Plus [ 32 ] | 45.41 | 62.84 | 54.13 | 50.75 | 69.80 | 60.27 | 54.17 | 91.20 | 36.82 |
| Model | SGC | Visual | Audio | Audio-Visual | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RefUse | CCC | Ref Subj. F1 | Ref S. F1 | Shot F1 | Shot tIoU | Sub F1 | Sub tIoU | Evt F1 | Evt tIoU | Spk F1 | EVSA F1 | ||
| Duration: 1 min | |||||||||||||
| Gemini 3.1-Pro | 97.46 | 74.13 | 38.41 | 83.09 | 86.54 | 77.90 | 69.76 | 67.55 | 47.73 | 63.09 | 70.65 | 91.15 | 53.30 |
| Gemini 2.5-Pro | 97.57 | 73.40 | 40.60 | 84.80 | 87.39 | 76.61 | 71.72 | 70.94 | 48.36 | 60.46 | 67.46 | 89.08 | 49.31 |
| Qwen3.5-Omni-Plus | 96.90 | 70.94 | 32.71 | 80.28 | 84.99 | 77.60 | 73.57 | 69.65 | 48.73 | 57.12 | 64.57 | 90.77 | 46.87 |
| Qwen3.5-Omni-Flash | 96.26 | 66.83 | 26.81 | 77.37 | 66.70 | 72.86 | 69.94 | 64.57 | 47.07 | 52.72 | 66.90 | 89.34 | 41.65 |
| Model | Visual | Audio | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Shot | Subject | Scene | Subshot | Dialogue | Non-dialogue | ||||
| Recall | Precision | Recall | Precision | Recall | Precision | ||||
| Duration: 1 min | |||||||||
| Gemini 3.1-Pro | 62.84 | 46.74 | 68.18 | 52.62 | 78.01 | 53.90 | 71.75 | 91.15 | 43.39 |
| Gemini 2.5-Pro | 59.70 | 54.97 | 57.11 | 57.66 | 67.97 | 54.19 | 72.40 | 89.08 | 40.03 |
| Qwen3.5-Omni-Plus | 58.92 | 48.27 | 62.41 | 51.83 | 71.96 | 51.88 | 70.08 | 90.77 | 38.23 |
| Model | Text Metric (Prose) | Native Constraints | Human Audit |
|---|---|---|---|
| Pair 1: Qwen3.5-Omni-Flash vs. Qwen3-Omni | |||
| Weaker (Qwen3-Omni) | 52.2 | 37.5 | 21.6 |
| Stronger (Qwen3.5-Omni-Flash) | 47.8 | 62.5 | 78.4 |
| Pair 2: Qwen3.5-Omni-Plus vs. Qwen3.5-Omni-Flash | |||
| Weaker (Qwen3.5-Omni-Flash) | 41.7 | 21.4 | 29.9 |
| Stronger (Qwen3.5-Omni-Plus) | 58.3 | 78.6 | 70.1 |
| Model | Track | SGC | RefUse | CCC | Ref Subj | Shot F1 | Shot tIoU | Evt F1 | Evt tIoU | Speaker | EVSA |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemini 3.1-Pro [ 15 ] | Native | 97.36 | 70.39 | 34.43 | 78.60 | 76.24 | 70.74 | 63.32 | 71.92 | 91.18 | 51.46 |
| Post-hoc | 99.19 | 46.91 | 5.20 | 57.11 | 97.39 | 95.62 | 40.73 | 56.65 | 87.36 | 38.29 | |
| Qwen3-Omni-Instruct [ 49 ] | Native | 89.02 | 40.96 | 9.85 | 54.84 | 49.98 | 67.92 | 41.70 | 40.99 | 88.35 | 13.33 |
| Post-hoc | 96.61 | 40.93 | 9.62 | 49.21 | 96.44 | 95.59 | 12.82 | 54.40 | 81.65 | 12.23 | |
| ASID-Captioner-7B [ 25 ] | Post-hoc | 96.17 | 39.58 | 9.65 | 41.89 | 96.51 | 95.73 | 11.09 | 54.83 | 87.85 | 10.87 |
| Model | Track | Shot | Subj-R | Subj-P | Scene-R | Scene-P | Subsh-R | Subsh-P | Dialogue | Non-dial |
|---|---|---|---|---|---|---|---|---|---|---|
| Gemini 3.1-Pro [ 15 ] | Native | 59.05 | 43.78 | 67.07 | 51.78 | 75.62 | 53.29 | 70.92 | 91.18 | 41.49 |
| Post-hoc | 35.84 | 24.75 | 75.90 | 24.72 | 82.59 | 36.15 | 62.94 | 87.36 | 21.42 | |
| Qwen3-Omni-Instruct [ 49 ] | Native | 34.63 | 40.31 | 62.02 | 36.07 | 70.73 | 29.42 | 57.84 | 88.35 | 20.51 |
| Post-hoc | 28.67 | 28.49 | 71.82 | 29.32 | 77.37 | 33.77 | 50.75 | 81.65 | 13.15 | |
| ASID-Captioner-7B [ 25 ] | Post-hoc | 22.87 | 25.24 | 65.66 | 32.96 | 64.35 | 33.25 | 40.17 | 87.85 | 10.37 |
| Sweep | Metric | Setting | Gemini 3.1-Pro | Gemini 2.5-Pro | Qwen3.5- Omni-Plus | Qwen3.5- Omni-Flash | Qwen3-Omni- Instruct | Qwen3-Omni- Captioner |
|---|---|---|---|---|---|---|---|---|
| Non-dialogue tIoU | Event F1 | 0.10 | 64.02 | 59.54 | 54.63 | 51.43 | 42.29 | 41.88 |
| 0.20 | 63.32 | 58.52 | 53.92 | 50.68 | 41.70 | 41.31 | ||
| 0.40 | 61.86 | 56.34 | 52.38 | 48.95 | 40.71 | 40.07 | ||
| Shot tIoU | Shot F1 | 0.20 | 77.84 | 72.32 | 72.96 | 67.81 | 54.22 | 55.34 |
| 0.30 | 76.24 | 70.93 | 71.25 | 65.65 | 49.98 | 52.29 | ||
| 0.50 | 64.07 | 60.87 | 60.37 | 52.79 | 37.41 | 38.82 |
| Data Sources | URL | License |
| FunQA | https://github.com/Jingkang50/FunQA | MIT |
| LLaVA-Video-178K | https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K | Apache-2.0 |
| Tarsier2-Recap | https://huggingface.co/datasets/omni-research/Tarsier2-Recap-585K | Research Only |
| Video-MME | https://github.com/MME-Benchmarks/Video-MME | Research Only |
| VisionRewardDB-Video | https://huggingface.co/datasets/zai-org/VisionRewardDB-Video | Apache-2.0 |
| LongVideoBench | https://huggingface.co/datasets/longvideobench/LongVideoBench | CC BY-NC-SA 4.0 |