Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models
Organizations: The University of Melbourne
Abstract
Video understanding is increasingly performed by multi-stage LLM agents that separate temporal grounding, visual observation, and reasoning. Yet these stages are typically evaluated on different benchmarks and distributions, making it difficult to determine where hallucinations originate. We first organize existing benchmarks around these stages and show that their scores provide inconsistent diagnostic signals: stronger stage-level performance does not reliably imply lower downstream hallucination, and even benchmarks targeting the same capability can disagree. We therefore introduce a causal stage-intervention protocol that overwrites individual stages while holding the downstream task fixed. Across 60,008 runs on three video-agent architectures, we find that grounding is the dominant source of downstream error, with roughly four times the causal impact of corrupting visual observations. Successful grounding depends primarily on locating the correct region rather than precise temporal overlap, explaining why standard mIoU metrics poorly predict downstream reliability. We further find that incorrect evidence is substantially more harmful than missing evidence. Finally, auditing existing benchmarks against these interventions reveals that their scores do not reliably predict causal cascade sensitivity and can fail under distribution shift. These results motivate intervention-based, stage-aware evaluation for trustworthy video agents.
Figures & tables
| Stage | Benchmark | Scale | Metric |
| Grounding | NExT-GQA ( Xiao et al., 2024b ) | 990 / 5,553 | Acc, mIoU |
| CG-Bench ( Chen et al., 2025a ) | 1,219 / 12,129 | Acc, mIoU | |
| Watching | LVBench (T.G.) ( Wang et al., 2025 ) | 72 / 179 | Acc |
| Argus ( Rawal et al., 2025a ) | 500 / 500 | ||
| Reasoning | Argus ( Rawal et al., 2025a ) | 500 / 500 | |
| ELV-Halluc ( Lu et al., 2025 ) | 200 / 3,600 | Acc, Diff, SAH |
| Method | Grounding | Watching | W H | Hallucination | |||||||||
| CG-Bench | NeXT-GQA | LVBench* | ARGUS | ELV-Halluc | VideoHallucer | ||||||||
| Acc | mIoU | Acc | mIoU | Acc | Acc | Diff | SAH | Basic | Halluc | Both | |||
| Qwen2-VL-7B | 31.37 | - | 74.59 | - | 54.79 | 55.52 | 82.09 | 8.1 | 5.8 | 6.1 | 76.42 | 64.25 | 46.58 |
| Qwen2.5-VL-7B | 32.24 | - | 72.98 | - | 58.10 | 51.19 | 81.57 | 14.2 | 4.5 | 5.1 | 74.58 | 68.83 | 49.58 |
| Video-R1 | 35.32 | 1.44 | 75.38 | 21.14 | 59.82 | 66.15 | 83.58 | 15.2 | -0.3 | -0.3 | 81.08 | 55.50 | 42.58 |
| VideoMind-7B | 38.40 | 7.10 | 76.79 | 31.40 | 54.79 | 55.63 | 83.06 | 11.4 | 4.2 | 4.6 | 76.25 | 64.58 | 46.83 |
| Grounding condition | EVA | VideoMind | Video-R1 |
| oracle (GT clue span) | 46.6 [44.2, 48.9] | 49.3 [47.0, 51.6] | 47.2 [44.9, 49.6] |
| predicted (agent as deployed) | 34.1 [31.8, 36.4] | 35.9 [33.6, 38.2] | — |
| full (whole video) | 27.0 [25.1, 28.9] | 34.8 [32.5, 37.1] | 33.3 [31.1, 35.4] |
| random (length-matched) | 29.0 [27.0, 31.1] | 31.4 [29.3, 33.4] | 30.6 [28.7, 32.6] |
| off-target (length-matched) | 27.2 [25.3, 29.1] | 30.6 [28.4, 33.0] | 30.7 [28.6, 32.8] |
| Headroom decomposition (paired contrasts, pp) | |||
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Corruption | EVA | VideoMind | Video-R1 | |
| none (clean, oracle window) | 0 | 43.2 [39.3, 47.0] | 49.5 [45.5, 53.5] | 46.5 [42.6, 50.4] |
| omission (frames dropped) | 0.25 | 46.7 [41.1, 52.2] | 50.0 [44.3, 55.6] | 48.0 [42.2, 53.9] |
| 0.50 | 41.0 [37.2, 44.8] | 45.3 [41.4, 49.3] | 47.2 [43.3, 50.9] | |
| 0.75 | 40.3 [35.3, 45.4] | 45.7 [40.0, 51.3] | 43.7 [38.1, 49.3] | |
| fabrication (frames replaced) | 0.25 | 42.7 [37.2, 48.4] | 50.3 [44.9, 55.7] | 45.0 [39.7, 50.6] |
| 0.50 | 40.7 [36.9, 44.5] | 45.5 [41.4, 49.5] | 41.2 [37.3, 45.0] |
| Oracle window dilation | EVA | VideoMind | Video-R1 |
| 45.7 [40.3, 51.2] | 50.3 [44.6, 55.9] | 46.7 [40.7, 52.4] | |
| 46.7 [41.6, 51.7] | 46.7 [41.1, 52.3] | 44.3 [39.0, 49.8] | |
| 43.3 [38.0, 48.6] | 46.0 [40.0, 52.0] | 45.0 [39.7, 50.3] | |
| 38.7 [33.6, 43.7] | 45.3 [39.6, 51.0] | 39.3 [34.0, 44.8] |
| Configuration | Sens. (pp) | 95% CI | Configuration | Sens. (pp) | 95% CI |
| EVA | 16.8 | [11.0, 22.5] | VideoMind | 13.8 | [9.2, 18.4] |
| EVA | 17.8 | [11.9, 23.5] | Video-R1 | 11.5 | [6.8, 16.3] |
| EVA | 15.0 | [10.0, 20.1] | Video-R1 | 16.3 | [11.9, 20.7] |
| VideoMind | 15.8 | [10.5, 21.0] | Video-R1 | 15.5 | [10.5, 20.5] |
| VideoMind | 16.5 | [11.6, 21.4] |
| Method | Size | long-acc. | mIoU | rec.@IoU | acc.@IoU | Averaged Token Usage | ||
| Visual | Total | Vis. Frac. | ||||||
| EVA ( =4, =16) | 7B | 36.93 | 4.41 | 4.97 | 2.63 | 8.1K | 9.2K | 85.8% |
| EVA ( =4, =32) | 7B | 39.00 | 4.47 | 4.90 | 2.57 | 13.2K | 14.7K | 88.3% |
| EVA ( =4, =64) | 7B | 41.00 | 4.83 | 5.35 | 2.69 | 16.5K | 18.5K | 88.3% |
| EVA ( =6, =16) | 7B | 38.17 | 4.58 | 5.28 | 2.93 | 8.2K | 9.5K | 84.5% |
| EVA ( =6, =32) | 7B | 40.20 | 4.51 | 4.99 | 2.59 | 13.3K | 14.9K | 87.8% |
| Method | Size | IoU | IoP | Long Acc | Acc@GQA | Averaged Token Usage | ||||||
| R@0.3 | R@0.5 | mIoU | R@0.3 | R@0.5 | mIoP | Visual | Total | Vis. Frac. | ||||
| EVA ( =4, =16) | 7B | 35.47 | 17.46 | 27.34 | 37.44 | 20.17 | 29.83 | 69.43 | 14.84 | 3.2K | 4.0K | 79.9% |
| EVA ( =4, =32) | 7B | 34.51 | 16.60 | 26.93 | 36.10 | 19.20 | 29.21 | 68.99 | 14.26 | 5.9K | 7.0K | 84.8% |
| EVA ( =4, =64) | 7B | 33.46 | 16.35 | 26.62 | 34.84 | 18.59 | 28.69 | 68.95 | 13.80 | 10.1K | 11.6K | 86.7% |
| EVA ( =6, =16) | 7B | 35.37 | 17.25 | 27.29 | 37.04 | 20.06 | 29.82 | 69.25 | 14.76 | 3.2K | 4.0K | 80.0% |
| EVA ( =6, =32) | 7B | 34.55 | 16.39 | 26.87 | 36.16 | 19.14 | 29.22 | 69.27 | 13.92 | 5.9K | 6.9K | 84.7% |
| Model | Entity | Event | Key Info. | Reasoning | Summarization | Overall |
| Recognition | Understanding | Retrieval | Avg. | |||
| VideoMind-7B | 62.90 | 57.61 | 62.50 | 29.73 | 54.14 | 54.79 |
| VideoMind-2B | 61.29 | 58.70 | 58.33 | 37.84 | 50.00 | 55.25 |
| Video-R1 | 58.06 | 65.22 | 66.67 | 45.95 | 42.86 | 59.82 |
| MARC-3B | 38.71 | 30.43 | 58.33 | 29.73 | 14.29 | 35.16 |
| EVA | 66.13 | 56.52 | 70.83 | 54.05 | 57.14 | 60.27 |
| Model | Param. | |||
| Video-R1 | 16 | – | 0.6825 | 0.8384 |
| 32 | – | 0.6615 | 0.8358 | |
| 64 | – | 0.6651 | 0.8328 | |
| VideoMind-7B | 16 | – | 0.5518 | 0.8299 |
| 32 | – | 0.5563 | 0.8306 | |
| Category | N | Mean IoU | Acc (IoU 0.1) | Acc (IoU 0.1) | |
| 2D Spatial Perception | 180 | 0.052 | 33.1 | 46.2 | +13.1 |
| Entity Cognition | 300 | 0.046 | 27.5 | 23.8 | + 3.7 |
| Entity Perception | 600 | 0.071 | 31.0 | 38.1 | + +7.0 |
| Event Cognition | 300 | 0.087 | 40.3 | 51.4 | +11.1 |
| Event Perception | 600 | 0.069 | 30.6 | 43.0 | +12.4 |
| Hallucination | 240 | 0.054 | 38.8 | 59.0 | +20.2 |
| Models | LLM Size | Visual Details | Object | Action | Declarative Content | Avg Acc | Avg Diff. | SAH Ratio | ||||||||
| In. | Out. | Diff. | In. | Out. | Diff. | In. | Out. | Diff. | In. | Out. | Diff. | |||||
| Existing Results from ELV-Halluc | ||||||||||||||||
| InternVL3-1B | 0.5B | 8 | 11 | 3 | 8.7 | 11 | 2.3 | 8.7 | 12.5 | 3.8 | 11.3 | 8.3 | -3 | 9.9 | 1.5 | 1.6 |
| InternVL3-2B | 1.5B | 7 | 15.5 | 8.5 | 8.7 | 17.2 | 8.5 | 7.2 | 10.5 | 3.3 | 10 | 13 | 3 | 11.1 | 5.8 | 6.3 |
| SmolVLM-2.2B | 1.7B | 0 | 0 | 0 | 3 | 5 | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.5 | 0.5 |
| Qwen2.5VL-3B | 3B | 2.2 | 10.5 | 8.3 | 7.7 | 13.8 | 6.1 | 5 | 8 | 3 | 6 | 6 | 0 | 7.4 | 4.3 | 4.5 |
| Method | Size | Overall | Obj-Rel | Temporal | Sem-Det | Ext-Fact | Ext-NF | Fact-Det | Interact |
| Video-R1 ( =16) | 7B | 81.08 | 72.00 | 55.11 | 93.00 | 91.50 | 89.50 | 98.00 | 69.35 |
| Video-R1 ( =32) | 7B | 71.67 | 28.00 | 51.70 | 84.00 | 91.50 | 92.00 | 94.00 | 67.74 |
| Video-R1 ( =64) | 7B | 64.08 | 4.50 | 60.80 | 52.00 | 91.00 | 93.00 | 94.00 | 70.16 |
| VideoMind ( =16) | 7B | 74.75 | 81.00 | 55.68 | 89.00 | 89.50 | 89.50 | 34.00 | 54.03 |
| VideoMind ( =32) | 7B | 76.25 | 82.50 | 56.25 | 92.00 | 90.00 | 90.00 | 38.00 | 55.65 |
| VideoMind ( =64) | 7B | 79.00 | 82.00 | 61.93 | 88.50 | 92.50 | 92.50 | 60.00 | 54.84 |
| Method | Size | Overall | Obj-Rel | Temporal | Sem-Det | Ext-Fact | Ext-NF | Fact-Det | Interact |
| Video-R1 ( =16) | 7B | 55.50 | 68.50 | 90.34 | 63.50 | 14.00 | 55.50 | 50.00 | 43.55 |
| Video-R1 ( =32) | 7B | 45.50 | 24.00 | 86.36 | 54.50 | 14.00 | 55.00 | 49.00 | 40.32 |
| Video-R1 ( =64) | 7B | 41.08 | 5.00 | 85.23 | 37.50 | 14.50 | 54.00 | 64.00 | 45.97 |
| VideoMind ( =16) | 7B | 63.75 | 76.50 | 88.64 | 75.00 | 26.00 | 63.50 | 41.00 | 69.35 |
| VideoMind ( =32) | 7B | 64.58 | 76.50 | 88.64 | 74.50 | 33.00 | 67.00 | 43.00 | 59.68 |
| VideoMind ( =64) | 7B | 62.33 | 77.00 | 88.64 | 71.50 | 28.00 | 66.00 | 41.00 | 53.23 |
| Method | Size | Overall | Obj-Rel | Temporal | Sem-Det | Ext-Fact | Ext-NF | Fact-Det | Interact |
| Video-R1 ( =16) | 7B | 42.58 | 52.00 | 48.86 | 58.00 | 12.00 | 48.00 | 49.00 | 29.03 |
| Video-R1 ( =32) | 7B | 34.08 | 20.00 | 43.18 | 46.50 | 12.50 | 49.50 | 47.00 | 23.39 |
| Video-R1 ( =64) | 7B | 29.17 | 3.50 | 51.14 | 19.00 | 11.50 | 47.50 | 59.00 | 30.65 |
| VideoMind ( =16) | 7B | 44.58 | 61.00 | 46.59 | 64.50 | 21.00 | 54.50 | 11.00 | 32.26 |
| VideoMind ( =32) | 7B | 46.83 | 61.50 | 49.43 | 67.00 | 27.50 | 58.00 | 13.00 | 27.42 |
| VideoMind ( =64) | 7B | 46.58 | 62.50 | 53.98 | 60.50 | 24.00 | 59.50 | 22.00 | 23.39 |