MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification
Organizations: The University of Sydney
Abstract
A core challenge in short-video fact-checking is identifying which evidence is sufficient to support a verification conclusion. Existing approaches either give the verifier all available evidence, introducing noise, or select evidence by topical relevance, which conflates relatedness with sufficiency. We identify evidential sufficiency as the selection criterion: whether a subset of evidence is adequate to support a confident verdict without redundancy. We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources. We propose a two-layer verification framework that separates claim-video consistency, assessed from internal evidence, from factual verdict determination, which additionally requires external corroboration. On top of it, a sufficiency-driven greedy search assembles evidence until a sufficiency threshold is met and outputs insufficient when the candidate pool is exhausted, rather than forcing a verdict. With Claude Sonnet 4, the method reaches a Macro-F1 of 0.510 using 4.5 evidence units on average (16% of the full evidence set), statistically indistinguishable from the full-evidence baseline (0.518 with 27.7 units), while significantly improving recognition of insufficient cases over the same search without abstention. The efficiency result replicates with GPT-5.5 and holds only partially with an open-weight Qwen2.5-72B verifier. Ablations show that external evidence is indispensable for factual determination, while internal video evidence grounds the verdict in claim-video consistency. These findings suggest that evidence-efficient verification is achievable, and that explicit abstention is needed when evidence is genuinely inadequate.
Figures & tables
| Verdict | Count | Proportion |
|---|---|---|
| Supported | 97 | 49.7% |
| Refuted | 39 | 20.0% |
| Insufficient | 59 | 30.3% |
| Total | 195 | 100% |
| Modality | Code | Total Units | Avg / Sample | Proportion |
|---|---|---|---|---|
| Visual keyframes | V | 1,207 | 6.2 | 21.9% |
| ASR transcripts | A | 2,847 | 14.6 | 51.7% |
| External search | E | 1,456 | 7.5 | 26.4% |
| Total | 5,510 | 28.3 | 100% |
| Benchmark | Media | Samples | Lang | Labels | Evidence (granularity) | Selection focus |
|---|---|---|---|---|---|---|
| FEVER | Text | 185K | EN | 3-way | Text sentences | No |
| AVeriTeC | Text | 4.6K | EN | 4-way | QA pairs + URLs | No |
| MOCHEG | Multimodal web | 15.6K | EN | 3-way | Paragraphs + images | No |
| FakeSV | Short video | 3.7K ∗ | ZH | Binary | V + A + M (per sample) | No |
| MiniVer-V | Short video | 195 | ZH/EN | 3-way | V + A + E (per unit) | Yes |
| Method | Macro-F1 | Accuracy | Avg Units | Abst. Prec. | Abst. Recall | OverConf. |
|---|---|---|---|---|---|---|
| B1 Full | 0.518 | 56.9% | 27.7 | 0.500 | 0.254 | 0.746 |
| B2 Top-3 | 0.466 | 49.7% | 3.0 | 0.378 | 0.525 | 0.475 |
| B2 Top-5 | 0.459 | 50.3% | 5.0 | 0.370 | 0.339 | 0.661 |
| B2 Top-7 | 0.460 | 51.3% | 7.0 | 0.350 | 0.237 | 0.763 |
| B3 CoT | 0.453 | 57.4% | 9.4 | 1.000 | 0.051 | 0.949 |
| B4 Greedy | 0.469 | 52.3% | 4.1 | 0.333 | 0.203 | 0.797 |
| Setting | Macro-F1 | Acc. | Units | F1 [95% CI] |
|---|---|---|---|---|
| B5 (proposed) | 0.510 | 54.4% | 4.5 | — |
| w/o Abstention (= B4) | 0.469 | 52.3% | 4.1 | 0.041 [ 0.10, +0.02] |
| Component ablations on B4 ( F1 relative to B4) | ||||
| w/o External | 0.273 | 32.3% | 7.3 | 0.196 [ 0.28, 0.11] |
| w/o V+A (E-only) | 0.484 | 53.3% | 3.5 | +0.016 [ 0.05, +0.08] |
| w/o ASR | 0.465 | 51.3% | 4.6 | 0.004 [ 0.07, +0.06] |
| Metric | Value |
|---|---|
| Avg Jaccard similarity | 0.882 |
| Avg Precision (B5 human / B5) | 0.899 |
| Avg Recall (B5 human / human) | 0.970 |
| Exact match rate | 71.8% |
| Human B5 rate | 92.8% |
| Avg human set size | 3.74 |
| Verifier | B1 F1 | B5 F1 | B5/B1 | Units | Ins-F1 B4 B5 |
|---|---|---|---|---|---|
| Claude Sonnet 4 | 0.518 | 0.510 | 98.5% | 4.5 (16%) | 0.253 0.383 |
| GPT-5.5 | 0.586 | 0.559 | 95.3% | 4.9 (18%) | 0.407 0.508 |
| Qwen2.5-72B | 0.580 | 0.489 | 84.3% | 2.4 (9%) | 0.413 0.423 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Set Size | Count | Proportion | Cumulative |
|---|---|---|---|
| 1 | 32 | 16.4% | 16.4% |
| 2 | 47 | 24.1% | 40.5% |
| 3 | 31 | 15.9% | 56.4% |
| 4 | 31 | 15.9% | 72.3% |
| 5 | 18 | 9.2% | 81.5% |
| 6 | 9 | 4.6% | 86.2% |
| Supported | Refuted | Insufficient | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | P | R | F1 | P | R | F1 | P | R | F1 |
| B1 Full | 0.632 | 0.742 | 0.682 | 0.471 | 0.615 | 0.533 | 0.500 | 0.254 | 0.337 |
| B2 Top-3 | 0.679 | 0.546 | 0.606 | 0.371 | 0.333 | 0.351 | 0.378 | 0.525 | 0.440 |
| B2 Top-5 | 0.606 | 0.649 | 0.627 | 0.405 | 0.385 | 0.395 | 0.370 | 0.339 | 0.354 |
| B2 Top-7 | 0.615 | 0.691 | 0.650 | 0.413 | 0.487 | 0.447 | 0.350 | 0.237 | 0.283 |
| B3 CoT | 0.596 | 0.866 | 0.706 | 0.490 | 0.641 | 0.556 | 1.000 | 0.051 | 0.097 |
| Backbone | Method | Macro-F1 | Accuracy | Units | Ins-F1 |
|---|---|---|---|---|---|
| Claude Sonnet 4 | B1 Full | 0.518 | 0.569 | 27.7 | 0.337 |
| B2 Top-5 | 0.459 | 0.503 | 5.0 | 0.354 | |
| B4 Greedy | 0.469 | 0.523 | 4.1 | 0.253 | |
| B5 Abstention | 0.510 | 0.544 | 4.5 | 0.383 | |
| GPT-5.5 (closed) | B1 Full | 0.586 | 0.651 | 27.7 | 0.464 |
| B2 Top-5 | 0.551 | 0.605 | 5.0 | 0.404 |
| Method | Macro-F1 | Units | Macro-F1 [95% CI] | Ins-F1 [95% CI] | CV- Macro-F1 |
|---|---|---|---|---|---|
| Claude Sonnet 4 ( ) | |||||
| B1 Full | 0.518 | 27.7 | 0.008 [ 0.08, +0.06] | +0.046 [ 0.08, +0.18] | — |
| B2 Top-3 | 0.466 | 3.0 | +0.044 [ 0.03, +0.12] | 0.056 [ 0.18, +0.07] | — |
| B2 Top-5 | 0.459 | 5.0 | +0.051 [ 0.02, +0.13] | +0.029 [ 0.11, +0.17] | — |
| B2 Top-7 | 0.460 | 7.0 | +0.050 [ 0.02, +0.12] | +0.101 [ 0.02, +0.22] | — |
| B3 CoT | 0.453 | 9.4 | +0.057 [ 0.02, +0.13] | +0.287 [+0.16, +0.41] ∗ | — |
| True Pred | Supported | Refuted | Insufficient |
|---|---|---|---|
| Supported (97) | 66 | 8 | 23 |
| Refuted (39) | 7 | 17 | 15 |
| Insufficient (59) | 30 | 6 | 23 |
| True Pred | Supported | Refuted | Insufficient |
|---|---|---|---|
| Supported (97) | 72 | 16 | 9 |
| Refuted (39) | 9 | 24 | 6 |
| Insufficient (59) | 33 | 11 | 15 |
| Method | Consistent | Inconsistent | Unclear |
|---|---|---|---|
| B1 Full | 168 (86.2%) | 18 (9.2%) | 9 (4.6%) |
| B2 Top-5 | 117 (60.0%) | 9 (4.6%) | 69 (35.4%) |
| B3 CoT | 141 (72.3%) | 54 (27.7%) | 0 (0.0%) |
| B4 Greedy | 83 (42.6%) | 13 (6.7%) | 99 (50.8%) † |
| B5 Abstention | 79 (40.5%) | 14 (7.2%) | 102 (52.3%) † |
| Setting | F1-Sup | F1-Ref | F1-Ins | Macro-F1 |
| B5 (proposed) | 0.660 | 0.486 | 0.383 | 0.510 |
| B4 (= B5 w/o Abstention) | 0.654 | 0.500 | 0.253 | 0.469 |
| B4 w/o External | 0.096 | 0.270 | 0.453 | 0.273 |
| B4 w/o V+A (E-only) | 0.660 | 0.453 | 0.340 | 0.484 |
| B4 w/o ASR | 0.638 | 0.494 | 0.263 | 0.465 |
| B4 w/o Pruning | 0.660 | 0.506 | 0.265 | 0.477 |
| Topic | Count | Proportion |
|---|---|---|
| International Politics & Military | 45 | 23.1% |
| Domestic News & Society | 30 | 15.4% |
| Health & Medicine | 29 | 14.9% |
| Science & Technology | 24 | 12.3% |
| Natural Disaster & Accident | 22 | 11.3% |
| Food Safety & Consumer | 14 | 7.2% |
| Platform | Count | Proportion |
|---|---|---|
| Douyin | 112 | 57.4% |
| Xiaohongshu | 64 | 32.8% |
| X (Twitter) | 8 | 4.1% |
| Bilibili | 3 | 1.5% |
| 3 | 1.5% | |
| News outlets (CNN, BBC, Xinhua, Southern Daily) | 5 | 2.6% |
| Modality | Total | Mean/Sample | Proportion | Source |
|---|---|---|---|---|
| Visual (V) | 1,207 | 6.2 | 21.9% | Claude Sonnet 4 |
| ASR (A) | 2,847 | 14.6 | 51.7% | Whisper (base) |
| External (E) | 1,456 | 7.5 | 26.4% | Web Search |
| Total | 5,510 | 28.3 | 100% |
| Modality | Units Selected | Proportion | Contrast with Pool |
|---|---|---|---|
| Visual (V) | 116 | 15.9% | 21.9% in pool |
| ASR (A) | 237 | 32.5% | 51.7% in pool |
| External (E) | 376 | 51.6% | 26.4% in pool |
| Total | 729 | 100% |
| Method | Avg (All) | Avg (Correct) | Avg (Incorrect) |
|---|---|---|---|
| B1 Full | 27.7 | 27.7 | 27.7 |
| B2 Top-3 | 3.0 | 3.0 | 3.0 |
| B2 Top-5 | 5.0 | 5.0 | 5.0 |
| B2 Top-7 | 7.0 | 7.0 | 7.0 |
| B3 CoT | 9.4 | 9.2 | 9.6 |
| B4 Greedy | 4.1 | 3.5 | 4.8 |