DataVista: Diagnosing Multimodal LLMs on Data Video Understanding
Organizations: The Hong Kong University of Science and Technology (Guangzhou) · Zhejiang University
Abstract
Data video is a media form that integrates data visualization with video narrative, widely adopted in news reporting and business analysis. Compared with general video understanding, data video understanding places greater emphasis on accurately reading data from animated charts, integrating evidence across charts and time, and understanding how narrative organization and visual design communicate information. Yet existing benchmarks target either general videos or static charts, and data video understanding has not been systematically evaluated. We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains. Systematic evaluation of 19 mainstream MLLMs shows that the best-performing model, Gemini-3.1-Pro, achieves 70.0% overall accuracy, still far below human expert performance, with models performing worst on Causal Reasoning and Narrative Structure. Increasing frame counts and adding subtitles mainly benefit data perception and temporal reasoning, with limited gains in narrative understanding. Further analysis of model responses identifies typical failure modes in chart reading, evidence judgment, and instruction understanding. The benchmark is available at https://github.com/HKUSTDial/DataVista.
Figures & tables
| Level 1 | Level 2 | Level 3 | Avg Acc | Overall | |||||
| Model | w/o Sub. | w/ Sub. | w/o Sub. | w/ Sub. | w/o Sub. | w/ Sub. | w/o Sub. | w/ Sub. | |
| Human Baseline | |||||||||
| Human Expert | 97.1 | 89.0 | 86.8 | 91.7 | – | ||||
| Closed-Source Models | |||||||||
| Gemini-3.1-Pro | 78.9 | 81.8 | 66.4 | 71.5 | 55.1 | 57.9 | 68.2 | 71.7 | 70.0 |
| GPT-5 | 76.5 | 78.0 | 61.2 | 66.7 | 55.9 | 58.1 | 66.0 | 68.8 | 67.4 |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Level | Screened | Filtered | Sent to review | Filtered (%) |
|---|---|---|---|---|
| L1 | 9,306 | 6,369 | 2,937 | 68.4 |
| L2 | 9,876 | 7,845 | 2,031 | 79.4 |
| L3 | 9,715 | 7,590 | 2,125 | 78.1 |
| Total | 28,897 | 21,804 | 7,093 | 75.5 |
| Model | Retained | Filtered out | Difference [95% CI] |
|---|---|---|---|
| Gemini-3.1-Pro | 55.6 | 85.6 | +30.0 [25.2, 35.1] |
| GPT-5 | 51.5 | 82.0 | +30.5 [23.6, 37.5] |
| Claude-Sonnet-4.6 | 43.9 | 76.0 | +32.2 [24.9, 39.5] |
| Qwen3-VL-32B-Instruct | 36.5 | 60.0 | +23.4 [15.0, 31.8] |
| InternVL3.5-38B | 42.5 | 65.4 | +22.9 [16.6, 29.3] |
| Theme | Subclass Topic | Videos | Pct.(%) |
| Economy (404) | Market Dashboard | 177 | 18.4 |
| Macro Comparison | 115 | 12.0 | |
| Earnings & Equities | 91 | 9.5 | |
| Ranking & Country Race | 16 | 1.7 | |
| Other | 5 | 0.5 | |
| Culture (173) | Ranking Race | 105 | 10.9 |
| Chart Type | Count | % | Chart Type | Count | % | Chart Type | Count | % |
|---|---|---|---|---|---|---|---|---|
| Bar Chart | 624 | 28.2 | Pictograph / Isotype | 59 | 2.7 | Bubble Chart | 17 | 0.8 |
| Line Chart | 409 | 18.5 | Flow Diagram | 59 | 2.7 | Other | 16 | 0.7 |
| Infographic | 266 | 12.0 | Timeline | 56 | 2.5 | Heatmap | 15 | 0.7 |
| Table | 228 | 10.3 | Annotated Image | 42 | 1.9 | Treemap | 11 | 0.5 |
| Pie / Donut | 112 | 5.1 | Scatter Plot | 33 | 1.5 | Network Graph | 6 | 0.3 |
| Choropleth Map | 99 | 4.5 | Area Chart | 26 | 1.2 |
| Baseline | L1 | L2 | L3 | Overall |
|---|---|---|---|---|
| Random selection | 24.8 | 19.5 | 16.4 | 20.8 |
| Most-frequent answers | 30.8 | 18.3 | 31.5 | 27.5 |
| Model | w/o Sub. | w/ Sub. | Overall |
|---|---|---|---|
| Closed-Source Models | |||
| Gemini-3.1-Pro | 68.2 [67.0, 69.7] | 71.7 [70.5, 73.1] | 70.0 [68.8, 71.3] |
| GPT-5 | 66.0 [64.5, 67.5] | 68.8 [67.5, 70.2] | 67.4 [66.1, 68.8] |
| Gemini-3-Flash | 64.1 [62.6, 65.7] | 68.5 [67.2, 70.0] | 66.3 [65.0, 67.7] |
| GPT-4o | 56.1 [54.9, 57.7] | 58.0 [56.8, 59.4] | 57.1 [56.0, 58.4] |
| Claude-Sonnet-4.6 | 61.3 [60.0, 62.8] | 65.8 [64.5, 67.2] | 63.6 [62.4, 64.9] |
| Model | Economy | Society | Science | Politics | Culture | Avg Acc | Domain mean |
|---|---|---|---|---|---|---|---|
| Gemini-3.1-Pro | 72.1 | 78.2 | 77.1 | 67.9 | 65.9 | 71.7 | 72.2 |
| GPT-5 | 69.2 | 78.8 | 71.9 | 62.8 | 63.1 | 68.8 | 69.2 |
| Claude-Sonnet-4.6 | 67.7 | 68.6 | 66.7 | 66.0 | 58.1 | 65.8 | 65.4 |
| Qwen3-VL-32B-Instruct | 61.9 | 67.3 | 57.3 | 62.8 | 47.5 | 59.8 | 59.4 |
| InternVL3.5-38B | 58.5 | 54.5 | 53.1 | 60.3 | 46.9 | 55.6 | 54.7 |
| Model | Input | L1 | L2 | L3 | Overall |
|---|---|---|---|---|---|
| Gemini-3.1-Pro | 50 frames | 78.9 | 66.4 | 55.1 | 68.2 |
| 50 frames + subtitles | 81.8 | 71.5 | 57.9 | 71.7 | |
| Video + audio | 84.6 | 72.3 | 59.4 | 73.5 | |
| Gemini-3-Flash | 50 frames | 75.6 | 59.7 | 52.3 | 64.1 |
| 50 frames + subtitles | 78.0 | 68.8 | 55.1 | 68.5 | |
| Video + audio | 80.3 | 66.1 | 55.9 | 68.9 |
| Level 1 | Level 2 | Level 3 | Avg Acc | Overall | |||||
|---|---|---|---|---|---|---|---|---|---|
| Model | w/o Sub. | w/ Sub. | w/o Sub. | w/ Sub. | w/o Sub. | w/ Sub. | w/o Sub. | w/ Sub. | |
| Qwen3-VL-8B-Instruct | 59.0 | 64.7 | 28.5 | 33.9 | 38.0 | 38.9 | 44.1 | 48.2 | 46.2 |
| Qwen3-VL-8B-Thinking | 62.5 | 69.6 | 51.5 | 53.2 | 42.7 | 45.8 | 53.4 | 57.8 | 55.6 |
| Qwen3-VL-32B-Instruct | 68.5 | 71.6 | 39.0 | 48.5 | 50.5 | 54.2 | 54.7 | 59.8 | 57.3 |
| Qwen3-VL-32B-Thinking | 72.3 | 77.4 | 55.3 | 64.4 | 48.0 | 50.2 | 60.2 | 65.5 | 62.9 |
| Model | Frames | Overall | L1 | L2 | L3 |
|---|---|---|---|---|---|
| Qwen3-VL-32B | 50 | 59.8 | 71.6 | 48.5 | 54.2 |
| 100 | 62.4 | 76.3 | 50.8 | 54.2 | |
| Change | +2.6 | +4.7 | +2.3 | 0.0 | |
| InternVL3.5-38B | 50 | 55.6 | 65.6 | 48.1 | 48.9 |
| 100 | 55.0 | 67.0 | 46.1 | 46.7 | |
| Change | 0.6 | +1.4 | 2.0 | 2.2 |
| Task | EM | Precision | Recall | F1 | Gold | Selected |
|---|---|---|---|---|---|---|
| Without Subtitles | ||||||
| Causal Reasoning | 27.9 | 94.3 | 71.8 | 79.9 | 2.7 | 2.1 |
| Argument Synthesis | 37.3 | 93.5 | 73.8 | 80.6 | 3.0 | 2.3 |
| Narrative Structure | 42.7 | 92.2 | 78.3 | 82.7 | 2.4 | 2.1 |
| Visual Communication Intent | 53.9 | 95.0 | 81.0 | 85.3 | 2.2 | 1.9 |
| Counterfactual Analysis | 62.6 | 93.1 | 86.6 | 88.1 | 2.3 | 2.1 |
| Two Correct Options | Three Correct Options | |||
| Task | EM | F1 | EM | F1 |
| Without Subtitles | ||||
| Causal Reasoning | 40.0 | 74.8 | 23.4 | 81.8 |
| Argument Synthesis | 37.1 | 69.8 | 40.3 | 83.9 |
| Narrative Structure | 50.4 | 81.7 | 32.9 | 84.1 |
| Visual Communication Intent | 58.1 | 85.3 | 38.1 | 84.2 |
| w/o sub | w/ sub | |||||||
| Model | 3 min | 3–6 min | 6–9 min | 9 min | 3 min | 3–6 min | 6–9 min | 9 min |
| Closed-Source Models | ||||||||
| Gemini-3.1-Pro | 70.2 | 70.8 | 64.5 | 66.7 | 71.4 | 74.9 | 73.7 | 67.0 |
| GPT-5 | 68.3 | 68.6 | 67.9 | 59.4 | 68.3 | 72.4 | 67.0 | 66.9 |
| Gemini-3-Flash | 64.1 | 68.3 | 60.4 | 62.6 | 65.3 | 72.1 | 69.6 | 67.0 |
| GPT-4o | 57.1 | 57.8 | 54.8 | 54.9 | 54.4 | 58.4 | 59.9 | 59.7 |