Data video is a media form that integrates data visualization with video narrative, widely adopted in news reporting and business analysis. Compared with general video understanding, data video understanding places greater emphasis on accurately reading data from animated charts, integrating evidence across charts and time, and understanding how narrative organization and visual design communicate information. Yet existing benchmarks target either general videos or static charts, and data video understanding has not been systematically evaluated. We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains. Systematic evaluation of 19 mainstream MLLMs shows that the best-performing model, Gemini-3.1-Pro, achieves 70.0% overall accuracy, still far below human expert performance, with models performing worst on Causal Reasoning and Narrative Structure. Increasing frame counts and adding subtitles mainly benefit data perception and temporal reasoning, with limited gains in narrative understanding. Further analysis of model responses identifies typical failure modes in chart reading, evidence judgment, and instruction understanding. The benchmark is available at https://github.com/HKUSTDial/DataVista.
Figures & tables
Figure 1: Three-level capability hierarchy and construction pipeline of DataVista .
Figure 2: Overview of the DataVista construction process, including data collection, human-AI collaborative annotation, and question design.
Figure 3: Overview of the DataVista dataset. (a) Video category hierarchy showing the distribution across five major domains and their sub-categories. (b) Distribution of video durations in minutes. (c) Question types organized by evaluation level.
Level 1
Level 2
Level 3
Avg Acc
Overall
Model
w/o Sub.
w/ Sub.
w/o Sub.
w/ Sub.
w/o Sub.
w/ Sub.
w/o Sub.
w/ Sub.
Human Baseline
Human Expert
97.1
89.0
86.8
91.7
–
Closed-Source Models
Gemini-3.1-Pro
78.9
81.8
66.4
71.5
55.1
57.9
68.2
71.7
70.0
GPT-5
76.5
78.0
61.2
66.7
55.9
58.1
66.0
68.8
67.4
Table 1: Exact-match accuracy (%) on DataVista . w/o Sub. and w/ Sub. use the same 50 sampled frames without audio, without and with subtitles; Avg Acc weights the three levels by question count, and Overall pools both conditions. Human evaluation uses complete videos with original audio.
Figure 4: Exact-match accuracy (%) of five models under question-only and frame inputs, by level. Colored segments show the change after adding subtitles (percentage points).
Figure 5: Exact-match accuracy across frame budgets without subtitles or audio.
Figure 6: Per-type exact-match accuracy (%) of five representative models: (a) without subtitles and (b) with subtitles.
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: An example L1 question in DataVista (Data Fact Reading). The question requires extracting a specific numerical value from a trend chart. The correct answer (C) is directly grounded in the visual label for the year 2022, while distractors correspond to the historical baseline (A), sector-specific tax reporting (B), and the US policy shift (D).
Figure 8: An example L2 question in DataVista . The video argues that headlines regarding Pakistan and Iran helped spark a late-day market recovery. The task is to select all essential chart observations (A, B, D) that support the argument.
Figure 9: An example L3 question in DataVista (Narrative Structure). This task evaluates the model’s ability to synthesize global narrative themes. The model must distinguish the sector portrayed as still resilient (B: healthcare and leisure/hospitality) from the other narrative roles: the software-employment downturn (A), the decline in construction employment (C), and manufacturing, construction, and trade, which the video explicitly describes as under pressure (D).
Figure 10: The human review interface for video screening. Annotators can watch the video, view AI-generated annotations, and confirm or correct the data video classification.
Level
Screened
Filtered
Sent to review
Filtered (%)
L1
9,306
6,369
2,937
68.4
L2
9,876
7,845
2,031
79.4
L3
9,715
7,590
2,125
78.1
Total
28,897
21,804
7,093
75.5
Appendix
Table 2: Text-only screening results by capability level.
Model
Retained
Filtered out
Difference [95% CI]
Gemini-3.1-Pro
55.6
85.6
+30.0 [25.2, 35.1]
GPT-5
51.5
82.0
+30.5 [23.6, 37.5]
Claude-Sonnet-4.6
43.9
76.0
+32.2 [24.9, 39.5]
Qwen3-VL-32B-Instruct
36.5
60.0
+23.4 [15.0, 31.8]
InternVL3.5-38B
42.5
65.4
+22.9 [16.6, 29.3]
Appendix
Table 3: Text-only accuracy (%) on retained and filtered-out questions. Differences are filtered-out minus retained, in percentage points, with 95% video-cluster bootstrap confidence intervals.
Theme
Subclass Topic
Videos
Pct.(%)
Economy (404)
Market Dashboard
177
18.4
Macro Comparison
115
12.0
Earnings & Equities
91
9.5
Ranking & Country Race
16
1.7
Other
5
0.5
Culture (173)
Ranking Race
105
10.9
Appendix
Table 4: Distribution of 5 major themes and 22 subclass topics in DataVista .
Figure 11: Distribution of video durations (30-second bins).
Figure 12: Distribution of video view counts ( log10 scale).
Figure 13: Monthly distribution of video upload dates.
Chart Type
Count
%
Chart Type
Count
%
Chart Type
Count
%
Bar Chart
624
28.2
Pictograph / Isotype
59
2.7
Bubble Chart
17
0.8
Line Chart
409
18.5
Flow Diagram
59
2.7
Other
16
0.7
Infographic
266
12.0
Timeline
56
2.5
Heatmap
15
0.7
Table
228
10.3
Annotated Image
42
1.9
Treemap
11
0.5
Pie / Donut
112
5.1
Scatter Plot
33
1.5
Network Graph
6
0.3
Choropleth Map
99
4.5
Area Chart
26
1.2
Appendix
Table 5: Distribution of chart types across 961 videos (2,210 total entries; ≈ 2.30 per video). Count is the occurrence count for each type; % is its share of all entries. “Other” subsumes miscellaneous and unclassified types.
Figure 14: An example L1 question in DataVista (Data Fact Reading)
Figure 15: An example L1 question in DataVista (Chart Element Recognition).
Figure 16: An example L2 question in DataVista (Argument Synthesis).
Figure 17: An example L2 question in DataVista (Cross-Temporal Change).
Figure 18: An example L3 question in DataVista (Visual Communication Intent).
Figure 19: An example L3 question in DataVista (Counterfactual Analysis).
Figure 20: An example L3 question in DataVista (Narrative Structure).
Baseline
L1
L2
L3
Overall
Random selection
24.8
19.5
16.4
20.8
Most-frequent answers
30.8
18.3
31.5
27.5
Appendix
Table 6: Exact-match accuracy (%) of the statistical baselines. Random selection reports the theoretical expectation; most-frequent answers use video-level five-fold cross-validation.
Model
w/o Sub.
w/ Sub.
Overall
Closed-Source Models
Gemini-3.1-Pro
68.2 [67.0, 69.7]
71.7 [70.5, 73.1]
70.0 [68.8, 71.3]
GPT-5
66.0 [64.5, 67.5]
68.8 [67.5, 70.2]
67.4 [66.1, 68.8]
Gemini-3-Flash
64.1 [62.6, 65.7]
68.5 [67.2, 70.0]
66.3 [65.0, 67.7]
GPT-4o
56.1 [54.9, 57.7]
58.0 [56.8, 59.4]
57.1 [56.0, 58.4]
Claude-Sonnet-4.6
61.3 [60.0, 62.8]
65.8 [64.5, 67.2]
63.6 [62.4, 64.9]
Appendix
Table 7: Exact-match accuracy (%) and 95% intervals.
Model
Economy
Society
Science
Politics
Culture
Avg Acc
Domain mean
Gemini-3.1-Pro
72.1
78.2
77.1
67.9
65.9
71.7
72.2
GPT-5
69.2
78.8
71.9
62.8
63.1
68.8
69.2
Claude-Sonnet-4.6
67.7
68.6
66.7
66.0
58.1
65.8
65.4
Qwen3-VL-32B-Instruct
61.9
67.3
57.3
62.8
47.5
59.8
59.4
InternVL3.5-38B
58.5
54.5
53.1
60.3
46.9
55.6
54.7
Appendix
Table 8: Per-domain exact-match accuracy (%) with subtitles. Domain mean gives equal weight to the five domains.
Model
Input
L1
L2
L3
Overall
Gemini-3.1-Pro
50 frames
78.9
66.4
55.1
68.2
50 frames + subtitles
81.8
71.5
57.9
71.7
Video + audio
84.6
72.3
59.4
73.5
Gemini-3-Flash
50 frames
75.6
59.7
52.3
64.1
50 frames + subtitles
78.0
68.8
55.1
68.5
Video + audio
80.3
66.1
55.9
68.9
Appendix
Table 9: Exact-match accuracy (%) across input conditions. The frame conditions use the same 50 frames, without audio; audiovisual input includes the complete video and original audio.
Level 1
Level 2
Level 3
Avg Acc
Overall
Model
w/o Sub.
w/ Sub.
w/o Sub.
w/ Sub.
w/o Sub.
w/ Sub.
w/o Sub.
w/ Sub.
Qwen3-VL-8B-Instruct
59.0
64.7
28.5
33.9
38.0
38.9
44.1
48.2
46.2
Qwen3-VL-8B-Thinking
62.5
69.6
51.5
53.2
42.7
45.8
53.4
57.8
55.6
Qwen3-VL-32B-Instruct
68.5
71.6
39.0
48.5
50.5
54.2
54.7
59.8
57.3
Qwen3-VL-32B-Thinking
72.3
77.4
55.3
64.4
48.0
50.2
60.2
65.5
62.9
Appendix
Table 10: Exact-match accuracy (%) of Qwen3-VL Instruct and Thinking variants.
Model
Frames
Overall
L1
L2
L3
Qwen3-VL-32B
50
59.8
71.6
48.5
54.2
100
62.4
76.3
50.8
54.2
Change
+2.6
+4.7
+2.3
0.0
InternVL3.5-38B
50
55.6
65.6
48.1
48.9
100
55.0
67.0
46.1
46.7
Change
− 0.6
+1.4
− 2.0
− 2.2
Appendix
Table 11: Exact-match accuracy (%) at 50 and 100 frames with subtitles.
Task
EM
Precision
Recall
F1
Gold
Selected
Without Subtitles
Causal Reasoning
27.9
94.3
71.8
79.9
2.7
2.1
Argument Synthesis
37.3
93.5
73.8
80.6
3.0
2.3
Narrative Structure
42.7
92.2
78.3
82.7
2.4
2.1
Visual Communication Intent
53.9
95.0
81.0
85.3
2.2
1.9
Counterfactual Analysis
62.6
93.1
86.6
88.1
2.3
2.1
Appendix
Table 12: Multi-select performance (%). Exact match is computed on multi-select items only. Gold and Selected denote the mean numbers of gold and predicted options. Empty or invalid predictions receive zero scores.
Two Correct Options
Three Correct Options
Task
EM
F1
EM
F1
Without Subtitles
Causal Reasoning
40.0
74.8
23.4
81.8
Argument Synthesis
37.1
69.8
40.3
83.9
Narrative Structure
50.4
81.7
32.9
84.1
Visual Communication Intent
58.1
85.3
38.1
84.2
Appendix
Table 13: Multi-select EM and F1 (%) by the number of correct options, using the same aggregation as Table 12 .
w/o sub
w/ sub
Model
≤ 3 min
3–6 min
6–9 min
> 9 min
≤ 3 min
3–6 min
6–9 min
> 9 min
Closed-Source Models
Gemini-3.1-Pro
70.2
70.8
64.5
66.7
71.4
74.9
73.7
67.0
GPT-5
68.3
68.6
67.9
59.4
68.3
72.4
67.0
66.9
Gemini-3-Flash
64.1
68.3
60.4
62.6
65.3
72.1
69.6
67.0
GPT-4o
57.1
57.8
54.8
54.9
54.4
58.4
59.9
59.7
Appendix
Table 14: Model accuracy (%) across four video-duration intervals, reported separately under w/o sub and w/ sub .
Figure 21: An instruction-understanding error of Gemini-3.1-Pro. The model is distracted by on-chart text and substitutes the printed one-year change for the requested time-window increase.
Figure 22: An instruction-understanding error of Gemini-3.1-Pro. The question admits only facts shown on the chart, yet the model still treats the narrated 2034 revenue target of $2 billion as supporting evidence.