Data videos communicate data insights through dynamic charts, voice narration, and synchronized animations, and have become a widely adopted form of data storytelling. However, producing them requires expertise in data analysis, narrative design, and video editing. Static visualization tools lack narrative and animation capabilities; authoring tools rely on pre-prepared charts rather than raw data; and pixel-level models generate videos end-to-end but cannot guarantee data accuracy or provenance. End-to-end automatic generation faces two core challenges: how to uniformly represent charts, narration, and animations together with their temporal relationships, and how to efficiently search a vast design space for narrative-coherent compositions. We present DataMagic, which authors data videos from raw tabular data through declarative multi-agent orchestration. First, the declarative specification DVSpec unifies charts, narration, and animations with data-bound references and declarative synchronization, ensuring data provenance and automatic audio-visual alignment. Second, a "Generate-then-Orchestrate" multi-agent strategy generates candidate scenes in parallel and then optimizes narrative coherence through global orchestration. DVSpec provides a shared state for three complementary interaction modes, bridging full automation with fine-grained human control. Evaluations on 109 real-world samples show that even the most advanced LLM (e.g., GPT-5) achieves only 2.13/5 with execution success rates between 48.62% and 86.24%; DataMagic improves quality to 3.89 (+83%) with success rates above 95%, with the most significant gains in animation and narrative dimensions. A user study shows that, compared to a conversational LLM workflow, DataMagic improves creation efficiency (79.7% reduction in task time) and reduces perceived cognitive load. Project page: https://github.com/HKUSTDial/DataMagic.
Figures & tables
Figure 1 : DVSpec structure and rendering flow.
Figure 2 : A DVSpec specification example for one scene, where the numbered markers (1–4) link each declarative field (chart, data, narration, and emphasis animation) to its effect in the rendered scene.
Figure 3 : System architecture of DataMagic .
Figure 4 : DataMagic web interface, showing (A–D) the main UI components and (D1–D3) three key interaction flows.
Evaluation Dimensions
Method
Exec Rate (%)
Intent
Insight
Narrative
Animation
Aesthetic
Avg. Score
Direct Generation Methods
DeepSeek-V3.2
48.62%
1.95
1.98
1.88
1.65
2.09
1.91
Gemini-2.5-Pro
66.06%
2.38
2.25
2.07
1.94
2.44
2.22
GPT-5
86.24%
2.36
2.22
2.05
1.84
2.17
2.13
Claude-Sonnet-4
84.40%
2.28
2.01
1.98
1.91
2.71
2.18
Table 1 : End-to-End Video Generation Performance Comparison.
Variant
Intent
Insight
Narrative
Animation
Aesthetic
Avg. Score
DataMagic (Full)
3.79
3.37
3.84
4.39
4.05
3.89
w/o Story Planner
3.42
3.16
3.26
3.79
3.58
3.44
w/o Orchestration
3.32
3.21
3.42
3.95
3.79
3.54
Table 2 : Ablation study of DataMagic variants.
Figure 5 : User-study results ( N=12 ). DataMagic reduced task time from 39.2 to 8.0 min (79.7%) and lowered ratings on five NASA-TLX workload dimensions, with no significant Performance difference. Error bars: 95% CIs; Wilcoxon signed-rank tests: ∗p<.05 , ∗∗p<.01 , ∗∗∗p<.001 .
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Comparative Analysis of Video Generation Quality Across Different Methods
Figure 7 : A representative high-quality case generated by DataMagic (Agriculture, overall score: 4.6). Four key frames are shown with their corresponding narration (italicized). Underlined spans in the narration are the index triggers in DVSpec that fire the co-occurring visual highlights, demonstrating precise audio-visual synchronization without manual timeline adjustment.
Figure 8 : High-quality Case 1 (Environmental Science): High-fidelity tracking of meteorological fluctuations and solar radiation swings.
Figure 9 : High-quality Case 2 (Gender-based Compensation): Visualization of complex gender-based compensation disparities across divisions.
Figure 10 : Low-quality Case 1 (Agriculture): Analysis of layout truncation and numerical discrepancies in DataMagic under extreme data distributions.
Figure 11 : Low-quality Case 2 (Energy): Analysis of systemic failures in baseline models across layout logic and cross-modal alignment.
Source
Datasets
Samples
Small (<100 rows)
Med. (100-1K)
Large (>1K rows)
T2R-bench
36
85
4
14
18
DAComp-DA
24
24
2
7
15
Total
60
109
6
21
33
Appendix
Table 3 : Dataset sources and scale distribution.
Dimension
Min
Median
Max
Mean
Rows
23
1,436
150,000
17,065.58
Columns
3
16
555
32.42
Numeric columns
1
8
552
24.25
Categorical columns
0
7
57
8.17
Appendix
Table 4 : Dataset dimension statistics.
Query Type
Count
Percentage
Typical Patterns
Trend analysis
12
11.0%
Time-series changes; Periodic patterns
Comparative analysis
15
13.8%
Cross-category comparison; Ranking analysis
Distribution analysis
18
16.5%
Value distribution; Proportion
Correlation analysis
33
30.2%
Variable relationships; Correlation
Comprehensive analysis
31
28.4%
Multi-dimensional analysis
Total
109
100.0%
–
Appendix
Table 5 : Query type distribution of the sample dataset.
Domains
Sub-domains
Technology and Engineering
Electronics and Automation Manufacturing; Academic Research; Energy Production and Power Systems; Automotive Industry
Environmental Management
Environmental Protection; Agriculture and Forestry; Resource Management
Transportation Logistics
Communication and Digital Infrastructure; Transportation Networks and Logistics Management
Social Policy Administration
Education Policy and Public Education; Government Administration and Public Sector Services; Labor and Employment Administration; Healthcare Systems and Public Health; Demographics and Social Development
Commercial Services and Markets
Retail Trade and E-commerce Platforms; Tourism and Hospitality Services; Digital Entertainment and Gaming; Food and Beverage Services; Real Estate and Housing Market; Business Management and Supply Chain
Financial Economics
Economic Development and International Trade; Banking and Financial Services
Appendix
Table 6 : Domains and sub-domains.
Figure 12 : The online evaluation platform used for human expert validation. The interface integrates the original user query, synchronized video playback, a multi-dimensional scoring guide, and a feedback module for documenting qualitative observations.