Product-centric advertisement video generation aims to create promotional videos that preserve fine-grained product identity while presenting selling points through coherent multi-shot narratives. However, this emerging task remains underexplored due to the lack of large-scale advertisement-specific datasets and comprehensive evaluation frameworks. To address this gap, we introduce \textbf{AdSpark}, a large-scale dataset and benchmark for product-centric advertisement video generation, based on data from a major e-commerce platform. \textit{AdSpark-300K} contains approximately 300K reference image--prompt--video triplets, comprising a real-world subset and a synthetic subset. Each sample provides structured advertisement annotations, including product identity annotations, selling-point descriptions, creative plans, and aligned audio scripts, enabling models to learn product preservation and advertisement-oriented visual storytelling. We further propose \textit{AdSpark-Bench}, a diagnostic benchmark that evaluates generated advertisements across six dimensions, including visual quality, product fidelity, instruction adherence, temporal coherence, audio alignment, and advertisement effectiveness. Based on AdSpark-Bench, we evaluate representative models, revealing key challenges in product preservation, multi-shot storytelling, and selling-point visualization. Experiments with AdSpark-300K-finetuned models further validate the effectiveness of our dataset. AdSpark provides a unified dataset and benchmark for future research, and we will release the dataset upon acceptance.
Figures & tables
Figure 1: Overview of AdSpark-300K and AdSpark-Bench. AdSpark-300K is a large-scale, high-quality product-centric advertisement video dataset spanning real-world and synthetic advertisements, with structured annotations across multiple dimensions. AdSpark-Bench provides a comprehensive evaluation of general video quality and advertisement-specific capabilities.
Dataset
Domain
Task
Multi- shot
Creative Plans
Selling Points
Identity Annotation
Audio Script
Clips
Resolution
Average Length(s)
MSRVTT Xu et al. (2016)
Open
T2V
✗
✗
✗
✗
✗
10K
240P
14.4
WebVid-10M Bain et al. (2021)
Open
T2V
✗
✗
✗
✗
✗
10M
360P
18.7
HD-VG-130M Wang et al. (2023a)
Open
T2V
✗
✗
✗
✗
✗
130M
720P
4.9
Panda-70M Chen et al. (2024)
Open
T2V
✗
✗
✗
✗
✗
70M
720P
8.6
InternVid Wang et al. (2023b)
Open
T2V
✗
✗
✗
✗
✗
234M
720P
11.7
OpenHumanVid Li et al. (2024)
Human
T2V
✗
✗
✗
✗
✗
52.3M
720P
4.9
Table 1: Comparison with existing video generation datasets. Most existing datasets focus on generic text-to-video generation (T2V), subject-to-video generation (S2V), or image-to-video generation (I2V), while lacking advertisement-specific annotations. In contrast, AdSpark-300K targets PC-AVG, providing structured advertisement annotations, including product identity, selling points, creative plans (style, scene, and shot design), and audio scripts.
Benchmark
Visual Quality
Script Adherence
Temporal Coherence
Product Fidelity
Shot Compliance
Audio Alignment
Selling-point Realization
Advertisement Effectiveness
Make-a-Video-Eval Singer et al. (2022)
✓
✓
✗
✗
✗
✗
✗
✗
FETV Liu et al. (2024b)
✓
✓
✓
✗
✗
✗
✗
✗
T2VScore Wu et al. (2024)
✓
✓
✓
✗
✗
✗
✗
✗
EvalCrafter Liu et al. (2024a)
✓
✓
✓
✗
✗
✗
✗
✗
VBench Huang et al. (2023)
✓
✓
✓
✗
✗
✗
✗
✗
VBench++ Huang et al. (2024)
✓
✓
✓
✗
✗
✗
✗
✗
Table 2: Comparison of AdSpark-Bench with existing video generation benchmarks. Existing benchmarks focus on general video quality, while AdSpark-Bench further evaluates product fidelity, selling-point realization, and advertisement effectiveness. ✓ indicates that the corresponding dimension is evaluated, but less comprehensively than in AdSpark-Bench.
Figure 2: Overview of the AdSpark-300K construction pipeline, comprising real-world advertisement curation and synthetic advertisement construction, each involving tailored stages for comprehensive and unified annotation.
Figure 3: Statistics of AdSpark-300K, including (a) video duration distribution, (b) audio density distribution, (c) shot count distribution, and (d) video frame quality distribution via MUSIQ Ke et al. (2021) , with MUSIQ scores normalized to [0, 1]. Additional statistics are provided in Fig. 5 .
Visual Quality
Product Fidelity
Instruction Adherence
Temporal Coherence
Audio Alignment
Ad Effectiveness
Method
Aes.
Img.
Subj.
Reg.
Text
X-shot
Struct.
Exec.
Content
Gme
Motion
Trans.
Temp.
BG Audio
Script
Attr.
Creat.
Narr.
ViduQ2
34.56
69.63
70.16
66.49
25.69
2.85
35.77
63.75
77.94
48.99
57.44
2.73
2.86
–
–
62.40
55.80
3.02
ViduQ3
38.26
70.68
60.57
58.72
33.60
84.76
71.26
60.57
83.75
48.22
48.93
78.19
41.62
50.85
76.91
62.79
55.39
81.05
Seedance 2.0
38.10
70.01
63.04
59.79
34.31
94.99
77.14
61.82
85.20
49.45
52.03
90.05
84.19
49.48
80.23
64.11
58.57
91.46
HappyHorse-1.1
39.56
69.60
59.98
57.26
31.04
93.27
89.41
64.22
80.43
48.69
54.49
86.72
95.43
52.40
83.55
64.94
59.00
88.52
Pixverse V5
42.08
74.49
61.78
60.12
18.52
2.75
35.48
67.21
76.94
48.37
57.93
2.11
2.29
–
–
61.12
53.05
2.53
Table 3: Quantitative comparison of state-of-the-art video generation models on AdSpark-Bench, covering proprietary, open-source, and AdSpark-finetuned models. The best results are highlighted in bold , while the second-best results are underlined . Submetric abbreviations correspond to the metrics introduced in the benchmark section, in the same order. Dashes (–) in the audio alignment columns indicate unavailable or invalid audio outputs and are excluded from evaluation.
Figure 4: Qualitative comparison on AdSpark-Bench. We present video frames generated by different models. All methods are conditioned on the corresponding reference image and advertisement prompt. Prompts are abbreviated for readability.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Input
Removed
Retained
Quality Filtering
600,000
199,870
400,130
Suitability Filtering
400,130
49,290
350,840
Category Filtering
350,840
2,314
348,526
Appendix
Table 4: Statistics of the three-stage SKU filtering pipeline. The number of removed SKUs is computed relative to the preceding stage.
Figure 5: Additional statistics of AdSpark-300K. The figure summarizes the distributions of (a) shot-level camera motions, (b) shot types, (c) the number of selling points per product, and (d) advertisement prompt lengths, illustrating the diversity of camera design, product presentation, commercial content, and prompt complexity in the dataset.
Figure 6: Representative reference images from AdSpark-Bench. We show benchmark reference products across diverse categories, including home appliances, electronics, clothing, furniture, food, personal care, and industrial supplies.
Figure 7: Qualitative examples from AdSpark-300K. Each row presents a reference product image and paired advertisement video frames sampled in temporal order.
Figure 8: Examples of advertisement videos and their structured prompts in AdSpark-300K. For each product, the prompt describes the overall scene style, shot-level visual planning, camera type and motion, temporal segments, and synchronized audio cues, including sound effects and voiceover scripts.
Figure 9: Qualitative comparison of different models for a clothes-drying rack product.
Figure 10: Qualitative comparison of different models for a wooden desk product.
Figure 11: Qualitative comparison of different models for a jewelry product.
Metric
Spearman ρ↑
Shot Execution Alignment
0.861
Content Alignment
0.833
Transition Naturalness
0.797
Advertisement Attractiveness
0.719
Creative Quality
0.752
Narrative Coherence
0.774
Appendix
Table 5: Correlation between GPT-assisted metrics and human rankings.
Training Data
Real
Synthetic
Img.
Subj.
Struct.
Trans.
Script
Creat.
LTX-2 (Base)
–
–
66.17
50.39
46.29
33.66
74.71
56.05
Real-100K
100K
0
68.03
64.73
70.68
72.76
78.26
57.61
Synthetic-100K
0
100K
68.44
58.17
71.79
80.09
81.38
57.12
Synthetic-200K
0
200K
68.69
61.59
76.23
87.26
82.62
58.47
AdSpark-300K
100K
200K
69.57
65.01
77.53
88.09
84.41
59.29
Appendix
Table 6: Ablation on the composition and scale of AdSpark-300K. All variants are initialized from the same LTX-2 checkpoint and trained with identical optimization settings unless otherwise specified. The best results are highlighted in bold .
Figure 12: Prompt for VLM-based Quality Filtering.
Figure 13: Prompt for real-world advertisement annotation.
Figure 14: Unified structured annotation schema used for both real-world and synthetic samples in AdSpark-300K.
Figure 15: Prompt for synthetic advertisement planning with GPT-5.5. The production prompt is condensed to retain its core product-grounding, selling-point visualization, shot-planning, and audio-alignment instructions.
Figure 16: Controlled vocabulary for shot and camera-motion planning in synthetic advertisement construction.
Figure 17: Prompt for synthetic advertisement plan review with Gemini-3.1-Pro-Preview.
Figure 18: Prompt for synthetic advertisement plan review with Gemini-3.1-Pro-Preview (continued).
Figure 19: An annotation example from AdSpark-300K. The structured annotation specifies product identity, selling points, scene and style designs, shot-level creative plans, aligned audio scripts, and generation prompts.
Figure 20: An annotation example from AdSpark-300K (continued).
Figure 21: An annotation example from AdSpark-300K (continued).
Figure 22: An annotation example from AdSpark-Bench. The structured annotation specifies product identity, selling points, scene and style designs, shot-level creative plans, aligned audio scripts, and generation prompts.
Figure 23: An annotation example from AdSpark-Bench (continued).
Figure 24: An annotation example from AdSpark-Bench (continued).
Figure 25: Prompt for GPT-5.5-based shot execution alignment.
Figure 26: Prompt for GPT-5.5-based scene and style alignment.
Figure 27: Prompt for GPT-5.5-based selling-point realization.
Figure 28: Prompt for GPT-5.5-based transition naturalness evaluation.
Figure 29: Prompt for GPT-5.5-based advertisement attractiveness evaluation.
Figure 30: Prompt for GPT-5.5-based creative quality evaluation.
Figure 31: Prompt for GPT-5.5-based narrative coherence evaluation.