Product-centric advertisement video generation aims to create promotional videos that preserve fine-grained product identity while presenting selling points through coherent multi-shot narratives. However, this emerging task remains underexplored due to the lack of large-scale advertisement-specific datasets and comprehensive evaluation frameworks. To address this gap, we introduce \textbf{AdSpark}, a large-scale dataset and benchmark for product-centric advertisement video generation, based on data from a major e-commerce platform. \textit{AdSpark-300K} contains approximately 300K reference image--prompt--video triplets, comprising a real-world subset and a synthetic subset. Each sample provides structured advertisement annotations, including product identity annotations, selling-point descriptions, creative plans, and aligned audio scripts, enabling models to learn product preservation and advertisement-oriented visual storytelling. We further propose \textit{AdSpark-Bench}, a diagnostic benchmark that evaluates generated advertisements across six dimensions, including visual quality, product fidelity, instruction adherence, temporal coherence, audio alignment, and advertisement effectiveness. Based on AdSpark-Bench, we evaluate representative models, revealing key challenges in product preservation, multi-shot storytelling, and selling-point visualization. Experiments with AdSpark-300K-finetuned models further validate the effectiveness of our dataset. AdSpark provides a unified dataset and benchmark for future research, and we will release the dataset upon acceptance.
Figures & tables
Figure 1: Overview of AdSpark-300K and AdSpark-Bench. AdSpark-300K is a large-scale, high-quality product-centric advertisement video dataset spanning real-world and synthetic advertisements, with structured annotations across multiple dimensions. AdSpark-Bench provides a comprehensive evaluation of general video quality and advertisement-specific capabilities.
Dataset
Domain
Task
Multi- shot
Creative Plans
Selling Points
Identity Annotation
Audio Script
Clips
Resolution
Average Length(s)
MSRVTT Xu et al. (2016)
Open
T2V
✗
✗
✗
✗
✗
10K
240P
14.4
WebVid-10M Bain et al. (2021)
Open
T2V
✗
✗
✗
✗
✗
10M
360P
18.7
HD-VG-130M Wang et al. (2023a)
Open
T2V
✗
✗
✗
✗
✗
130M
720P
4.9
Panda-70M Chen et al. (2024)
Open
T2V
✗
✗
✗
✗
✗
70M
720P
8.6
InternVid Wang et al. (2023b)
Open
T2V
✗
✗
✗
✗
✗
234M
720P
11.7
OpenHumanVid Li et al. (2024)
Human
T2V
✗
✗
✗
✗
✗
52.3M
720P
4.9
Table 1: Comparison with existing video generation datasets. Most existing datasets focus on generic text-to-video generation (T2V), subject-to-video generation (S2V), or image-to-video generation (I2V), while lacking advertisement-specific annotations. In contrast, AdSpark-300K targets PC-AVG, providing structured advertisement annotations, including product identity, selling points, creative plans (style, scene, and shot design), and audio scripts.
Benchmark
Visual Quality
Script Adherence
Temporal Coherence
Product Fidelity
Shot Compliance
Audio Alignment
Selling-point Realization
Advertisement Effectiveness
Make-a-Video-Eval Singer et al. (2022)
✓
✓
✗
✗
✗
✗
✗
✗
FETV Liu et al. (2024b)
✓
✓
✓
✗
✗
✗
✗
✗
T2VScore Wu et al. (2024)
✓
✓
✓
✗
✗
✗
✗
✗
EvalCrafter Liu et al. (2024a)
✓
✓
✓
✗
✗
✗
✗
✗
VBench Huang et al. (2023)
✓
✓
✓
✗
✗
✗
✗
✗
VBench++ Huang et al. (2024)
✓
✓
✓
✗
✗
✗
✗
✗
Table 2: Comparison of AdSpark-Bench with existing video generation benchmarks. Existing benchmarks focus on general video quality, while AdSpark-Bench further evaluates product fidelity, selling-point realization, and advertisement effectiveness. ✓ indicates that the corresponding dimension is evaluated, but less comprehensively than in AdSpark-Bench.
Figure 2: Overview of the AdSpark-300K construction pipeline, comprising real-world advertisement curation and synthetic advertisement construction, each involving tailored stages for comprehensive and unified annotation.
Figure 3: Statistics of AdSpark-300K, including (a) video duration distribution, (b) audio density distribution, (c) shot count distribution, and (d) video frame quality distribution via MUSIQ Ke et al. (2021) , with MUSIQ scores normalized to [0, 1]. Additional statistics are provided in Fig. 5 .
Visual Quality
Product Fidelity
Instruction Adherence
Temporal Coherence
Audio Alignment
Ad Effectiveness
Method
Aes.
Img.
Subj.
Reg.
Text
X-shot
Struct.
Exec.
Content
Gme
Motion
Trans.
Temp.
BG Audio
Script
Attr.
Creat.
Narr.
ViduQ2
34.56
69.63
70.16
66.49
25.69
2.85
35.77
63.75
77.94
48.99
57.44
2.73
2.86
–
–
62.40
55.80
3.02
ViduQ3
38.26
70.68
60.57
58.72
33.60
84.76
71.26
60.57
83.75
48.22
48.93
78.19
41.62
50.85
76.91
62.79
55.39
81.05
Seedance 2.0
38.10
70.01
63.04
59.79
34.31
94.99
77.14
61.82
85.20
49.45
52.03
90.05
84.19
49.48
80.23
64.11
58.57
91.46
HappyHorse-1.1
39.56
69.60
59.98
57.26
31.04
93.27
89.41
64.22
80.43
48.69
54.49
86.72
95.43
52.40
83.55
64.94
59.00
88.52
Pixverse V5
42.08
74.49
61.78
60.12
18.52
2.75
35.48
67.21
76.94
48.37
57.93
2.11
2.29
–
–
61.12
53.05
2.53
Table 3: Quantitative comparison of state-of-the-art video generation models on AdSpark-Bench, covering proprietary, open-source, and AdSpark-finetuned models. The best results are highlighted in bold , while the second-best results are underlined . Submetric abbreviations correspond to the metrics introduced in the benchmark section, in the same order. Dashes (–) in the audio alignment columns indicate unavailable or invalid audio outputs and are excluded from evaluation.
Figure 4: Qualitative comparison on AdSpark-Bench. We present video frames generated by different models. All methods are conditioned on the corresponding reference image and advertisement prompt. Prompts are abbreviated for readability.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Input
Removed
Retained
Quality Filtering
600,000
199,870
400,130
Suitability Filtering
400,130
49,290
350,840
Category Filtering
350,840
2,314
348,526
Appendix
Table 4: Statistics of the three-stage SKU filtering pipeline. The number of removed SKUs is computed relative to the preceding stage.
Figure 5: Additional statistics of AdSpark-300K. The figure summarizes the distributions of (a) shot-level camera motions, (b) shot types, (c) the number of selling points per product, and (d) advertisement prompt lengths, illustrating the diversity of camera design, product presentation, commercial content, and prompt complexity in the dataset.
Figure 6: Representative reference images from AdSpark-Bench. We show benchmark reference products across diverse categories, including home appliances, electronics, clothing, furniture, food, personal care, and industrial supplies.
Figure 7: Qualitative examples from AdSpark-300K. Each row presents a reference product image and paired advertisement video frames sampled in temporal order.
Figure 8: Examples of advertisement videos and their structured prompts in AdSpark-300K. For each product, the prompt describes the overall scene style, shot-level visual planning, camera type and motion, temporal segments, and synchronized audio cues, including sound effects and voiceover scripts.
Figure 9: Qualitative comparison of different models for a clothes-drying rack product.
Figure 10: Qualitative comparison of different models for a wooden desk product.
Figure 11: Qualitative comparison of different models for a jewelry product.
Metric
Spearman ρ↑
Shot Execution Alignment
0.861
Content Alignment
0.833
Transition Naturalness
0.797
Advertisement Attractiveness
0.719
Creative Quality
0.752
Narrative Coherence
0.774
Appendix
Table 5: Correlation between GPT-assisted metrics and human rankings.
Training Data
Real
Synthetic
Img.
Subj.
Struct.
Trans.
Script
Creat.
LTX-2 (Base)
–
–
66.17
50.39
46.29
33.66
74.71
56.05
Real-100K
100K
0
68.03
64.73
70.68
72.76
78.26
57.61
Synthetic-100K
0
100K
68.44
58.17
71.79
80.09
81.38
57.12
Synthetic-200K
0
200K
68.69
61.59
76.23
87.26
82.62
58.47
AdSpark-300K
100K
200K
69.57
65.01
77.53
88.09
84.41
59.29
Appendix
Table 6: Ablation on the composition and scale of AdSpark-300K. All variants are initialized from the same LTX-2 checkpoint and trained with identical optimization settings unless otherwise specified. The best results are highlighted in bold .
Figure 12: Prompt for VLM-based Quality Filtering.
Figure 13: Prompt for real-world advertisement annotation.
Figure 14: Unified structured annotation schema used for both real-world and synthetic samples in AdSpark-300K.
Figure 15: Prompt for synthetic advertisement planning with GPT-5.5. The production prompt is condensed to retain its core product-grounding, selling-point visualization, shot-planning, and audio-alignment instructions.
Figure 16: Controlled vocabulary for shot and camera-motion planning in synthetic advertisement construction.
Figure 17: Prompt for synthetic advertisement plan review with Gemini-3.1-Pro-Preview.
Figure 18: Prompt for synthetic advertisement plan review with Gemini-3.1-Pro-Preview (continued).
Figure 19: An annotation example from AdSpark-300K. The structured annotation specifies product identity, selling points, scene and style designs, shot-level creative plans, aligned audio scripts, and generation prompts.
Figure 20: An annotation example from AdSpark-300K (continued).
Figure 21: An annotation example from AdSpark-300K (continued).
Figure 22: An annotation example from AdSpark-Bench. The structured annotation specifies product identity, selling points, scene and style designs, shot-level creative plans, aligned audio scripts, and generation prompts.
Figure 23: An annotation example from AdSpark-Bench (continued).
Figure 24: An annotation example from AdSpark-Bench (continued).
Figure 25: Prompt for GPT-5.5-based shot execution alignment.
Figure 26: Prompt for GPT-5.5-based scene and style alignment.
Figure 27: Prompt for GPT-5.5-based selling-point realization.
Figure 28: Prompt for GPT-5.5-based transition naturalness evaluation.
Figure 29: Prompt for GPT-5.5-based advertisement attractiveness evaluation.
Figure 30: Prompt for GPT-5.5-based creative quality evaluation.
Figure 31: Prompt for GPT-5.5-based narrative coherence evaluation.
Generating realistic and user-preferred advertisements is a key challenge in e-commerce. Existing approaches utilize multiple independent models driven by click-through-rate (CTR) to controllably create attractive image or text advertisements. However, their pipelines lack cross-modal perception and rely on CTR that only reflects average preferences. Therefore, we explore jointly generating personalized image-text advertisements from historical click behaviors. We first design a Unified Advertisement Generative model (Uni-AdGen) that employs a single autoregressive framework to produce both advertising images and texts. By incorporating a foreground perception module and instruction tuning, Uni-AdGen enhances the realism of the generated content. To further personalize advertisements, we equip Uni-AdGen with a coarse-to-fine preference understanding module that effectively captures user interests from noisy multimodal historical behaviors to drive personalized generation. Additionally, we construct the first large-scale Personalized Advertising image-text dataset (PAd1M) and introduce a Product Background Similarity (PBS) metric to facilitate training and evaluation. Extensive experiments show that our method outperforms baselines in general and personalized advertisement generation. Our project is available at https://github.com/JD-GenX/Uni-AdGen.
Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realistic clips from multimodal conditions, yet they do not optimize how a product should be transformed into an effective advertisement or how future generation should be improved from online business feedback. To close this loop, we propose AgenticGen, a reward-guided agentic framework that decomposes advertising video generation into two trainable reasoning stages, strategy selection and draft generation, thereby exposing optimization targets that online business feedback can supervise. AgenticGen learns a performance-based reward from accumulated online feedback and a complementary rubric-based reward aligned with human quality standards, then uses them to supervise policy optimization. DPO first moves the agentic policies toward online preferences, and GRPO further refines both stages with process and outcome rewards. Offline experiments validate the reward models and successive policy optimization. Online A/B experiments in the TikTok advertising system show that AgenticGen after DPO and GRPO improves CTR by 2.72%, CVR by 2.63%, and Advv by 9.61% over the SFT baseline.
Xingyuan Bu, Chengru Song, Hao Zhou +9
ByteDance · Tsinghua University · The Chinese University of Hong Kong
Creative Generation (CG) leverages generative models to automatically produce advertising content that highlights product features, and it has been a significant focus of recent research. However, while CG has advanced considerably, most efforts have concentrated on generating advertising text and images, leaving Creative Video Generation (CVG) relatively underexplored. This gap is largely due to two major challenges faced by Text-to-Video (T2V) models: (a) \textbf{ambiguous semantic alignment}, where models struggle to accurately correlate product selling points with creative video content, and (b) \textbf{inadequate motion adaptability}, resulting in unrealistic movements and distortions. To address these challenges, we develop a comprehensive Advertising Creative Knowledge Base (ACKB) as a foundational resource and propose a knowledge-driven approach (KD-CVG) to overcome the knowledge limitations of existing models. KD-CVG consists of two primary modules: Semantic-Aware Retrieval (SAR) and Multimodal Knowledge Reference (MKR). SAR utilizes the semantic awareness of graph attention networks and reinforcement learning feedback to enhance the model's comprehension of the connections between selling points and creative videos. Building on this, MKR incorporates semantic and motion priors into the T2V model to address existing knowledge gaps. Extensive experiments have demonstrated KD-CVG's superior performance in achieving semantic alignment and motion adaptability, validating its effectiveness over other state-of-the-art methods. The code and dataset will be open source at https://kdcvg.github.io/KDCVG/.