cs.CVOct 8, 2026

From Video Clips to Creation Trajectory: Sora100K for AI-Native Video Creation

Authors: Sicong Yang, Ruihuan Yang, Jian Lu, Jianfei Yuan, Xiaodong Cun, Xiuli Bi

Organizations: Chongqing University of Posts and Telecommunications, Chongqing, China · GVC Lab, Great Bay University, Dongguan, China

Abstract

AI-Native video creation is shifting from isolated video clips toward iterative video creation workflows. However, existing datasets remain largely video clips, representing video generation and editing as separate tasks rather than connected stages of a video creation workflow. In this paper, we introduce Sora100K, a dataset that represents the AI-Native video creation workflow as a structured video creation trajectory. Specifically, we first identify video creation trajectories and decompose them into three subsets according to their structural roles: text-to-video generation records as roots, single-turn video editing records as editing edges, and multi-turn video editing records as complete trajectories. Then, we use a VLM to assign semantic annotations for generation roots and editing-operation annotations for editing edges. A strict construction pipeline further reconstructs source-to-edit lineage, editing order, and intermediate video states while ensuring data quality. Finally, we perform lightweight adaptation on LTX-2 models to assess the supervision value of Sora100K. The results show improvements in visual quality, multi-shot generation, and cross-shot consistency, while successive-turn evaluation reveals that following multi-turn editing instructions remains challenging. Sora100K establishes a new data foundation for AI-Native video creation beyond isolated video clips and toward structured video creation trajectory. The dataset and supplementary materials are publicly available at https://huggingface.co/datasets/ysicong/Sora100K.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VideoX-Qwen: Data-Centric Instruction-Based Video Editing

    Sep 22, 2026JJiahang Li, Dingbao Shao, Xinyu Chen +8Video EditingInstruction-Based Video Editing

  2. JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation

    Jun 2, 2026Yinan Chen, Chuming Lin, Zhennan Chen +12Video Editing BenchmarksTraining Data Curation

  3. Bridging Creative Intent and Visual Quality: Creator-Driven Recurrent Video Generation with Agentic Feedback Loops

    Jun 17, 2026Denis Savytski, Aiden Lei, Heding Liu +4Long Video GenerationHuman-in-the-Loop AI