cs.AI · 2609.12769 Copy arXiv ID · Sep 14, 2026 Save Unified Agentic Video Editing Across Levels of Complexity and Creativity Authors: Surabhi S. Nath , Kim Ferres , Milan Petrović , Lion Schulz
Abstract Editing is a core component of video production, requiring creative planning and decisions under multiple constraints. Here, we report methods for agentic tooling for automated video editing across three tasks varying in editorial goal, complexity and creativity, namely scene previews, video summaries and cinematic trailers. We evaluate the outputs and discuss implications for automation and agency.
Explore similar work Jun 22, 2026 · Hengji Zhou, Lingxuan Huang, Jian Wang +4 Agentic Video Generation Video Editing
May 31, 2026 · Lecheng Yan, Yichong Zhang, Ben Pan +5 Video Editing Agentic Video Generation
May 18, 2026 · Yongsheng Yu, Ziyun Zeng, Zhiyuan Xiao +4 Video Editing Text-To-Video Models
Jun 22, 2026 · cs.CV J/K move · Enter open · S save
Hengji Zhou, Lingxuan Huang, Jian Wang, Bing Zhou +3
1Harbin Institute of Technology, Shenzhen · 2South China University of Technology · 3The University of Hong Kong · 4Snap Inc
Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-specific tasks. They face two critical limitations: i) inability to handle diverse video comprehension and editing operations, and ii) lack of long-video understanding for coherent narrative creation. We propose VideoAgent, an all-in-one agentic framework addressing these challenges through two key innovations. First, we develop automated video shot creation with shot planning agents for coherent narratives and cross-modal retrieval for aligned visual content. Second, we design a multi-agent orchestration framework integrating over thirty specialized editing agents. Intent parsing filters relevant tools while textual-gradient graph optimization assembles complex editing pipelines. Extensive experiments on our newly-proposed VideoEdit benchmark and public datasets demonstrate VideoAgent's superiority over existing multimodal LLMs and agentic systems. VideoAgent achieves 87-95% orchestration success rates while reducing API costs by 60%. Human evaluation across six video categories shows VideoAgent produces professional-quality content approaching human-level performance, with ratings only 4% below human-created videos. We release our code at https://github.com/HKUDS/VideoAgent.