AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from selecting dialog takes and shaping interview footage into a story to cutting commercials from product shots, voiceovers and graphics. Every task provides a brief, source assets, a container and a set of tests. A task is resolved when the output passes every test. The tests check the delivery format, the content and the brief's explicit requirements, and include a quality test calibrated on 2,582 blind judgments by 43 video editors. We evaluate 16 agents that pair frontier models with coding-agent harnesses such as Codex, Claude Code and OpenCode. The best, GPT-6 Astra in Codex with curated editorial guidance, resolves only 15 of the 56 tasks (26.8%), and the average agent resolves 14.0%. Human editors prefer the reference edit in 83.5% of judgments. Most unresolved runs (562 of 771) fail only the quality test: agents perceive footage through stills and transcripts and check their renders for defects, not craft. We release the tasks, verifier and per-run results at https://timelinebench.tensortest.com.
Figures & tables
Model
Harness
Tests
Res. %
Human %
GPT-6 Astra
OpenCode
48
21.4
22.3
Claude Fable 5.1
OpenCode
50
17.9
22.9
Claude Opus 5
OpenCode
46
17.9
19.0
GPT-5.6 Sol
OpenCode
45
17.9
17.6
Grok 4.6
OpenCode
39
12.5
17.0
Gemini 3.8 Flash
OpenCode
42
10.7
13.1
Table 1: Main results , one run per task. Tests : runs, of 56, passing every test but the quality test. Res. : resolution rate. Human : human win-or-tie rate. With CIs in Table 8 .
Ability
Evidence
Reading
Delivery and explicit requirements
857 of 863 videos meet the specification; 762 pass every brief test
Reliable
Sound
Required lines missed: 5.9% (transcripts); 41.5% for computer use
Via tools only
Timeline hygiene
Repeats, outtakes in 1–7% of top agents’ losses; 18.5% for computer use
Mostly reliable
Checking one’s own edit
Output checked in 86–100% of runs; 5.5% of fixes concern pacing, story, shots
Defects only
Judging one’s own edit
93% of runs claim success, including 95.5% of edits that fail a test
Absent
Story assembly
Drives agent wins; narrative scenes are the best collection (28.7%)
Emerging
Table 2: Current agents as editors. Evidence from Section 6 and Appendix J .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Total
Median per task
Delivery
Collection
Tasks
source hours
source minutes
reference seconds
source/ reference
video files
format
fps
window (s)
EditStock
11
18.8
69.5
67.1
57.3
96
L; HD, 4K
23.976, 24
29–305
Cinestudy
15
10.3
27.0
110.9
14.6
2
L; HD
23.976, 24, 25
35–315
UGC
15
1.4
5.7
39.1
8.4
50
P; HD
25
25–55
Commercial
15
2.4
9.6
63.1
8.9
76
L; HD
24
44–125
All
56
33.0
12.4
60.3
12.0
55.5
41 L, 15 P
23.976, 24, 25
25–315
Appendix
Table 3: Composition of Timeline-Bench by collection. Source hours sum each task’s primary moving-image footage. Medians are per task; source minutes and the source/reference ratio exclude Beauty of Delhi , which is built from 603 stills. Video files count every video container in a task’s inputs. L/P: landscape/portrait; HD: 1920 × 1080 or 1080 × 1920; 4K: 3840 × 2160 or 4096 × 2160.
Content test
Measurement
Limit
Mostly silent
share of the runtime that is silent (no audio track also fails)
0.50
Mostly frozen
share of the runtime with a frozen picture
0.60
Repeated footage
duration of footage shown more than once
1 s
Dead air
longest internal silence
5 s
Channel imbalance
level difference between the left and right channels
6 dB
Faces cut by the frame edge
share of frames with a face in which a face is cut by the edge
0.30
Appendix
Table 4: Content tests. Each content test fails a video whose measurement exceeds the limit.
Model
Model identifier
Provider route
Reasoning setting
OpenCode 1.18.31
Gemini 3.1 Pro Preview
google/gemini-3.1-pro-preview
Google API
thinkingLevel: high
Gemini 3.8 Flash
google/gemini-3.8-flash
Google API
thinkingLevel: high
Claude Fable 5.1
anthropic/claude-fable-5-1
Anthropic API
adaptive thinking, effort: max
Claude Opus 5
anthropic/claude-opus-5
Anthropic API
adaptive thinking, effort: max
GPT-5.6 Sol
openai/gpt-5.6-sol
OpenAI API
reasoningEffort: max
Appendix
Table 5: Agent configurations. Model identifiers are the exact strings sent to the provider. OpenRouter routes are restricted to the named provider with fallbacks disabled. Every reasoning setting is the provider’s maximum and also applies to helper and subagent roles. DeepSeek Flash is the deepseek-flash alias, served by DeepSeek-V4.1-Flash.
3.11 virtual environment with Pillow 11.3.0, NumPy 2.2.6, SciPy 1.15.3, soundfile 0.13.1, OpenCV (headless) 4.12.0.88 and pypdf 6.0.0
FFmpeg and ffprobe
5.1.9 (Debian 12 package)
Remotion
4.0.524 with its CLI and media packages; React 19.3.0
HyperFrames
0.8.40 with GSAP 3.15.0
Appendix
Table 6: Linux execution environment of the 15 coding agents.
Median runtime
Mean cost
Model
Harness / guidance
per run (min)
per run ($)
GPT-6 Astra
OpenCode
26.2
9.67
Claude Fable 5.1
OpenCode
60.6
16.96
Claude Opus 5
OpenCode
51.4
13.83
GPT-5.6 Sol
OpenCode
27.1
5.96
Grok 4.6
OpenCode
25.9
3.46
Appendix
Table 7: Runtime and model cost per agent, over all 56 runs of each agent. An asterisk marks token-based cost estimates for Codex CLI.
Model
Harness / guidance
Resolution rate % [95% CI]
Human win-or-tie % [95% CI]
GPT-6 Astra
OpenCode
21.4 [11.6, 34.4]
22.3 [14.8, 29.8]
Claude Fable 5.1
OpenCode
17.9 [8.9, 30.4]
22.9 [15.0, 30.9]
Claude Opus 5
OpenCode
17.9 [8.9, 30.4]
19.0 [11.6, 26.4]
GPT-5.6 Sol
OpenCode
17.9 [8.9, 30.4]
17.6 [9.4, 25.8]
Grok 4.6
OpenCode
12.5 [5.2, 24.1]
17.0 [10.1, 23.9]
Gemini 3.8 Flash
OpenCode
10.7 [4.0, 21.9]
13.1 [7.3, 18.9]
Appendix
Table 8: Main results with 95% confidence intervals , one run per task ( Table 1 ).
Contrast
n
Δ W/T [95% CI]
p
Resolved ( p )
GPT-5.6 Sol: Codex CLI − OpenCode
55
− 4.2 [ − 13.2, + 4.7]
1.00
5 vs 10 (0.63)
GPT-6 Astra: Codex CLI − OpenCode
53
+ 0.9 [ − 7.5, + 9.4]
1.00
12 vs 12 (1.00)
Claude Fable 5.1: Claude Code − OpenCode
56
+ 0.6 [ − 9.6, + 10.8]
1.00
10 vs 10 (1.00)
Claude Opus 5: Claude Code − OpenCode
51
+ 5.9 [ − 2.9, + 14.7]
1.00
13 vs 10 (1.00)
GPT-6 Astra: curated guidance − none
56
+ 0.6 [ − 8.4, + 9.6]
1.00
15 vs 12 (1.00)
GPT-6 Astra: computer use − code
53
− 16.4 [ − 24.0, − 8.7]
< 0.001
2 vs 12 (0.038)
Appendix
Table 9: Matched contrasts , first minus second. n : tasks with both edits judged. Δ W/T: difference in human win-or-tie rate over these tasks, in points, with 95% CI and Holm-adjusted paired-permutation p . Resolved : tasks resolved by each agent (Holm-adjusted exact McNemar p ).
Category
AP (%)
RS (%)
Either (%)
Overall polish
13.8
46.9
52.8 [42.9, 63.0]
Shot selection
10.0
29.3
34.2 [24.7, 42.9]
Text and graphics
8.9
29.4
31.8 [24.8, 38.9]
Transitions and effects
6.7
26.6
29.2 [18.4, 40.3]
Story and structure
6.9
25.1
27.6 [19.8, 36.0]
Sound design
5.1
20.9
23.1 [15.7, 30.6]
Appendix
Table 10: Why human editors prefer the reference edit , by category: share of the 1,101 notes on judgments preferring the reference edit that name a problem of the agent edit (AP), a strength of the reference edit (RS), or either (95% interval from resampling human editors).
Benchmark
Input
Output scored
Evaluation
Human judgment of outputs
VEBench ( Deng et al., 2026a )
Edited videos and a question
Answer, clip choice or time span
Accuracy; temporal overlap
None
MEDit-Bench ( Ogata et al., 2026 )
One long video and an editing message
Cut list
Temporal overlap with professional edits
User study on a subset (1,620 evaluations)
AgenticVBench ( Cao et al., 2026 )
Source videos and a brief or storyboard
Video with a manifest or report
Programmatic verifiers; 1,069 binary rubric items for repurposing
Expert rubric grading; three-editor human baseline on a subset
CutVerse ( Hu et al., 2026 )
GUI application state and an objective
GUI trajectory
Milestone checks
Not reported
ProSoftArena ( Ai et al., 2026 )
Real desktop and a task
Final state or artifact
Execution scripts; subjective comparison with human work on creative tasks
Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-specific tasks. They face two critical limitations: i) inability to handle diverse video comprehension and editing operations, and ii) lack of long-video understanding for coherent narrative creation. We propose VideoAgent, an all-in-one agentic framework addressing these challenges through two key innovations. First, we develop automated video shot creation with shot planning agents for coherent narratives and cross-modal retrieval for aligned visual content. Second, we design a multi-agent orchestration framework integrating over thirty specialized editing agents. Intent parsing filters relevant tools while textual-gradient graph optimization assembles complex editing pipelines. Extensive experiments on our newly-proposed VideoEdit benchmark and public datasets demonstrate VideoAgent's superiority over existing multimodal LLMs and agentic systems. VideoAgent achieves 87-95% orchestration success rates while reducing API costs by 60%. Human evaluation across six video categories shows VideoAgent produces professional-quality content approaching human-level performance, with ratings only 4% below human-created videos. We release our code at https://github.com/HKUDS/VideoAgent.
Hengji Zhou, Lingxuan Huang, Jian Wang +4
1Harbin Institute of Technology, Shenzhen · 2South China University of Technology · 3The University of Hong Kong +1
Editing a long-form video from heterogeneous footage requires more than selecting clips: an agent must preserve narrative intent across material preparation, timeline construction, post-production, and revision while leaving enough evidence to diagnose failures. We present \textbf{Crayotter}, an open-source multimodal multi-agent system for prompt-driven video editing. Crayotter organizes production into three phases: coverage-aware material preparation, artifact-based editing research, and tool-grounded timeline execution. Each phase externalizes inspectable artifacts, including coverage reports, multimodal analyses, editing blueprints, tool calls, and intermediate renders. These artifacts make an editing run traceable and allow failed segments to be diagnosed and selectively revised instead of requiring a full restart. We evaluate Crayotter on 23 editing themes against CapCut-Mate and CutClaw. Under human evaluation, Crayotter achieves an average score of 3.40/5, compared with 2.44 and 1.70 for the two baselines, with consistent gains in theme alignment, narrative coherence, and editing smoothness. We additionally describe a replayable trajectory schema and verifiable reward design that prepare these workflows for future policy optimization. Code, traces, and examples are publicly available at https://github.com/idwts/Crayotter.
Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative the editor wishes to convey. Benchmarks for a closely related task, video summarization, reduce editorial intent to a single, message-agnostic notion of saliency and thus do not account for this diversity. For evaluating message-driven video editing, we present \textbf{MEDit-Bench}, a dataset and benchmark, which pairs long-form videos with multiple editing messages and multiple professionally produced edits per message, demonstrating that different messages yield substantially different edits from the same source. We define an automatic evaluation protocol based on temporal alignment metrics, and find that an LLM-as-a-judge preference, a natural proxy for narrative quality, is unreliable for this task due to severe position bias. We additionally annotate each message with ambiguity and contextfulness scores, and show that both dimensions negatively correlate with model performance, establishing message difficulty as a meaningful stratification factor. Experiments with state-of-the-art MLLMs and reinforcement fine-tuned baselines show that while strong models approach human temporal alignment at lenient thresholds, all models fall behind humans at stricter criteria. A human perceptual study further confirms a large quality gap, with professional human edits remaining consistently preferred over model outputs.