AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from selecting dialog takes and shaping interview footage into a story to cutting commercials from product shots, voiceovers and graphics. Every task provides a brief, source assets, a container and a set of tests. A task is resolved when the output passes every test. The tests check the delivery format, the content and the brief's explicit requirements, and include a quality test calibrated on 2,582 blind judgments by 43 video editors. We evaluate 16 agents that pair frontier models with coding-agent harnesses such as Codex, Claude Code and OpenCode. The best, GPT-6 Astra in Codex with curated editorial guidance, resolves only 15 of the 56 tasks (26.8%), and the average agent resolves 14.0%. Human editors prefer the reference edit in 83.5% of judgments. Most unresolved runs (562 of 771) fail only the quality test: agents perceive footage through stills and transcripts and check their renders for defects, not craft. We release the tasks, verifier and per-run results at https://timelinebench.tensortest.com.
Figures & tables
Model
Harness
Tests
Res. %
Human %
GPT-6 Astra
OpenCode
48
21.4
22.3
Claude Fable 5.1
OpenCode
50
17.9
22.9
Claude Opus 5
OpenCode
46
17.9
19.0
GPT-5.6 Sol
OpenCode
45
17.9
17.6
Grok 4.6
OpenCode
39
12.5
17.0
Gemini 3.8 Flash
OpenCode
42
10.7
13.1
Table 1: Main results , one run per task. Tests : runs, of 56, passing every test but the quality test. Res. : resolution rate. Human : human win-or-tie rate. With CIs in Table 8 .
Ability
Evidence
Reading
Delivery and explicit requirements
857 of 863 videos meet the specification; 762 pass every brief test
Reliable
Sound
Required lines missed: 5.9% (transcripts); 41.5% for computer use
Via tools only
Timeline hygiene
Repeats, outtakes in 1–7% of top agents’ losses; 18.5% for computer use
Mostly reliable
Checking one’s own edit
Output checked in 86–100% of runs; 5.5% of fixes concern pacing, story, shots
Defects only
Judging one’s own edit
93% of runs claim success, including 95.5% of edits that fail a test
Absent
Story assembly
Drives agent wins; narrative scenes are the best collection (28.7%)
Emerging
Table 2: Current agents as editors. Evidence from Section 6 and Appendix J .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Total
Median per task
Delivery
Collection
Tasks
source hours
source minutes
reference seconds
source/ reference
video files
format
fps
window (s)
EditStock
11
18.8
69.5
67.1
57.3
96
L; HD, 4K
23.976, 24
29–305
Cinestudy
15
10.3
27.0
110.9
14.6
2
L; HD
23.976, 24, 25
35–315
UGC
15
1.4
5.7
39.1
8.4
50
P; HD
25
25–55
Commercial
15
2.4
9.6
63.1
8.9
76
L; HD
24
44–125
All
56
33.0
12.4
60.3
12.0
55.5
41 L, 15 P
23.976, 24, 25
25–315
Appendix
Table 3: Composition of Timeline-Bench by collection. Source hours sum each task’s primary moving-image footage. Medians are per task; source minutes and the source/reference ratio exclude Beauty of Delhi , which is built from 603 stills. Video files count every video container in a task’s inputs. L/P: landscape/portrait; HD: 1920 × 1080 or 1080 × 1920; 4K: 3840 × 2160 or 4096 × 2160.
Content test
Measurement
Limit
Mostly silent
share of the runtime that is silent (no audio track also fails)
0.50
Mostly frozen
share of the runtime with a frozen picture
0.60
Repeated footage
duration of footage shown more than once
1 s
Dead air
longest internal silence
5 s
Channel imbalance
level difference between the left and right channels
6 dB
Faces cut by the frame edge
share of frames with a face in which a face is cut by the edge
0.30
Appendix
Table 4: Content tests. Each content test fails a video whose measurement exceeds the limit.
Model
Model identifier
Provider route
Reasoning setting
OpenCode 1.18.31
Gemini 3.1 Pro Preview
google/gemini-3.1-pro-preview
Google API
thinkingLevel: high
Gemini 3.8 Flash
google/gemini-3.8-flash
Google API
thinkingLevel: high
Claude Fable 5.1
anthropic/claude-fable-5-1
Anthropic API
adaptive thinking, effort: max
Claude Opus 5
anthropic/claude-opus-5
Anthropic API
adaptive thinking, effort: max
GPT-5.6 Sol
openai/gpt-5.6-sol
OpenAI API
reasoningEffort: max
Appendix
Table 5: Agent configurations. Model identifiers are the exact strings sent to the provider. OpenRouter routes are restricted to the named provider with fallbacks disabled. Every reasoning setting is the provider’s maximum and also applies to helper and subagent roles. DeepSeek Flash is the deepseek-flash alias, served by DeepSeek-V4.1-Flash.
3.11 virtual environment with Pillow 11.3.0, NumPy 2.2.6, SciPy 1.15.3, soundfile 0.13.1, OpenCV (headless) 4.12.0.88 and pypdf 6.0.0
FFmpeg and ffprobe
5.1.9 (Debian 12 package)
Remotion
4.0.524 with its CLI and media packages; React 19.3.0
HyperFrames
0.8.40 with GSAP 3.15.0
Appendix
Table 6: Linux execution environment of the 15 coding agents.
Median runtime
Mean cost
Model
Harness / guidance
per run (min)
per run ($)
GPT-6 Astra
OpenCode
26.2
9.67
Claude Fable 5.1
OpenCode
60.6
16.96
Claude Opus 5
OpenCode
51.4
13.83
GPT-5.6 Sol
OpenCode
27.1
5.96
Grok 4.6
OpenCode
25.9
3.46
Appendix
Table 7: Runtime and model cost per agent, over all 56 runs of each agent. An asterisk marks token-based cost estimates for Codex CLI.
Model
Harness / guidance
Resolution rate % [95% CI]
Human win-or-tie % [95% CI]
GPT-6 Astra
OpenCode
21.4 [11.6, 34.4]
22.3 [14.8, 29.8]
Claude Fable 5.1
OpenCode
17.9 [8.9, 30.4]
22.9 [15.0, 30.9]
Claude Opus 5
OpenCode
17.9 [8.9, 30.4]
19.0 [11.6, 26.4]
GPT-5.6 Sol
OpenCode
17.9 [8.9, 30.4]
17.6 [9.4, 25.8]
Grok 4.6
OpenCode
12.5 [5.2, 24.1]
17.0 [10.1, 23.9]
Gemini 3.8 Flash
OpenCode
10.7 [4.0, 21.9]
13.1 [7.3, 18.9]
Appendix
Table 8: Main results with 95% confidence intervals , one run per task ( Table 1 ).
Contrast
n
Δ W/T [95% CI]
p
Resolved ( p )
GPT-5.6 Sol: Codex CLI − OpenCode
55
− 4.2 [ − 13.2, + 4.7]
1.00
5 vs 10 (0.63)
GPT-6 Astra: Codex CLI − OpenCode
53
+ 0.9 [ − 7.5, + 9.4]
1.00
12 vs 12 (1.00)
Claude Fable 5.1: Claude Code − OpenCode
56
+ 0.6 [ − 9.6, + 10.8]
1.00
10 vs 10 (1.00)
Claude Opus 5: Claude Code − OpenCode
51
+ 5.9 [ − 2.9, + 14.7]
1.00
13 vs 10 (1.00)
GPT-6 Astra: curated guidance − none
56
+ 0.6 [ − 8.4, + 9.6]
1.00
15 vs 12 (1.00)
GPT-6 Astra: computer use − code
53
− 16.4 [ − 24.0, − 8.7]
< 0.001
2 vs 12 (0.038)
Appendix
Table 9: Matched contrasts , first minus second. n : tasks with both edits judged. Δ W/T: difference in human win-or-tie rate over these tasks, in points, with 95% CI and Holm-adjusted paired-permutation p . Resolved : tasks resolved by each agent (Holm-adjusted exact McNemar p ).
Category
AP (%)
RS (%)
Either (%)
Overall polish
13.8
46.9
52.8 [42.9, 63.0]
Shot selection
10.0
29.3
34.2 [24.7, 42.9]
Text and graphics
8.9
29.4
31.8 [24.8, 38.9]
Transitions and effects
6.7
26.6
29.2 [18.4, 40.3]
Story and structure
6.9
25.1
27.6 [19.8, 36.0]
Sound design
5.1
20.9
23.1 [15.7, 30.6]
Appendix
Table 10: Why human editors prefer the reference edit , by category: share of the 1,101 notes on judgments preferring the reference edit that name a problem of the agent edit (AP), a strength of the reference edit (RS), or either (95% interval from resampling human editors).
Benchmark
Input
Output scored
Evaluation
Human judgment of outputs
VEBench ( Deng et al., 2026a )
Edited videos and a question
Answer, clip choice or time span
Accuracy; temporal overlap
None
MEDit-Bench ( Ogata et al., 2026 )
One long video and an editing message
Cut list
Temporal overlap with professional edits
User study on a subset (1,620 evaluations)
AgenticVBench ( Cao et al., 2026 )
Source videos and a brief or storyboard
Video with a manifest or report
Programmatic verifiers; 1,069 binary rubric items for repurposing
Expert rubric grading; three-editor human baseline on a subset
CutVerse ( Hu et al., 2026 )
GUI application state and an objective
GUI trajectory
Milestone checks
Not reported
ProSoftArena ( Ai et al., 2026 )
Real desktop and a task
Final state or artifact
Execution scripts; subjective comparison with human work on creative tasks