Recent image editing systems have achieved impressive semantic understanding, visual fidelity, and instruction-following ability, while video editing remains substantially more difficult and costly. In this paper, we present a simple alternative to end-to-end video editing: instead of training a monolithic video editor, we transform a strong image editor into a video editor through anchor-based generation. Our key insight is that video editing can be decomposed into two subproblems: editing a sparse set of keyframes and propagating those edits across time. Based on this observation, we propose Anchor-based Video Editing (AVE), a two-stage framework in which a powerful image editor first performs composed editing on selected keyframes, and a motion-guided image-to-video diffusion model then generates the final video by treating the edited keyframes as fixed anchors. This design directly inherits the strengths of modern image editors while avoiding expensive end-to-end video editing training. Experiments on IVEBench and VIE-Bench show that AVE achieves strong performance in instruction following, temporal consistency, and content fidelity. Further ablations reveal that final video editing quality is strongly correlated with the quality of the image editor, suggesting that future progress in video editing may come from stronger image editing foundations and lightweight transfer to video. Code is available at https://github.com/wangf3014/AVE.
Figures & tables
Paradigm
Consistency
Frame Quality
Instruction Following
Training Cost
End-to-end video editing
Good
Good
Depends on training
Expensive
Per-frame editing
Bad
Excellent
Bad
Cheap
Two-stage editing of AVE (ours)
Good
Excellent
Great and robust
Normal
Table 1: Intuitive comparison of three video editing paradigms. Our model attains favorable video consistency, frame-wise quality and instruction-following abilities with a normal cost.
Figure 1: Overview of our two-stage pipeline. In Stage-1, we edit a sparse set of keyframes using either open-source or API-based image editors. In Stage-2, an image-to-video diffusion model interpolates between the edited anchors while preserving the source motion via a lightweight motion encoder and anchor clamping.
Figure 2: Architecture of the lightweight motion encoder used in Stage-2. The encoder projects video latents into a token sequence and applies factorized spatial-then-temporal self-attention to extract motion features that condition the frozen image-to-video backbone.
Model
Total
Quality
Ins. Comp.
Fidelity
Motion
Sem. Consist.
Content
Ditto ( Bai et al., 2026 )
66.75
78.12
49.08
73.04
78.99
24.92
3.64
InsV2V ( Cheng et al., 2024 )
66.68
79.58
38.61
81.85
85.55
24.11
4.05
Lucy-Edit-Dev ( DecartAI Team, 2025 )
63.53
82.09
33.94
74.55
67.79
23.83
3.83
VACE ( Jiang et al., 2025 )
62.61
79.83
25.42
82.58
88.59
23.80
4.03
ICVE ( Liao et al., 2025 )
60.33
71.25
45.35
64.37
45.72
22.86
3.55
Omni-Video ( Tan et al., 2025 )
58.65
78.20
43.64
54.11
50.72
21.98
2.85
Table 2: IVEBench short subset results. Here we report the overall performance of total score, video quality, instruction compliance, video fidelity, and detailed metrics of motion smoothness, semantic consistency, and content fidelity and leave the remaining scores of IVEBench in Appendix. Our AVE model consistently obtains the best results among the evaluation metrics.
Model
Total
Quality
Ins. Comp.
Fidelity
Motion
Sem. Consist.
Content
Ditto ( Bai et al., 2026 )
65.93
77.95
47.85
71.98
72.82
23.42
3.69
InsV2V ( Cheng et al., 2024 )
65.72
80.24
37.41
79.50
67.88
24.06
4.13
Lucy-Edit-Dev ( DecartAI Team, 2025 )
64.89
82.10
31.51
81.05
73.46
23.92
4.13
VACE ( Jiang et al., 2025 )
61.61
80.12
26.73
77.98
88.40
23.56
3.74
ICVE ( Liao et al., 2025 )
58.77
71.96
40.38
63.98
48.40
22.60
3.48
Omni-Video ( Tan et al., 2025 )
57.09
77.84
42.44
51.00
55.11
21.90
2.59
Table 3: IVEBench long subset results. Similar to the short subset, our model attains the best results across the reported metrics. AVE has good robustness to long-video editing.
Method
Add
Swap / Change
Remove
Style / Tone
Ins. Follow
Pres.
Quality
Kling1.6 ( Kling Team et al., 2025 )
6.602
8.800
8.253
-
7.813
8.697
7.143
Kling-Omni ( Kling Team et al., 2025 )
9.181
9.194
9.133
9.341
9.518
9.346
8.751
Runway ( Runway, 2025 )
8.447
9.161
8.504
9.133
9.109
8.972
8.354
Pika ( Pika, 2026 )
-
7.408
-
-
7.542
7.847
6.837
MiniMax ( Zi et al., 2025 )
-
-
6.839
-
6.963
7.518
6.037
DiffuEraser ( Li et al., 2025 )
-
-
6.243
-
6.346
6.807
5.576
Table 4: Comparison results on VIE-Bench. We report the average score of each editing category (add, swap, remove, and style) and each evaluation metric (instruction following, motion preservation, and video quality). Results of closed-source commercial models are marked in gray .
Figure 3: Instruction-based video editing results produced by AVE. Guided by keyframe-level instructions and image-to-video extension, our method delivers high-fidelity visual details, strong temporal coherence, and robust performance across diverse editing tasks.
Stage-1 Image Editor
Total
Quality
Ins. Comp.
Fidelity
Motion
Sem. Consist.
Content
Nano Banana 2 (default)
68.75
82.15
48.31
75.79
88.95
24.68
4.16
GPT-5.1
68.10
81.43
47.11
75.76
88.74
24.10
4.02
Seedream 4.0
67.59
80.36
46.67
75.74
88.61
23.83
3.85
BAGEL (finetuned)
66.07
79.38
43.06
75.78
88.10
21.96
3.83
BAGEL (base)
58.54
68.12
37.84
69.66
77.21
20.34
3.42
Table 5: Ablation on the Stage-1 image editor on IVEBench-long. We replace only the composed image editor in Stage-1 and keep the entire Stage-2 pipeline fixed. The results show a strong correlation between image editing quality and final video editing performance.
Strategy
#Keyframes
Total
Quality
Ins. Comp.
Fidelity
Motion
Uniform
2
66.85
79.58
46.02
74.96
87.42
Uniform
⌊T/48⌋
67.73
80.84
47.12
75.24
88.21
Uniform
⌊T/24⌋
68.21
81.69
47.86
75.09
88.71
Uniform
⌊T/12⌋
67.94
81.27
47.40
75.15
88.38
Scene-first + max-gap (ours)
2
67.28
79.93
46.48
75.44
87.67
Scene-first + max-gap (ours)
⌊T/48⌋
68.09
81.16
47.56
75.55
88.46
Table 6: Ablation on keyframe selection strategies and keyframe budgets on IVEBench-long. Our default strategy combines scene-aware anchors from PySceneDetect with largest-gap completion. AVE is robust to the exact keyframe selection rule, while a moderate keyframe budget around ⌊T/24⌋ gives the best trade-off.
Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available. In this work, we introduce VINCIE-NExT, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data. VINCIE-NExT decomposes video editing into a structured chain of composable sub-tasks (Video -> Image -> Image -> Video), routing editing intent through the image domain and enabling scalable joint training from heterogeneous image and video corpora under a unified diffusion objective. An image editing pair, synthesized by the model or supplied by the user, is prepended as an in-context visual demonstration that serves as a spatial appearance blueprint for every output frame. To ground appearance edits across the interleaved context, we introduce a novel position encoding that links image demonstrations and video frames in a shared spatial coordinate system, enabling pixel-faithful propagation of appearance changes to every output frame. Chain-of-Editing further provides principled test-time scaling: by executing the sub-task chain as progressive diffusion stages, editing quality can be improved by investing additional compute without retraining. Comprehensive experiments on OpenVE-Bench demonstrate the state-of-the-art performance across diverse editing categories, with ablations confirming the effectiveness of each component.
Leigang Qu, Feng Cheng, Ziyan Yang +7
National University of Singapore · ByteDance Seed · University of Science and Technology of China
Video editing aims to modify input videos according to user intent. Recently, end-to-end training methods have garnered widespread attention, constructing paired video editing data through video generation or editing models. However, compared to image editing, the high annotation costs of video data severely constrain the scale, quality, and task diversity of video editing datasets when relying on video generative models or manual annotation. To bridge this gap, we propose LIVE, a joint training framework that leverages large-scale, high-quality image editing data alongside video datasets to bolster editing capabilities. To mitigate the domain discrepancy between static images and dynamic videos, we introduce a frame-wise token noise strategy, which treats the latents of specific frames as reasoning tokens, leveraging large pretrained video generative models to create plausible temporal transformations. Moreover, through cleaning public datasets and constructing an automated data pipeline, we adopt a two-stage training strategy to anneal video editing capabilities. Furthermore, we curate a comprehensive evaluation benchmark encompassing over 60 challenging tasks that are prevalent in image editing but scarce in existing video datasets. Extensive comparative and ablation experiments demonstrate that our method achieves state-of-the-art performance. The source code will be publicly available.
Weicheng Wang, Zhicheng Zhang, Zhongqi Zhang +6
Nankai University · Kuaishou Technology · Pengcheng Laboratory +1
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc, compared with 58.95 for the strongest evaluated training-free baseline, while obtaining competitive results on IVEBench. A user study further shows a 51.8% overall preference for EditVid over 7 competing methods.
Adheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen +6