Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available. In this work, we introduce VINCIE-NExT, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data. VINCIE-NExT decomposes video editing into a structured chain of composable sub-tasks (Video -> Image -> Image -> Video), routing editing intent through the image domain and enabling scalable joint training from heterogeneous image and video corpora under a unified diffusion objective. An image editing pair, synthesized by the model or supplied by the user, is prepended as an in-context visual demonstration that serves as a spatial appearance blueprint for every output frame. To ground appearance edits across the interleaved context, we introduce a novel position encoding that links image demonstrations and video frames in a shared spatial coordinate system, enabling pixel-faithful propagation of appearance changes to every output frame. Chain-of-Editing further provides principled test-time scaling: by executing the sub-task chain as progressive diffusion stages, editing quality can be improved by investing additional compute without retraining. Comprehensive experiments on OpenVE-Bench demonstrate the state-of-the-art performance across diverse editing categories, with ablations confirming the effectiveness of each component.
Figures & tables
Figure 1 : Qualitative showcase of VINCIE-NExT across diverse video editing scenarios. Each case presents frames extracted from the source video and the edited output. Our method enables temporally consistent stylization, dynamic background replacement, identity and attribute editing, scene transformation, and object removal, while preserving motion, structure, and visual coherence across frames.
Figure 2 : Overview of our proposed VINCIE-NExT framework. The video editing process is decomposed into a structured chain of sub-tasks ( Video→Image→Image→Video ). The image editing pair (Is,It) serves as an in-context visual demonstration that transfers image editing priors to video generation via a unified diffusion backbone.
Methods
#Reso.
Overall
Global Style
Background Change
Local Change
Local Remove
Local Add
Subtitle Edit
Creative Edit
Camera Edit
Runway Aleph
1280×720
3.65
3.72
2.62
4.18
4.16
2.78
3.62
3.64
4.53
VACE [ 28 ]
1280×720
1.57
1.49
1.55
2.07
1.46
1.26
1.48
1.47
1.62
OmniVideo [ 59 ]
640×352
1.31
1.11
1.18
1.14
1.14
1.36
1.00
2.26
1.00
InsViE [ 72 ]
720×480
1.53
2.20
1.06
1.48
1.36
1.17
2.18
2.02
1.09
Lucy-Edit [ 60 ]
1280×704
2.15
2.27
1.57
3.20
1.75
2.30
1.61
2.86
1.61
ICVE [ 37 ]
384×240
2.07
2.22
1.62
2.57
2.51
1.97
2.09
2.41
1.11
Table 1: Performance comparison of video editing methods across eight editing categories, evaluated by Gemini 2.5 Pro, with automatic ratings on a 1–5 scale following the OpenVE-Bench protocol. #Reso. denotes the output resolution. Grey rows denote closed-source commercial models.
Data Composition
CoE
Overall
Global Style
Background Change
Local Change
Local Remove
Local Add
Subtitle Edit
Creative Edit
Camera Edit
I2I
V2V*
V ↔ I ↔ I
✓
1.14
1.29
1.00
1.05
1.05
1.00
1.60
1.20
1.04
✓
✓
1.05
1.05
1.00
1.01
1.03
1.00
1.27
1.06
1.00
✓
1.34
1.68
1.00
1.12
1.03
1.06
2.23
2.09
1.03
✓
✓
1.80
2.58
1.86
2.03
1.46
1.48
1.62
2.30
1.15
✓
2.49
3.47
1.96
2.67
2.71
2.08
3.40
2.03
1.29
Table 2: Ablation study on data composition and Chain-of-Editing (CoE) on OpenVE-Bench, evaluated by Gemini 2.5 Pro. Rows with grey background indicate that CoE is applied at inference time. I2I, V2V*, and V ↔ I ↔ I are the three training data sources. Training on V ↔ I ↔ I data alone already enables meaningful video editing, and CoE provides consistent gains when this chain data is present. The full combination of all three data types with CoE achieves the best overall score. All rows use the model after the first training stage at 256×256 resolution.
Figure 3 : Impact of intermediate image editing quality on CoE video editing performance (OpenVE-Bench, scored by Gemini 2.5 Pro). Stronger external editors ( Qwen-Image-Edit [ 69 ] and Nano Banana [ 61 ] ) consistently outperform self-editing, with complementary strengths highlighting the modular advantage of CoE.
Figure 6
Figure 7 : Qualitative results on global scene stylization ( left , summer-to-snow) and localized compositional editing ( right , injecting fire/water/wind).
Figure 8 : Comparison of V2V direct inference, CoE with self editing, and CoE with an external editor on subject replacement ( left ) and removal ( right ).
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Architecture
DiT backbone
3B MM-DiT, initialized from the same T2V-pretrained backbone as VINCIE [ 51 ]
VAE
Video VAE inflated from SD3 [ 14 ] ; spatial ×8 , temporal ×4 , 16 channels
Patch size (t,h,w)
(1,2,2)
Text encoder
Flan-T5 [ 11 ]
Trainable / frozen
DiT trained; VAE and text encoder frozen
Training data
Appendix
Table 3: Implementation details of VINCIE-NExT.
Methods
#Reso.
Overall
Global Style
Background Change
Local Change
Local Remove
Local Add
Subtitle Edit
Creative Edit
Camera Edit
Runway Aleph
1280×720
3.50
3.47
2.84
3.88
3.88
2.79
3.50
3.23
4.48
VACE [ 28 ]
1280×720
1.17
1.41
1.16
1.43
1.00
1.05
1.02
1.13
1.16
OmniVideo [ 59 ]
640×352
1.02
1.02
1.00
1.00
1.00
1.00
1.16
1.00
1.00
InsViE [ 72 ]
720×480
1.40
2.25
1.23
1.60
1.00
1.23
1.22
1.68
1.02
Lucy-Edit [ 60 ]
1280×704
1.95
2.17
2.20
3.30
1.03
2.37
1.06
2.35
1.14
ICVE [ 37 ]
384×240
2.25
2.35
1.86
2.91
2.68
2.27
2.04
1.94
1.38
Appendix
Table 4: Performance comparison of video editing methods across eight editing categories, evaluated by Seed1.6-VL . Scores are automatic MLLM ratings on a 1–5 scale following the OpenVE-Bench protocol. #Reso. denotes the output resolution. Grey rows denote closed-source commercial models.
TDF3D-RoPE
CoE
Overall
Global Style
Background Change
Local Change
Local Remove
Local Add
Subtitle Edit
Creative Edit
Camera Edit
1.95
2.78
2.02
2.17
2.15
1.71
1.39
2.09
1.04
✓
2.30
2.83
1.69
2.67
2.33
1.97
3.12
2.41
1.33
✓
2.53
3.06
1.94
2.88
2.85
1.97
3.49
2.69
1.29
✓
✓
2.65
3.64
2.15
3.03
3.11
1.82
3.20
2.89
1.22
Appendix
Table 5: Ablation study on TDF3D-RoPE, which fuses inter-shot and intra-shot position encodings so that spatially corresponding patches across shots share zero relative position, enabling content-driven pixel-level correspondence. Evaluated on OpenVE-Bench with Gemini 2.5 Pro.
TPR
CoE
Overall
Global Style
Background Change
Local Change
Local Remove
Local Add
Subtitle Edit
Creative Edit
Camera Edit
2.44
3.30
2.01
2.78
2.41
1.84
3.27
2.48
1.33
✓
2.53
3.25
2.37
2.68
2.63
1.99
3.36
2.58
1.22
✓
2.53
3.06
1.94
2.88
2.85
1.97
3.49
2.69
1.29
✓
✓
2.65
3.64
2.15
3.03
3.11
1.82
3.20
2.89
1.22
Appendix
Table 6: Ablation study on Temporal Position Randomization (TPR) for image editing training data, evaluated on OpenVE-Bench with Gemini 2.5 Pro.
I2I
V2V*
V ↔ I ↔ I
SQ
PQ
O
✓
3.61
5.94
3.67
✓
3.65
5.90
3.65
✓
4.65
5.88
4.43
✓
✓
4.54
5.82
4.42
✓
✓
✓
5.23
5.66
4.86
Appendix
Table 7: Ablation study on data composition evaluated on GEdit-Bench [ 39 ] , with evaluation metrics including SQ (Semantic Consistency), PQ (Perceptual Quality), and O (Overall Score).
Method
V2V
Creat.
Decomp.
IC
I2I
Overall
Creative Edit
Global Style
Local Remove
OpenVE-Edit
✓
✓
2.49
2.31
3.16
1.85
Pure V2V (ours)
✓
2.32
1.66
3.13
2.33
V2V* + CoE (ours)
✓
✓
2.57
1.88
3.32
2.87
Full data + CoE (ours)
✓
✓
✓
✓
2.65
2.89
3.64
3.11
Appendix
Table 8: Comparison with Pure V2V training on OpenVE-Bench (Gemini 2.5 Pro). V2V: paired video editing data. Creat.: Creative Edit V2V data. Decomp.: decomposed, interleaved chain-of-editing format of V2V. IC: our V ↔ I ↔ I in-context data. I2I: OmniEdit image editing data.
Method
Overall
Global Style
Background Change
Local Edit
Creative Edit
Pure V2V (w/o decomp., w/o CoE)
2.32
3.13
1.89
2.22
1.66
Full data + CoE
2.65
3.64
2.15
2.65
2.89
Δ
+0.33 ( +14.2% )
+0.51 ( +16.3% )
+0.26 ( +13.8% )
+0.43 ( +19.5% )
+1.23 ( +74.1% )
Appendix
Table 9: Per-category gains of full data + CoE over Pure V2V on OpenVE-Bench (Gemini 2.5 Pro). Local Edit averages Local Change, Local Remove, and Local Add.
Table 16
Method
imaging_quality
quality_avg
Lucy-Edit [ 60 ]
0.652
0.800
ICVE [ 37 ]
0.688
0.812
DITTO [ 1 ]
0.687
0.816
Ours
0.715
0.819
Appendix
Table 12: VBench objective metrics on OpenVE-Bench.
Steps
Time
Overall
Background Change
Creative Edit
Global Style
Local Add
Local Change
Local Remove
Subtitle Edit
1
15.02s
2.30
1.72
2.03
3.59
1.64
2.48
2.44
3.03
4
15.93s
2.55
2.29
2.49
3.60
1.80
2.61
2.85
3.28
16
18.69s
2.79
2.69
2.72
3.67
2.24
2.96
3.15
3.42
64
35.20s
2.81
2.76
2.97
3.55
2.10
3.05
3.06
3.57
Appendix
Table 13: Test-time scaling via the number of denoising steps of the intermediate image-editing stage, evaluated on OpenVE-Bench by Gemini 3.1 Flash Lite. Inference time is measured on one H200 GPU.
Figure 9 : Test-time scaling of CoE: overall score (line, left axis) and inference time (bars, right axis) as the number of denoising steps of the intermediate image-editing stage increases. Scored by Gemini 3.1 Flash Lite.
Method
Inference Time
AnyV2V [ 33 ]
11min 46s
ICVE [ 37 ]
118.1s
Lucy-Edit [ 60 ]
39.8s
Ours (Stage 1 + Stage 2)
18.4s + 35.1s = 53.5s
Appendix
Table 14: Inference latency on one H200 GPU (50 denoising steps for all methods).
Method
#Reso.
Overall
Background Change
Camera Edit
Creative Edit
Global Style
Local Add
Local Change
Local Remove
Subtitle Edit
AnyV2V [ 33 ]
360×640
1.62
1.38
1.16
2.17
3.05
1.29
1.56
1.16
1.37
Ours (Stage 1)
256×256
2.79
2.71
1.15
2.88
3.52
2.33
2.96
3.17
3.36
Ours (Stage 2)
480×640
3.26
3.16
2.26
3.76
3.95
2.55
3.62
3.45
3.37
Appendix
Table 15: Comparison with AnyV2V on OpenVE-Bench, evaluated by Gemini 3.0 Flash Lite.
Figure 10 : Qualitative comparison on background change editing.
Figure 11 : Qualitative comparison on creative edit editing.
Figure 12 : Qualitative comparison on global style editing.
Figure 13 : Qualitative comparison on local add editing.
Figure 14 : Qualitative comparison on local change editing.
Figure 15 : Qualitative comparison on local remove editing.
Figure 16 : Qualitative comparison on subtitle edit editing.
Figure 17 : Qualitative comparison of three inference modes of VINCIE-NExT on subject replacement ( left ) and subject removal ( right ): V2V direct inference, CoE with self-editing, and CoE with an external image editor.
Video editing aims to modify input videos according to user intent. Recently, end-to-end training methods have garnered widespread attention, constructing paired video editing data through video generation or editing models. However, compared to image editing, the high annotation costs of video data severely constrain the scale, quality, and task diversity of video editing datasets when relying on video generative models or manual annotation. To bridge this gap, we propose LIVE, a joint training framework that leverages large-scale, high-quality image editing data alongside video datasets to bolster editing capabilities. To mitigate the domain discrepancy between static images and dynamic videos, we introduce a frame-wise token noise strategy, which treats the latents of specific frames as reasoning tokens, leveraging large pretrained video generative models to create plausible temporal transformations. Moreover, through cleaning public datasets and constructing an automated data pipeline, we adopt a two-stage training strategy to anneal video editing capabilities. Furthermore, we curate a comprehensive evaluation benchmark encompassing over 60 challenging tasks that are prevalent in image editing but scarce in existing video datasets. Extensive comparative and ablation experiments demonstrate that our method achieves state-of-the-art performance. The source code will be publicly available.
Weicheng Wang, Zhicheng Zhang, Zhongqi Zhang +6
Nankai University · Kuaishou Technology · Pengcheng Laboratory +1
Recent image editing systems have achieved impressive semantic understanding, visual fidelity, and instruction-following ability, while video editing remains substantially more difficult and costly. In this paper, we present a simple alternative to end-to-end video editing: instead of training a monolithic video editor, we transform a strong image editor into a video editor through anchor-based generation. Our key insight is that video editing can be decomposed into two subproblems: editing a sparse set of keyframes and propagating those edits across time. Based on this observation, we propose Anchor-based Video Editing (AVE), a two-stage framework in which a powerful image editor first performs composed editing on selected keyframes, and a motion-guided image-to-video diffusion model then generates the final video by treating the edited keyframes as fixed anchors. This design directly inherits the strengths of modern image editors while avoiding expensive end-to-end video editing training. Experiments on IVEBench and VIE-Bench show that AVE achieves strong performance in instruction following, temporal consistency, and content fidelity. Further ablations reveal that final video editing quality is strongly correlated with the quality of the image editor, suggesting that future progress in video editing may come from stronger image editing foundations and lightweight transfer to video. Code is available at https://github.com/wangf3014/AVE.
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc, compared with 58.95 for the strongest evaluated training-free baseline, while obtaining competitive results on IVEBench. A user study further shows a 51.8% overall preference for EditVid over 7 competing methods.
Adheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen +6