Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available. In this work, we introduce VINCIE-NExT, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data. VINCIE-NExT decomposes video editing into a structured chain of composable sub-tasks (Video -> Image -> Image -> Video), routing editing intent through the image domain and enabling scalable joint training from heterogeneous image and video corpora under a unified diffusion objective. An image editing pair, synthesized by the model or supplied by the user, is prepended as an in-context visual demonstration that serves as a spatial appearance blueprint for every output frame. To ground appearance edits across the interleaved context, we introduce a novel position encoding that links image demonstrations and video frames in a shared spatial coordinate system, enabling pixel-faithful propagation of appearance changes to every output frame. Chain-of-Editing further provides principled test-time scaling: by executing the sub-task chain as progressive diffusion stages, editing quality can be improved by investing additional compute without retraining. Comprehensive experiments on OpenVE-Bench demonstrate the state-of-the-art performance across diverse editing categories, with ablations confirming the effectiveness of each component.
Figures & tables
Figure 1 : Qualitative showcase of VINCIE-NExT across diverse video editing scenarios. Each case presents frames extracted from the source video and the edited output. Our method enables temporally consistent stylization, dynamic background replacement, identity and attribute editing, scene transformation, and object removal, while preserving motion, structure, and visual coherence across frames.
Figure 2 : Overview of our proposed VINCIE-NExT framework. The video editing process is decomposed into a structured chain of sub-tasks ( Video→Image→Image→Video ). The image editing pair (Is,It) serves as an in-context visual demonstration that transfers image editing priors to video generation via a unified diffusion backbone.
Methods
#Reso.
Overall
Global Style
Background Change
Local Change
Local Remove
Local Add
Subtitle Edit
Creative Edit
Camera Edit
Runway Aleph
1280×720
3.65
3.72
2.62
4.18
4.16
2.78
3.62
3.64
4.53
VACE [ 28 ]
1280×720
1.57
1.49
1.55
2.07
1.46
1.26
1.48
1.47
1.62
OmniVideo [ 59 ]
640×352
1.31
1.11
1.18
1.14
1.14
1.36
1.00
2.26
1.00
InsViE [ 72 ]
720×480
1.53
2.20
1.06
1.48
1.36
1.17
2.18
2.02
1.09
Lucy-Edit [ 60 ]
1280×704
2.15
2.27
1.57
3.20
1.75
2.30
1.61
2.86
1.61
ICVE [ 37 ]
384×240
2.07
2.22
1.62
2.57
2.51
1.97
2.09
2.41
1.11
Table 1: Performance comparison of video editing methods across eight editing categories, evaluated by Gemini 2.5 Pro, with automatic ratings on a 1–5 scale following the OpenVE-Bench protocol. #Reso. denotes the output resolution. Grey rows denote closed-source commercial models.
Data Composition
CoE
Overall
Global Style
Background Change
Local Change
Local Remove
Local Add
Subtitle Edit
Creative Edit
Camera Edit
I2I
V2V*
V ↔ I ↔ I
✓
1.14
1.29
1.00
1.05
1.05
1.00
1.60
1.20
1.04
✓
✓
1.05
1.05
1.00
1.01
1.03
1.00
1.27
1.06
1.00
✓
1.34
1.68
1.00
1.12
1.03
1.06
2.23
2.09
1.03
✓
✓
1.80
2.58
1.86
2.03
1.46
1.48
1.62
2.30
1.15
✓
2.49
3.47
1.96
2.67
2.71
2.08
3.40
2.03
1.29
Table 2: Ablation study on data composition and Chain-of-Editing (CoE) on OpenVE-Bench, evaluated by Gemini 2.5 Pro. Rows with grey background indicate that CoE is applied at inference time. I2I, V2V*, and V ↔ I ↔ I are the three training data sources. Training on V ↔ I ↔ I data alone already enables meaningful video editing, and CoE provides consistent gains when this chain data is present. The full combination of all three data types with CoE achieves the best overall score. All rows use the model after the first training stage at 256×256 resolution.
Figure 3 : Impact of intermediate image editing quality on CoE video editing performance (OpenVE-Bench, scored by Gemini 2.5 Pro). Stronger external editors ( Qwen-Image-Edit [ 69 ] and Nano Banana [ 61 ] ) consistently outperform self-editing, with complementary strengths highlighting the modular advantage of CoE.
Figure 6
Figure 7 : Qualitative results on global scene stylization ( left , summer-to-snow) and localized compositional editing ( right , injecting fire/water/wind).
Figure 8 : Comparison of V2V direct inference, CoE with self editing, and CoE with an external editor on subject replacement ( left ) and removal ( right ).
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Architecture
DiT backbone
3B MM-DiT, initialized from the same T2V-pretrained backbone as VINCIE [ 51 ]
VAE
Video VAE inflated from SD3 [ 14 ] ; spatial ×8 , temporal ×4 , 16 channels
Patch size (t,h,w)
(1,2,2)
Text encoder
Flan-T5 [ 11 ]
Trainable / frozen
DiT trained; VAE and text encoder frozen
Training data
Appendix
Table 3: Implementation details of VINCIE-NExT.
Methods
#Reso.
Overall
Global Style
Background Change
Local Change
Local Remove
Local Add
Subtitle Edit
Creative Edit
Camera Edit
Runway Aleph
1280×720
3.50
3.47
2.84
3.88
3.88
2.79
3.50
3.23
4.48
VACE [ 28 ]
1280×720
1.17
1.41
1.16
1.43
1.00
1.05
1.02
1.13
1.16
OmniVideo [ 59 ]
640×352
1.02
1.02
1.00
1.00
1.00
1.00
1.16
1.00
1.00
InsViE [ 72 ]
720×480
1.40
2.25
1.23
1.60
1.00
1.23
1.22
1.68
1.02
Lucy-Edit [ 60 ]
1280×704
1.95
2.17
2.20
3.30
1.03
2.37
1.06
2.35
1.14
ICVE [ 37 ]
384×240
2.25
2.35
1.86
2.91
2.68
2.27
2.04
1.94
1.38
Appendix
Table 4: Performance comparison of video editing methods across eight editing categories, evaluated by Seed1.6-VL . Scores are automatic MLLM ratings on a 1–5 scale following the OpenVE-Bench protocol. #Reso. denotes the output resolution. Grey rows denote closed-source commercial models.
TDF3D-RoPE
CoE
Overall
Global Style
Background Change
Local Change
Local Remove
Local Add
Subtitle Edit
Creative Edit
Camera Edit
1.95
2.78
2.02
2.17
2.15
1.71
1.39
2.09
1.04
✓
2.30
2.83
1.69
2.67
2.33
1.97
3.12
2.41
1.33
✓
2.53
3.06
1.94
2.88
2.85
1.97
3.49
2.69
1.29
✓
✓
2.65
3.64
2.15
3.03
3.11
1.82
3.20
2.89
1.22
Appendix
Table 5: Ablation study on TDF3D-RoPE, which fuses inter-shot and intra-shot position encodings so that spatially corresponding patches across shots share zero relative position, enabling content-driven pixel-level correspondence. Evaluated on OpenVE-Bench with Gemini 2.5 Pro.
TPR
CoE
Overall
Global Style
Background Change
Local Change
Local Remove
Local Add
Subtitle Edit
Creative Edit
Camera Edit
2.44
3.30
2.01
2.78
2.41
1.84
3.27
2.48
1.33
✓
2.53
3.25
2.37
2.68
2.63
1.99
3.36
2.58
1.22
✓
2.53
3.06
1.94
2.88
2.85
1.97
3.49
2.69
1.29
✓
✓
2.65
3.64
2.15
3.03
3.11
1.82
3.20
2.89
1.22
Appendix
Table 6: Ablation study on Temporal Position Randomization (TPR) for image editing training data, evaluated on OpenVE-Bench with Gemini 2.5 Pro.
I2I
V2V*
V ↔ I ↔ I
SQ
PQ
O
✓
3.61
5.94
3.67
✓
3.65
5.90
3.65
✓
4.65
5.88
4.43
✓
✓
4.54
5.82
4.42
✓
✓
✓
5.23
5.66
4.86
Appendix
Table 7: Ablation study on data composition evaluated on GEdit-Bench [ 39 ] , with evaluation metrics including SQ (Semantic Consistency), PQ (Perceptual Quality), and O (Overall Score).
Method
V2V
Creat.
Decomp.
IC
I2I
Overall
Creative Edit
Global Style
Local Remove
OpenVE-Edit
✓
✓
2.49
2.31
3.16
1.85
Pure V2V (ours)
✓
2.32
1.66
3.13
2.33
V2V* + CoE (ours)
✓
✓
2.57
1.88
3.32
2.87
Full data + CoE (ours)
✓
✓
✓
✓
2.65
2.89
3.64
3.11
Appendix
Table 8: Comparison with Pure V2V training on OpenVE-Bench (Gemini 2.5 Pro). V2V: paired video editing data. Creat.: Creative Edit V2V data. Decomp.: decomposed, interleaved chain-of-editing format of V2V. IC: our V ↔ I ↔ I in-context data. I2I: OmniEdit image editing data.
Method
Overall
Global Style
Background Change
Local Edit
Creative Edit
Pure V2V (w/o decomp., w/o CoE)
2.32
3.13
1.89
2.22
1.66
Full data + CoE
2.65
3.64
2.15
2.65
2.89
Δ
+0.33 ( +14.2% )
+0.51 ( +16.3% )
+0.26 ( +13.8% )
+0.43 ( +19.5% )
+1.23 ( +74.1% )
Appendix
Table 9: Per-category gains of full data + CoE over Pure V2V on OpenVE-Bench (Gemini 2.5 Pro). Local Edit averages Local Change, Local Remove, and Local Add.
Table 16
Method
imaging_quality
quality_avg
Lucy-Edit [ 60 ]
0.652
0.800
ICVE [ 37 ]
0.688
0.812
DITTO [ 1 ]
0.687
0.816
Ours
0.715
0.819
Appendix
Table 12: VBench objective metrics on OpenVE-Bench.
Steps
Time
Overall
Background Change
Creative Edit
Global Style
Local Add
Local Change
Local Remove
Subtitle Edit
1
15.02s
2.30
1.72
2.03
3.59
1.64
2.48
2.44
3.03
4
15.93s
2.55
2.29
2.49
3.60
1.80
2.61
2.85
3.28
16
18.69s
2.79
2.69
2.72
3.67
2.24
2.96
3.15
3.42
64
35.20s
2.81
2.76
2.97
3.55
2.10
3.05
3.06
3.57
Appendix
Table 13: Test-time scaling via the number of denoising steps of the intermediate image-editing stage, evaluated on OpenVE-Bench by Gemini 3.1 Flash Lite. Inference time is measured on one H200 GPU.
Figure 9 : Test-time scaling of CoE: overall score (line, left axis) and inference time (bars, right axis) as the number of denoising steps of the intermediate image-editing stage increases. Scored by Gemini 3.1 Flash Lite.
Method
Inference Time
AnyV2V [ 33 ]
11min 46s
ICVE [ 37 ]
118.1s
Lucy-Edit [ 60 ]
39.8s
Ours (Stage 1 + Stage 2)
18.4s + 35.1s = 53.5s
Appendix
Table 14: Inference latency on one H200 GPU (50 denoising steps for all methods).
Method
#Reso.
Overall
Background Change
Camera Edit
Creative Edit
Global Style
Local Add
Local Change
Local Remove
Subtitle Edit
AnyV2V [ 33 ]
360×640
1.62
1.38
1.16
2.17
3.05
1.29
1.56
1.16
1.37
Ours (Stage 1)
256×256
2.79
2.71
1.15
2.88
3.52
2.33
2.96
3.17
3.36
Ours (Stage 2)
480×640
3.26
3.16
2.26
3.76
3.95
2.55
3.62
3.45
3.37
Appendix
Table 15: Comparison with AnyV2V on OpenVE-Bench, evaluated by Gemini 3.0 Flash Lite.
Figure 10 : Qualitative comparison on background change editing.
Figure 11 : Qualitative comparison on creative edit editing.
Figure 12 : Qualitative comparison on global style editing.
Figure 13 : Qualitative comparison on local add editing.
Figure 14 : Qualitative comparison on local change editing.
Figure 15 : Qualitative comparison on local remove editing.
Figure 16 : Qualitative comparison on subtitle edit editing.
Figure 17 : Qualitative comparison of three inference modes of VINCIE-NExT on subject replacement ( left ) and subject removal ( right ): V2V direct inference, CoE with self-editing, and CoE with an external image editor.