cs.CVSep 29, 2026

ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

Authors: Donghao Zhou, Haoyang He, Fan Zhang, Hao Yang, Guisheng Liu, Xin Gao, Zhongwei Wan, Xingyuan Bu, +5 more

Organizations: The Ohio State University

Abstract

Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic encoders, so they often fall short in working with implicit edits that require causal or semantic reasoning. To bridge this fundamental gap in video editing, we propose ThinkV2V, a reasoning-driven framework for complex instruction-guided video editing, explicitly activating MLLM thinking before visual generation. At its core, ThinkV2V builds on a practical MLLM-to-DiT architecture to turn explicit thinking over the source video and instruction into refined conditioning signals for video editing. Further, we equip it with a dedicated training and inference recipe, combining Progressive Curriculum Training, which gradually cultivates the model from basic editing to reasoning-intensive cases, with Inference-Time Thinking Scaling, which iteratively refines candidate prompts and selects the most reliable one, to better elicit reasoning in challenging editing scenarios. We also curate the ThinkV2V-150K dataset and introduce ThinkV2V-Bench to support training and evaluation of video editing with implicit intent and causal reasoning. Experimental results demonstrate the state-of-the-art performance of ThinkV2V on both complex and standard editing scenarios, in which our 5B-scale DiT model substantially outperforms larger 10B-scale baselines.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reasoning to Align: Implicit Reasoning in Diffusion Transformers for Video Editing

    May 23, 2026Yan Li, Lin Liu, Xiaopeng Zhang +1Video EditingDiffusion-Based Image Editing

  2. MiVE: Multiscale Vision-language features for reference-guided video Editing

    May 14, 2026Tong Wang, Meng Zou, Chengjing Wu +4Video EditingVision Encoders

  3. One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing

    Sep 3, 2026Adheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen +6Video EditingEdit