Most video editing methods focus on changing the appearance of the source video, while offering limited control over its dynamics. We introduce Reimagine Video Dynamics (RVD), a framework that disentangles a compact, editable dynamics token from visual context. We learn this token through self-supervised reconstruction: given the first frame as visual context, a renderer must recover the original video from the dynamics token, encouraging it to capture how the scene evolves rather than how it looks. This disentanglement allows video dynamics to be edited directly while preserving visual context. We develop a language-guided dynamics-token editor that transforms source dynamics into target dynamics, and train it with a scalable counterfactual video-pair pipeline and a two-stage training strategy. Extensive experiments show that RVD enables effective video dynamics editing, training-free retiming, and appearance-controlled re-rendering.
Figures & tables
Figure 1: By disentangling video dynamics from visual context, RVD enables video dynamics editing, training-free retiming, and appearance re-rendering.
Figure 2: RVD architecture. (a) Renderer: Self-supervised reconstruction disentangles a compact dynamics token from visual context. (b) Editor: A language-guided editor modifies the dynamics token, which is then rendered into a video with the target dynamics.
Figure 3: Changing the world appearance while keeping its dynamics. The same dynamics token is rendered with a different first frame and caption.
Figure 4: Qualitative comparison of video appearance re-rendering.
Metric
RVD
Ditto
InsV2V
OmniVideo2
STDF
VideoGrain
VINO
IF ↑
4.438
3.973
1.615
3.712
2.085
2.260
4.345
Pass (%) ↑
86.8
74.5
10.5
63.5
26.0
27.2
81.3
JEPA ↑
0.9050
0.8496
0.6933
0.8256
0.8688
0.8686
0.8706
Imaging ↑
69.80
68.37
61.77
70.46
62.23
65.49
67.09
Table 1: Quantitative comparison of video appearance re-rendering.
Figure 5: Dynamics-token analysis. (a) Temporal variation spans the spatial grid. (b) Different motion states exhibit distinct dynamics-token distributions.
Figure 6: Training-free retiming in dynamics-token space. Resampling the dynamics token along its temporal axis directly slows down or speeds up the rendered motion without additional training.
Figure 7: Counterfactual data pipeline.
Figure 8: Video dynamics editing with RVD.
Figure 9: Qualitative comparison of video dynamics editing.
Method
Edit success
Output-internal consistency
IF ↑
Pass (%) ↑
Background stability ↑
Temporal smoothness ↑
RVD (ours)
4.837
98.3
0.9711
0.9919
Ditto ( Bai et al., 2026 )
2.537
36.7
0.9552
0.9858
OmniVideo2 ( Yang et al., 2026 )
2.187
27.0
0.9531
0.9777
VINO ( Chen et al., 2026 )
3.957
75.0
0.9645
0.9866
Table 4: Quantitative comparison of video dynamics editing.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: Detailed renderer architecture. Frozen V-JEPA features pass through the trainable dynamics bottleneck. DynCrossAttention injects only the compressed dynamics token into the finetuned Wan I2V backbone; image, text, and noisy-latent conditions retain Wan’s native paths.
Figure 11: Reconstruction under controlled renderer conditions. The full and caption-free renderers retain the source event, whereas removing the dynamics token changes or suppresses its evolution.
Pairing
Flow trajectory ↑
Motion profile ↑
Matched source–output
0.727
0.501
Mismatched source–output
0.312
0.012
Appendix
Table 5: Direct motion preservation when visual context changes under fixed D . The mismatched control pairs each output with another source from the same edit type.
Figure 12: More renderer comparisons. Six appearance edits compare RVD with Ditto ( Bai et al., 2026 ) , OmniVideo2 ( Yang et al., 2026 ) , and VINO ( Chen et al., 2026 ) at aligned relative time; each row reports its instruction.
Figure 13: Additional training-free retiming results. At the same normalized time, the 0.5× output has progressed less and the 2× output has progressed further; the 1× reconstruction follows the source timing.
Figure 14: Detailed editor architecture. Text, image, and learned queries enter the finetuned MLLM independently. Its semantic memory guides the trainable token editor through cross-attention and AdaLN; the target-token branch and token loss are training-only.
Group
Generated direction
Before
Pass
Filtered
Pass rate
Action
activity → sleep
1,994
1,875
119
94.0%
move → stop
480
370
110
77.1%
run → walk
3,000
2,469
531
82.3%
reduce distance
2,000
1,927
73
96.4%
Interaction
break → intact
196
126
70
64.3%
cut → whole
3,000
2,262
738
75.4%
Appendix
Table 6: Counterfactual generation and VLM quality-control statistics.
Figure 15: More editor comparisons. Five dynamics edits compare RVD with Ditto ( Bai et al., 2026 ) , OmniVideo2 ( Yang et al., 2026 ) , and VINO ( Chen et al., 2026 ) at aligned relative time; each row reports its instruction.
Figure 16: Editor contribution under fixed target context. All conditions use the same target context, instruction, seed, and renderer; only the dynamics input changes.
Dynamics condition
IF ↑
Pass (%) ↑
DINO ↑
None (target context only)
4.563
90.2
0.864
Unedited source
3.259
59.8
0.808
Matched video, same task
3.866
73.2
0.826
Edited source (ours)
4.705
95.5
0.921
Appendix
Table 7: Editor contribution on 112 matched cases with fixed target context. Only the dynamics input changes; DINO measures appearance consistency.
Edit family
Dynamics condition
IF ↑
Pass (%) ↑
DINO ↑
Trajectory / rhythm
None
4.55
92.2
0.859
Unedited source
3.47
68.6
0.840
Matched video
3.59
68.6
0.801
Edited source
4.69
98.0
0.912
State / interaction
None
4.57
88.5
0.868
Unedited source
3.08
52.5
0.781
Appendix
Table 8: Editor contribution by edit family. Target context and sampling conditions are identical within each case.
Video editing has advanced substantially in recent years, with methods increasingly accounting for the visual consequences of edits, such as changes to shadows and occlusions. However, the physical consequences of edits, including changes to subsequent motion and interactions, remain less explored. We formulate this problem as physical counterfactual video editing (PCVE), which aims to generate a counterfactual video depicting the resulting motion and interactions given a source video, a physical edit, and its execution frame. PCVE is challenging because it requires understanding scene physics and inferring the downstream motion and interactions induced by a physical intervention, while paired factual and counterfactual data and dedicated evaluation metrics are lacking. We introduce VideoPhysEdit, a new training-free pipeline for PCVE in rigid-body scenes. It makes physical reasoning explicit through a novel physical scene reconstruction method that recovers a scene reproducing the observed motion and interactions under simulation, enabling the pipeline to apply physical edits as interventions and use the resulting trajectories to guide counterfactual video generation. We further construct PCVE-RigidBench, a synthetic benchmark with paired source and counterfactual target videos and physical ground truth, and introduce the Physical Edit Score. VideoPhysEdit achieves substantially higher physical edit accuracy than open-source methods and commercial models while maintaining competitive visual fidelity. Its Physical Edit Score is 0.376, the only positive score among the compared methods. Qualitative comparisons on real videos further show that VideoPhysEdit applies to real-world scenes and better depicts the downstream motion and interactions induced by the edits than the compared methods. Code: https://github.com/Hammour-steak/VideoPhysEdit
Conghan Yue, Yuanjie Chen, Yue Han +4
Institute of Trustworthy Embodied AI, Fudan University
Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration. Our method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation (SA-DMD), and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift. Extensive automatic and human evaluations show that JoyAI-Video-Edit substantially outperforms existing streaming editors and remains competitive with strong offline systems on both short and long videos. The complete system achieves end-to-end 720p video editing at approximately 30 FPS on a single Nvidia B200 GPU. Code is available at https://github.com/jd-opensource/JoyAI-Video-Edit.
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc, compared with 58.95 for the strongest evaluated training-free baseline, while obtaining competitive results on IVEBench. A user study further shows a 51.8% overall preference for EditVid over 7 competing methods.
Adheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen +6