TripleFlow: Training-Free Video Object Removal by Bridging Residual Editing and Native Generation
Authors: Songhe Wang, Lifu Wei, Shuolin Xu, Charles A. Kamhoua, David Miller
Organizations: CSE Department, Penn State University · Department of Computer Science, The University of British Columbia, Canada · Department of Computing, Bournemouth University, United Kingdom · DEVCOM Army Research Laboratory, Network Security Branch, Adelphi, MD · EE Department, Penn State University
Video object removal presents a uniquely difficult editing challenge. Because a removal prompt specifies only what to erase rather than what to generate, the model must infer and reconstruct a highly specific occluded background entirely from the surrounding context. Existing training-free methods struggle with this because their editing mechanisms act primarily as localized erasers. They fail to actively synthesize the missing background details and often leave behind ghosting artifacts. To solve this, we propose TripleFlow, a training-free framework that tightly couples erasure and generation. It coordinates a source flow, a residual flow, and a synthesis flow throughout the entire process. By reusing a single target prediction, the residual flow isolates and suppresses the object, while the synthesis flow independently reconstructs the occluded background. Crucially, TripleFlow injects this newly synthesized background back into the editing trajectory at every step. This continuous feedback loop ensures that the generated structures actively guide the removal process, achieving seamless completion that is spatiotemporally consistent with the unedited scene. Extensive evaluations across five challenging benchmarks demonstrate that TripleFlow establishes a new state-of-the-art, significantly outperforming existing baselines in both reconstruction fidelity and temporal consistency.
Figures & tables
Figure 1: TripleFlow performs high-quality zero-shot video object removal without additional training, producing clean and spatiotemporally coherent results in challenging scenarios. It removes not only the target object, but also associated effects such as cast shadows, mirror reflections, and gravity-induced motion.
Figure 2: Object removal with general-purpose video editors. Given an empty-road prompt, FlowDirector ( Li et al., 2025a ) and DNAEdit ( Xie et al., 2025 ) retain vehicles, while RF-Edit ( Wang et al., 2025 ) removes the car but substantially changes the road and surroundings.
Figure 3: TripleFlow pipeline. Our framework coordinates three coupled flows within a frozen video diffusion backbone at each noise step. The source flow preserves the observed scene. By reusing a shared target prediction, the residual flow suppresses the foreground object using a velocity difference while the synthesis flow independently reconstructs the occluded background. The framework injects this synthesized background back into the residual editing trajectory, followed by mask projection before the next step.
Method
DAVIS ↑
WIPER ↑
PROVE-M ↑
PROVE-H ↑
ROSE ↑
OmnimatteZero
2.917
3.188
2.576
3.167
3.260
ObjectWiper
2.056
2.375
2.232
2.389
2.625
ContextFlow
2.667
3.500
2.438
3.111
2.357
OmniEraser
3.000
3.188
3.125
2.889
2.929
TripleFlow (Ours)
3.394
3.576
3.154
3.275
3.325
Table 1: Video object removal results. CORE is computed as the mean of ObjectScore and AftereffectScore (higher is better).
Method
PROVE-M
ROSE
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
TripleFlow
24.636
0.8644
0.1724
27.387
0.9112
0.1053
OmnimatteZero
20.894
0.8106
0.2819
25.450
0.8687
0.1877
ObjectWiper
16.608
0.6616
0.4129
18.897
0.6893
0.3655
ContextFlow
17.951
0.7912
0.3144
22.015
0.8912
0.1599
OmniEraser
19.044
0.7881
0.3407
19.649
0.7848
0.2741
Table 2: Reconstruction quality against ground truth labels on PROVE-M and ROSE.
Figure 4: Left: Radar comparison across nine combined metrics from PROVE-M and ROSE. Further outward indicates better performance. Middle: First-place preference rates in the five-method human evaluation on ten showcase videos. Right: Performance-speed trade-off.
Figure 5: Video object removal under camera motion. TripleFlow removes foreground subjects and reconstructs temporally coherent backgrounds while preserving the original camera motion and scene structure.
Figure 6: Qualitative comparison with existing video object removal methods. TripleFlow achieves clean removal while preserving fine background details and scene geometry, as illustrated by the bookshelf contents and warehouse shelving. In contrast, OmnimatteZero, ObjectWiper, and ContextFlow exhibit residual objects, altered background appearance, or structural distortions in these examples.
Figure 7: Qualitative ablation results. The input frames mark the removal targets: shoes and their mirror reflection (top), and a yellow clock (bottom). Full TripleFlow removes the targets while preserving the surrounding scene. Ablating individual flows or editing control can leave object remnants or introduce visible artifacts.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Hole PSNR ↑
CORE-O ↑
CORE-A ↑
PROVE-M (16 cases)
Full N7
22.57
0.840
0.162
19.24
2.69
2.50
w/o Source Flow
19.33
0.778
0.213
9.60
2.00
2.06
w/o Residual Flow
20.11
0.808
0.198
10.43
2.00
2.00
w/o Synthesis Flow
22.28
0.834
0.170
17.34
2.44
2.31
w/o Editing Control
21.54
0.824
0.181
15.13
2.19
2.06
Appendix
Table 3: N7 ablations on the fixed 20% evaluation subsets. All methods are evaluated on the same 16 PROVE-M and 12 ROSE cases. PSNR, SSIM, LPIPS, and hole-region PSNR are averaged per video over frames 1 onward, then equally averaged across cases. CORE-O and CORE-A denote paired CORE object-removal and aftereffect scores, respectively. Higher is better except for LPIPS. The w/o Editing Control variant retains mask projection and first-frame anchoring.
Figure 8: Additional qualitative ablations on PROVE-M and ROSE. Each row compares the input, full model, and variants without the source flow, residual flow, synthesis flow, or editing control. Red boxes indicate the removal targets.
Figure 9: Remove the kayak from the canal.
Figure 10: Remove the egret from the wetland.
Figure 11: Remove the shoes in front of the mirror (hard case).
Figure 12: Remove the coffee machine from the kitchen counter.
Figure 13: Remove the cyclist and bicycle from the courtyard.
Figure 14: Remove the visitor from the art gallery.
Figure 15: Remove the horse from the farmyard.
Figure 16: Remove the person holding a newspaper.
Figure 17: Remove the person from the observatory.
Figure 18: Remove the bicycle from the plaza.
Figure 19: Remove the forklift and driver from the warehouse.
Figure 20: Remove the person in front of the bookcase.
Figure 21: Remove the person from the doorway.
Figure 22: Remove the person from the entryway.
Figure 23: Remove the bicycle in front of the bench.
Video object removal is a fundamental yet challenging task in video editing. Despite recent progress, existing methods typically fall into two categories. Traditional approaches based on optical flow or attention mechanisms often introduce noticeable artifacts and yield unnatural results. In contrast, diffusion-based methods improve visual realism but demand multiple denoising steps, limiting their practicality. To address these issues, we propose From-Draft-to-Draft-Free (D2DF), a framework that distills the ability of transforming coarse drafts into refined videos into a one-step video generation model. Within D2DF, a teacher model is trained to refine low-quality removal results ("drafts") into high-fidelity videos by multiple steps. Then, through Prior-Privileged Consistency Distillation (PPCD), we distill this capability into a student model that performs one-step removal conditioned on the draft. To eliminate draft dependency, we introduce a Self-Guided Fast Planting (SGFP) module based on our Temporal Masked Transformer that autonomously generates scene-consistent pseudo-drafts in latent space, enabling a fully draft-free one-step model. Extensive experiments show that both draft-conditioned and draft-free versions achieve state-of-the-art performance on multiple metrics, surpassing traditional and multi-step generative methods in both quality and efficiency. The denoising process for a single video takes only about 1 second.
Zizhao Chen, Ping Wei, Guang Dai +2
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University · SGIT AI Lab, State Grid Corporation of China · Baidu +1
Existing video object removal methods predominantly rely on diffusion models following a noise-to-data paradigm, where generation starts from uninformative Gaussian noise. This approach discards the rich structural and contextual priors present in the original input video. Consequently, such methods often lack sufficient guidance, leading to incomplete object erasure or the synthesis of implausible content that conflicts with the scene's physical logic. In this paper, we reformulate video object removal as a video-to-video translation task via a stochastic bridge model. Unlike noise-initialized methods, our framework establishes a direct stochastic path from the source video (with objects) to the target video (objects removed). This bridge formulation effectively leverages the input video as a strong structural prior, guiding the model to perform precise removal while ensuring that the filled regions are logically consistent with the surrounding environment. To address the trade-off where strong bridge priors hinder the removal of large objects, we propose a novel adaptive mask modulation strategy. This mechanism dynamically modulates input embeddings based on mask characteristics, balancing background fidelity with generative flexibility. Extensive experiments demonstrate that our approach significantly outperforms existing methods in both visual quality and temporal consistency. The project page is https://bridgeremoval.github.io/.
Zijie Lou, Xiangwei Feng, Jiaxin Wang +7
MT Lab, Meitu Inc., Beijing 100083, China · Beijing Jiaotong University, Beijing 100044, China
Video object removal frequently struggles to simultaneously eliminate target objects and their associated physical effects (e.g., smoke, reflections, light, and ripples) in out-of-domain scenarios due to complex spatiotemporal ambiguities. While existing methods primarily rely on spatial masks, they often fail to capture weakly correlated effects, and the potential of explicit textual guidance remains underexplored. Furthermore, a fundamental optimization conflict exists in removal models between high-level semantic generalization and precise pixel-level background preservation. To address these challenges, we propose GenEraser, a novel framework for generalized and high-fidelity video object and effect removal. First, we introduce a Multi-Conditional Mixture-of-Experts (MC-MoE) paired with Bipartite Text guidance to fully exploit the multimodal priors of Diffusion Transformers, significantly enhancing the identification of complex effects. Second, a Learnable Deep ``CFG'' Fusion mechanism (LD-CFG) is developed to adaptively balance the relative dominance of mask and textual conditions across diverse scenarios. Finally, we propose a Decoupled Expert Architecture, comprising a Locator and a Preserver, to mitigate the inherent trade-off between semantic generalization and pixel alignment. Extensive experiments demonstrate that our GenEraser surpasses recent state-of-the-art approaches, achieving significant quantitative improvements (e.g., 2.16 dB and 1.44 dB on the ROSE Benchmark and VOR-Eval, respectively) while maintaining exceptionally robust generalization in open-world scenarios. https://cyqii.github.io/GenEraser.github.io/
Yuqing Chen, Lin Liu, Haisu Wu +4
Tsinghua University, China · Huawei, China · Southeast University, China +2