ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing
Organizations: University of Rochester · Adobe Research
Abstract
Current video editors can insert objects but often struggle to make them participate in interactions such as being picked up or manipulated. We introduce ALIVE, a framework that makes inserted objects "alive" through coherent interactions with the source video's contents, using an edited first frame and an instruction naming only the added object. We curate 35,800 editing pairs combining 3D-rendered, model-generated, and real-world videos with general editing pairs from ROSE. Each pair differs in the target object's presence while preserving the surrounding action, teaching editors coordinated object behavior and source preservation. We further train a vision-language model (VLM) to predict interaction guidance from the same inputs. We introduce the ALIVE-interaction benchmark to assess interaction fidelity, source preservation, and visual coherence using a unified VLM-based protocol, and evaluate on the general video object insertion benchmark. Without VLM guidance, ALIVE improves Overall over the strongest evaluated baseline by 43.9% and 4.4% on the two benchmarks, respectively. VLM-predicted guidance further improves the ALIVE-interaction score by 0.95 points without additional user inputs.
Figures & tables
| VLM evaluation | Automatic metrics | ||||||||
| Method | Guidance | Overall | Task | Motion | PickScore | Frame CLIP | Video ViCLIP | TC-CLIP | TC-DINO |
| ALIVE-interaction benchmark (128 cases) | |||||||||
| Señorita | P1 | 60.24 | 53.52 | 36.52 | 19.95 | 24.04 | 19.09 | 0.9572 | 0.9494 |
| NovaEdit | P0 | 36.36 | 8.59 | 0.39 | 18.94 | 21.15 | 15.41 | 0.9744 | 0.9734 |
| I2VEdit | P0 | 56.55 | 48.44 | 33.79 | 19.68 | 23.48 | 18.61 | 0.9363 | 0.9132 |
| AnyV2V | P2 † | 42.75 | 44.53 | 31.25 | 19.77 | 24.12 | 19.48 | 0.9141 | 0.8916 |
| Training source | Pairs | Pooled | ALIVE-interaction benchmark ( ) | general video object insertion benchmark ( ) | ||||
|---|---|---|---|---|---|---|---|---|
| Overall | Overall | Task | Motion | Overall | Task | Motion | ||
| 3D-rendered videos | 14,362 | 74.85 | 73.08 | 76.37 | 62.89 | 77.06 | 78.16 | 71.12 |
| Real-world videos | 8,542 | 81.51 | 80.19 | 88.48 | 79.49 | 83.16 | 83.50 | 80.10 |
| Model-generated videos | 5,896 | 81.26 | 82.81 | 90.82 | 82.23 | 79.32 | 79.85 | 71.84 |
| General editing pairs (ROSE) | 7,000 | 70.93 | 60.12 | 60.74 | 45.12 | 84.36 | 88.35 | 81.31 |
| All four sources | 35,800 | 84.39 | 82.35 | 89.65 | 82.03 | 86.93 | 88.83 | 83.01 |
| Guidance | Task | ID | Motion | BG | Temp. | Visual | Overall | |
|---|---|---|---|---|---|---|---|---|
| P0 | 128 | 91.41 | 81.84 | 83.40 | 93.36 | 80.47 | 75.00 | 85.38 |
| P1 | 128 | 93.75 | 83.98 | 86.13 | 91.80 | 81.05 | 75.20 | 86.71 |
| P2 | 128 | 93.16 | 82.23 | 86.91 | 92.38 | 80.66 | 77.15 | 86.68 |
| P3 | 128 | 94.73 | 83.40 | 88.48 | 91.60 | 81.25 | 75.59 | 87.37 |
| P4 | 128 | 95.12 | 85.94 | 88.67 | 93.55 | 82.81 | 78.32 | 88.69 |
| VLM-enhanced P1 | ||||||||
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Output | Steps | Time (s) | Peak (GiB) |
| Señorita | 30 | 109.18 | 28.04 | |
| NovaEdit | 50 | 407.70 | 19.63 | |
| I2VEdit | 25 | 863.62 | 65.16 | |
| AnyV2V | 50 | 172.79 | 30.31 | |
| PropFly | 50 | 71.63 | 21.14 | |
| Wan 14B, P1 | 30 | 304.64 | 64.46 |
| Optimizer steps | Pairs | ALIVE-interaction benchmark | general video object insertion benchmark |
|---|---|---|---|
| 5k | 35,800 | 82.35 | 86.93 |
| 20k | 35,800 | 86.46 | 89.20 |
| Backbone | Steps | ALIVE-interaction benchmark | general video object insertion benchmark |
|---|---|---|---|
| Wan 2.2 TI2V (5B) | 5k | 71.85 | 81.54 |
| Wan 2.2 I2V (14B) | 5k | 89.82 | 89.44 |
| LTX-2.5 (22B) | 5k | 82.35 | 86.93 |
| Input | Role |
|---|---|
| Basic instruction | Identify the intended addition without prescribing an action sequence. |
| SOURCE | Establish non-target scene content, camera motion, and interaction context. |
| Edited first frame | Specify the intended object or person and initial edit. |
| REFERENCE | Provide identity, contact, occlusion, interaction, and natural-effects evidence. |
| CANDIDATE | Generated result to evaluate at every displayed timepoint. |
| Dimension | Criterion | Weight |
|---|---|---|
| Task fidelity | Complete addition and coherent participation in interactions supported by SOURCE and . | 0.25 |
| Object identity | Stable identity, geometry, and attributes consistent with . | 0.15 |
| Motion / propagation | Plausible motion or appropriate stability, contact, timing, and occlusion. | 0.20 |
| Background preservation | Faithful non-target source content and camera motion. | 0.15 |
| Temporal consistency | Stable appearance and smooth transitions without flicker or disappearance. | 0.15 |
| Visual quality | Realistic boundaries, lighting, and natural effects. | 0.10 |