Unifying Video Tasks via Spatiotemporal Analogy
Organizations: Cornell University · Meta
Abstract
Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse task taxonomy with a strict train-test split, ViGeo generalizes to unseen video manipulations and zero-shot modalities (e.g., event cameras). Finally, we identify task internalization, where a query format associated with a pretrained task overrides the demonstration, and show that this shortcut can be removed with a small amount of task-unrelated data, highlighting the need to decorrelate prompt format from task identity.
Figures & tables
| Category | Task Family | {\color[rgb]{1,0.3906,0.3047}A}/{\color[rgb]{0.375,0.8516,0.2148}C} (input) | {\color[rgb]{0.9961,0.6836,0}B}/{\color[rgb]{0.1758,0.5664,1}D} (target) | Layout | # | Split |
| Camera | Motion transfer | still frame | video @ traj. | 1 | Train + Eval | |
| View transfer | video @ pose 1 | video @ pose 2 | 2 | Train + Eval | ||
| Modal | Modality estimation | rgb video | modality video | 5 | Train | |
| Multi-modal quad | rgb video | 4-modal grid | 1 | Train | ||
| Segmentation | rgb video | mask / matted | 2 | Train | ||
| Guided generation | frozen rgb vid. modal | rgb video | 4 | Train |
| Optical Flow | Reconstruction Metrics | ||||
|---|---|---|---|---|---|
| Task | Method | EPE | PSNR | SSIM | LPIPS |
| Flow- Guided Generation | FloVD (CVPR 2025) | 3.771 | 12.363 | 0.240 | 0.539 |
| Go-with-the-flow (CVPR 2025) | 0.356 | 19.214 | 0.491 | 0.181 | |
| ViGeo | 0.185 | 22.365 | 0.671 | 0.106 | |
| Point- Guided Generation | Tora (CVPR 2025) | 0.612 | 0 6.923 | 0.366 | 0.230 |
| Wan-Move (NeurIPS 2025) | 0.289 | 17.048 | 0.365 | 0.157 | |
| Camera Metrics | |||||
|---|---|---|---|---|---|
| Task | Dataset | Method | RotErr | TransErr | CamErr |
| Camera Motion Transfer | DL3DV | CamCloneMaster (SIGGRAPH Asia 2025) | 3.367 | 13.941 | 15.408 |
| XFactor (ICLR 2026) | 1.066 | 0 5.378 | 0 5.761 | ||
| ViGeo | 2.173 | 14.612 | 15.203 | ||
| Camera Metrics | |||||
|---|---|---|---|---|---|
| Task | Dataset | Method | RotErr | TransErr | CamErr |
| Camera View Transfer | BridgeData-v2 | XFactor (ICLR 2026) | 36.131 | 0.857 | 1.235 |
| ViGeo | 2.382 | 0.093 | 0.117 | ||
| DROID | XFactor (ICLR 2026) | 64.824 | 1.200 | 1.959 | |
| ViGeo | 9.149 | 0.116 | 0.261 | ||
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Full Frame | Covisible | ||||||
|---|---|---|---|---|---|---|---|
| Task | Dataset | Method | PSNR | SSIM | LPIPS | PSNR | SSIM |
| Camera View Transfer | BRIDGE | XFactor (ICLR 2026) | 13.268 | 0.502 | 0.604 | 14.279 | 0.524 |
| ViGeo | 16.773 | 0.622 | 0.299 | 16.908 | 0.608 | ||
| DROID | XFactor (ICLR 2026) | 12.047 | 0.480 | 0.666 | 12.186 | 0.498 | |
| ViGeo | 14.470 | 0.564 | 0.377 | 14.416 | 0.569 | ||
| DROID ( ) | BRIDGE ( ) | |||
|---|---|---|---|---|
| Method | RotErr | TransDir | RotErr | TransDir |
| XFactor (ICLR 2026) | 68.92 | 63.64 | 39.48 | 49.12 |
| ViGeo | 31.58 | 18.77 | 27.32 | 22.79 |
| Depth-to-Flow | Flow-to-Depth | |||
|---|---|---|---|---|
| Method | MAE | SSIM | MAE | SSIM |
| ViGeo (base) | 0.375 | 0.335 | 0.333 | 0.257 |
| ViGeo (+ finetuning) | 0.142 | 0.818 | 0.108 | 0.793 |
| Task | Row Schema | Notes |
| Training tasks | ||
| est_depth | [ depth ] | |
| est_normal | [ normal ] | |
| est_semantic | [ semantic ] | per-pixel semantic class (NYU-40 palette). |
| est_flow | [ flow ] | |
| est_point | [ points ∗ † ] | ∗ static dots on a black image frame, indicating the initial points to be tracked. † RGB video with point tracks overlaid. |
| Task | Row schema | Notes |
| Training tasks | ||
| zoom-in_static | [full crop] | Static zoom-in to the same region |
| zoom-in_animated | [full dolly-in] | Animated zoom-in to the same region |
| zoom-out_static | [crop full] | — |
| zoom-out_animated | [crop dolly-out] | — |
| change_params | [orig degraded] | Samples a tuning curve over time and applies it to the original image; variants include saturation, temperature, tone, and chromatic magnitude |