LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation
Organizations: UC San Diego · University of Virginia · Meta · Amazon · Lambda
Abstract
We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often care not only about how the camera moves, but also about what the scene should look like at key future moments, especially the final frame. Existing camera controls specify viewpoint trajectories, while text prompts provide only coarse semantic guidance; neither precisely determines the content and spatial layout of future views. This limitation becomes particularly pronounced under large viewpoint changes, where the camera reveals regions that are not visible in the first frame. LIFT therefore uses the last-frame layout as an explicit control signal for the desired future scene. Since learning from such sparse layout guidance is substantially more challenging than conditioning on dense per-frame layouts, we introduce on-policy self-distillation (OPSD) to transfer the control capability of a dense-layout teacher to a last-frame-layout student. We further curate LIFT-Vista, a dataset featuring large viewpoint changes with camera and temporally consistent layout annotations. Experiments show that LIFT improves video quality, future-layout controllability, and camera controllability over other methods.
Figures & tables
| Method | Control | Video Quality | Camera Error | Semantic Consistency | |||||||
| Camera | Object Motion | Future Layout | FVD | FID | LPIPS | RotErr | TransErr | mIoU | |||
| Direct-a-Video | ✓ | ✓ | ✗ | 539.28 | 71.56 | 0.80 | 28.48 | 2.28 | 0.10 | 0.11 | 0.14 |
| MagicMotion | ✗ | ✓ | ✗ | 277.94 | 18.32 | 0.57 | 15.74 | 1.59 | 0.41 | 0.53 | 0.21 |
| GEN3C | ✓ | ✗ | ✗ | 89.59 | 15.60 | 0.51 | 3.91 | 2.62 | 0.18 | 0.38 | 0.19 |
| Uni3C | ✓ | ✗ | ✗ | 111.60 | 12.75 | 0.42 | 3.23 | 0.70 | 0.31 | 0.47 | 0.21 |
| LIFT | ✓ | ✓ | ✓ | 99.35 | 12.84 | 0.42 | 2.97 | 0.59 | 0.51 | 0.59 | 0.24 |
| Method | Training Sample (Steps BS) | Video Quality | Camera Error | Semantic Consistency | |||||
| FVD | FID | LPIPS | RotErr | TransErr | mIoU | ||||
| Direct last-frame SFT | 114.03 | 12.96 | 0.43 | 2.98 | 0.56 | 0.44 | 0.56 | 0.23 | |
| D2S-SFT | 110.53 | 12.99 | 0.42 | 3.02 | 0.51 | 0.47 | 0.59 | 0.23 | |
| Ours | 99.35 | 12.84 | 0.42 | 2.97 | 0.59 | 0.51 | 0.59 | 0.24 | |
| Ours w/o SFT anchor | – | 129.14 | 17.16 | 0.45 | 3.36 | 0.68 | 0.49 | 0.57 | 0.22 |
| OPSD Variant | Lastframe-Layout Mode | Camera-Only Mode | |||||||||
| FVD | FID | RotErr | TransErr | FVD | FID | RotErr | TransErr | ||||
| Teacher (Dense layout) | 123.50 | 13.72 | 2.92 | 0.51 | 0.60 | 0.60 | 0.23 | 147.20 | 14.25 | 3.85 | 0.63 |
| Student @ step0 | 137.10 | 14.08 | 3.58 | 0.58 | 0.40 | 0.56 | 0.22 | 147.20 | 14.25 | 3.85 | 0.63 |
| Lastframe-layout single-mode | 102.07 | 13.07 | 2.88 | 0.58 | 0.49 | 0.59 | 0.23 | 121.88 | 13.42 | 3.09 | 0.67 |
| Camera-only single-mode | 108.30 | 14.00 | 3.04 | 0.59 | 0.38 | 0.57 | 0.22 | 105.22 | 13.77 | 3.03 | 0.63 |
| Ours (dual-mode) | 99.35 | 12.84 | 2.97 | 0.59 | 0.51 | 0.59 | 0.24 | 95.66 | 13.18 | 3.30 | 0.63 |
| Lastframe-Layout Mode | Camera-Only Mode | ||||||||||
| FVD | FID | RotErr | TransErr | mIoU | FVD | FID | RotErr | TransErr | |||
| 0.5 | 102.63 | 13.40 | 2.936 | 0.620 | 0.4898 | 0.5862 | 0.2306 | 106.01 | 13.85 | 3.190 | 0.703 |
| 0.7 | 99.35 | 12.84 | 2.968 | 0.593 | 0.5094 | 0.5861 | 0.2351 | 95.66 | 13.18 | 3.304 | 0.632 |
| 0.9 | 93.05 | 12.96 | 2.978 | 0.653 | 0.5060 | 0.5832 | 0.2321 | 98.62 | 13.34 | 3.357 | 0.757 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Stage 1 | Stage 2 | Stage 3 |
| Training mode | Camera Control (SFT) | Dense Layout Control (SFT) | Dual-Mode OPSD |
| Base Model | Wan2.1-Fun-V1.1-1.3B- Control-Camera | Stage 1 | Stage 2 |
| Dataset size | 120,898 | 58,272 | 58,272 |
| Training steps | 8,000 | 4,000 | 500 |
| Global batch size | 32 | 32 | 16 |
| Learning rate |
| Layout given at | Video Quality | Camera Error | Semantic Consistency | |||||
| FVD | FID | LPIPS | RotErr | TransErr | mIoU | |||
| Dense | 98.84 | 12.58 | 0.398 | 2.544 | 0.594 | 0.6217 | 0.6081 | 0.2403 |
| 8 frames | 92.72 | 12.86 | 0.406 | 2.772 | 0.574 | 0.5758 | 0.5907 | 0.2384 |
| 4 frames | 95.86 | 12.79 | 0.407 | 2.751 | 0.559 | 0.5764 | 0.5859 | 0.2379 |
| Last frame only † | 99.35 | 12.84 | 0.418 | 2.968 | 0.593 | 0.5094 | 0.5861 | 0.2351 |
| Setting | Direct LF-SFT | D2S-SFT | OPSD (Ours) |
| Starting checkpoint | Stage 1 | Stage 1 | Stage 1 |
| Training schedule | 8K LF | 4K Dense + 2K 8-frame + 2K LF | 4K Dense + 500 OPSD |
| Global batch size | 32 | 32 | 32 (Dense) / 16 (OPSD) |
| Learning rate | (Dense) / (OPSD) | ||
| Total training steps | 8,000 | 8,000 | 4,500 |
| Sample updates (Steps BS) | 256K | 256K | 136K |
| Method | Video Quality | Camera Error | Semantic Consistency | Time | |||||
| FVD | FID | LPIPS | RotErr | TransErr | mIoU | s/step | |||
| All 50 rollout states | 127.24 | 15.02 | 0.441 | 3.467 | 0.624 | 0.4786 | 0.5740 | 0.2234 | 544.82 |
| Equal 10 states | 120.23 | 15.17 | 0.439 | 3.228 | 0.589 | 0.4807 | 0.5777 | 0.2223 | 185.28 |
| Late 10 states | 197.63 | 17.82 | 0.494 | 6.347 | 0.906 | 0.3089 | 0.5227 | 0.2024 | 183.56 |
| Prefix 10 states | 127.02 | 17.77 | 0.442 | 2.961 | 0.544 | 0.4922 | 0.5664 | 0.2219 | 117.21 |
| Prefix 10 states + SFT anchor | 102.07 | 13.07 | 0.422 | 2.879 | 0.583 | 0.4934 | 0.5883 | 0.2313 | 125.20 |
| Method | FVD | FID | LPIPS | RotErr | TransErr |
| w/o filter | 278.85 | 32.88 | 0.52 | 15.62 | 4.05 |
| w/ filter | 266.58 | 30.99 | 0.51 | 13.14 | 3.87 |
| Ours | Uni3C | MagicMotion | |
| Preference | 65.96% | 22.34% | 11.70% |