cs.CVOct 1, 2026

Bootstrapping Video Interaction Generation with Synthetic State Transitions

Authors: Jiho Jang, Jinyoung Kim, Nojun Kwak, Kyungjune Kim

Organizations: Seoul National University · Independent Researcher · Sejong University

Abstract

While recent video generative models can synthesize high-fidelity videos, they struggle to portray plausible physical interactions and the resulting state transitions, a critical bottleneck for applications in robotics and VR/AR. To address this, we introduce a framework to generate a scalable synthetic dataset of controllable interactions. Our pipeline leverages a structured taxonomy and state-of-the-art image editing models to create explicit start' and end' state images, which serve as visual anchors for the interaction. To generate a seamless video utilizing these anchors, we propose State-Guided Sampling (SGS), a novel sampling technique that mitigates artifacts common in naive conditional generation. Furthermore, we develop and validate a new automated evaluation system that aligns with human judgments to ensure data quality. Experiments show that fine-tuning a base model on our dataset significantly enhances its ability to generate plausible interactions.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Event-Driven Video Generation

    Mar 12, 2026Chika Maduabuchi, Jindong WangLong Video GenerationIterative Denoising Methods

  2. GraphVid: Interactive Graph-Controllable Video Generation

    Jul 23, 2026Vedant Shah, Onkar Susladkar, Tushar Prakash +5Video DatasetHuman Motion Generation

  3. SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation

    May 11, 2026Liangyang Ouyang, Ruicong Liu, Caixin Kang +2Human Motion GenerationDramadirector