cs.CVSep 29, 2026

WeLike2Party! In-Context Motion Transfer for Multi-Human Image Animation

Authors: Sangeyl Lee, Seunghyun Shin, Seungho Park, Wooseok Jeon, Hae-Gon Jeon

Organizations: Department of Artificial Intelligence, Yonsei University · AI Graduate School, GIST

Abstract

Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skeletons or parametric body meshes and struggle to preserve identity-motion binding under inter-person occlusion. To address this limitation, we propose WeLike2Party, a multi-human animation framework built on direct in-context video conditioning without explicit pose or mesh extraction at inference. We further introduce Reference Asymmetric RoPE Conditioning to preserve fine-grained appearance details, and Identity Binding Supervision to associate each reference identity with its intended motion trajectory. To support cross-identity training, we construct MotionTwin, a large-scale synthetic dataset comprising 14.4K cross-identity video pairs with shared subject and camera motions, totaling 84.3 hours of photorealistic video. We additionally present MotionTwin-Bench, a cross-identity benchmark specifically designed to evaluate subject-level visual fidelity and identity-motion binding. Extensive experiments on MotionTwin-Bench and real-world videos demonstrate that WeLike2Party outperforms recent state-of-the-art methods in subject-level visual fidelity, identity-motion binding, and overall perceptual quality, particularly in multi-person interactions with substantial occlusion.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Vera: Identity-Faithful Human Subject-to-Video Generation

    Jul 22, 2026Yulong Xu, Xinyue Liu, Shujuan Li +6Video Dataset

  2. MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

    Jul 15, 2026Zhongyi Zhang, Guangyuan Wang, Li Hu +6Category-Level Object Pose EstimationUnified Framework

  3. Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy

    Jun 5, 2026Yuan Zeng, Yujia Shi, Yuhao Yang +4Beyond Retrieval