cs.ROSep 27, 2026

SocialHumanoid: Towards Expressive Humanoid Behavior via One-Step Co-Speech Motion Generation

Authors: Chengqun Yang, Tengjie Zhu, Liang Xu, Fulong Liu, Guanzhu Ren, Yitong Xing, Xuefeng Lu, Fei Shi, +5 more

Organizations: Shanghai Jiao Tong University · ZTE Corporation

Abstract

Humanoid robots are increasingly expected to serve as embodied social agents that communicate naturally with humans through face-to-face interaction. During such communication, humanoid robots require body behaviors that are synchronized with speech, affectively expressive, and suitable for real-time execution. However, existing co-speech methods are primarily developed for digital humans and lack joint support for affective control and low-latency continuous generation on physical embodiments. To bridge this gap, we present SocialHumanoid, a system for expressive humanoid behavior via one-step co-speech motion generation. Given response speech and a specified affective condition, SocialHumanoid generates each full-body motion window in a single forward pass and connects successive windows through motion-history conditioning. The generated human motion is further converted online into embodiment-compatible robot references and tracked by a whole-body controller for physical execution. To provide explicit supervision for affective body expression, we further introduce AffectMoCap, a 4-hour dataset captured from two professional actors, containing synchronized speech, body motion, fine-grained hand motion, and emotion annotations. On BEAT2, SocialHumanoid achieves the best FGD among the compared generation methods, competitive speech-motion synchrony, and approximately 6×6\times faster inference than GestureLSM under the same protocol. Perceptual evaluations further show that training with AffectMoCap improves affect recognition from generated body motion, while real-robot experiments demonstrate continuous affect-conditioned behavior and stable long-horizon execution. Our project page is https://rex0191.github.io/SocialHumanoid/.

Figures & tables

Explore similar work

CardsList
  1. ECHO-G: Embodied Co-speech Humanoid mOtion Generation

    Sep 30, 2026Yizhao Li, Pusen Gao, Ming Wang +3Co-Speech Gesture GenerationHuman Motion Generation

  2. PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation

    Jun 18, 2026Zhangzhao Liang, Xiaofen Xing, Mingyue Yang +2Human Motion GenerationHuman Motion

  3. WaveSync: Constrained Wavefront Optimization for Synchronized Co-Speech Gestures in Humanoid Robots

    Jun 15, 2026Thang Tran Viet, Thanh Nguyen Canh, Gia Huy Uong +4Co-Speech Gesture GenerationImportance