SocialHumanoid: Towards Expressive Humanoid Behavior via One-Step Co-Speech Motion Generation
Organizations: Shanghai Jiao Tong University · ZTE Corporation
Abstract
Humanoid robots are increasingly expected to serve as embodied social agents that communicate naturally with humans through face-to-face interaction. During such communication, humanoid robots require body behaviors that are synchronized with speech, affectively expressive, and suitable for real-time execution. However, existing co-speech methods are primarily developed for digital humans and lack joint support for affective control and low-latency continuous generation on physical embodiments. To bridge this gap, we present SocialHumanoid, a system for expressive humanoid behavior via one-step co-speech motion generation. Given response speech and a specified affective condition, SocialHumanoid generates each full-body motion window in a single forward pass and connects successive windows through motion-history conditioning. The generated human motion is further converted online into embodiment-compatible robot references and tracked by a whole-body controller for physical execution. To provide explicit supervision for affective body expression, we further introduce AffectMoCap, a 4-hour dataset captured from two professional actors, containing synchronized speech, body motion, fine-grained hand motion, and emotion annotations. On BEAT2, SocialHumanoid achieves the best FGD among the compared generation methods, competitive speech-motion synchrony, and approximately faster inference than GestureLSM under the same protocol. Perceptual evaluations further show that training with AffectMoCap improves affect recognition from generated body motion, while real-robot experiments demonstrate continuous affect-conditioned behavior and stable long-horizon execution. Our project page is https://rex0191.github.io/SocialHumanoid/.
Figures & tables
| Dataset | Body Scope | Affect Label | Scale |
|---|---|---|---|
| Trinity Mocap [ 59 ] | Full body (3D) | – | 4 h |
| S2G-2D [ 60 ] | Upper body (2D) | – | 60 h |
| TED-2D [ 61 ] | Upper body (2D) | – | 97 h |
| TWH Mocap [ 62 ] | Full body (3D) | – | 20 h |
| TED-3D [ 1 ] | Upper body (3D) | – | 97 h |
| S2G-3D [ 63 ] | Upper body (3D) | – | 38 h |
| Methods | FGD ( ) | BC ( ) | Diversity ( ) | NFE ( ) |
|---|---|---|---|---|
| Ground-Truth | – | 0.703 | 11.97 | – |
| HA2G [ 4 ] | 12.32 | 0.677 | 8.626 | 30 |
| DisCo [ 27 ] | 9.417 | 0.643 | 9.912 | 1 |
| CaMN [ 3 ] | 6.644 | 0.676 | 10.86 | 1 |
| DiffSHEG [ 8 ] | 7.141 | 0.743 | 8.21 | 25 |
| TalkShow [ 26 ] | 6.209 | 0.695 | 13.47 | 64 |
| Methods | AIST (s) ( ) |
|---|---|
| DiffSHEG [ 8 ] | 4.4136 |
| TalkShow [ 26 ] | 0.9279 |
| ProbTalk [ 29 ] | 0.1472 |
| EMAGE [ 7 ] | 0.1121 |
| MambaTalk [ 30 ] | 0.1580 |
| SynTalker [ 36 ] | 10.8981 |
| Source | Stimulus | Label Source | Match Acc (%) ( ) |
|---|---|---|---|
| BEAT2 [ 7 ] | Motion only | BEAT emotion label | 76.4 |
| BEAT2 [ 7 ] | Audio + motion | BEAT emotion label | 78.2 |
| ZeroEGGS [ 20 ] | Motion only | Dataset style label | 86.8 |
| ZeroEGGS [ 20 ] | Audio + motion | Dataset style label | 84.2 |
| AffectMoCap (ours) | Motion only | Actor emotion label | 88.6 |
| AffectMoCap (ours) | Audio + motion | Actor emotion label | 92.4 |
| System setting | Naturalness ( ) | Emotion Fit ( ) | Delay Impact ( ) |
|---|---|---|---|
| w/o affective condition | 4.06 | 3.64 | 3.18 |
| w/o AffectMoCap fine-tuning | 3.94 | 3.30 | 3.22 |
| w/o speech playback | 4.12 | 3.62 | 3.56 |
| MeanFlow retraining and sampling | 4.10 | 3.98 | 3.28 |
| SocialHumanoid | 4.23 | 3.92 | 3.14 |
| Variant | FGD ( ) | BC ( ) | Diversity ( ) | NFE ( ) |
|---|---|---|---|---|
| w/o temporal attention | 4.648 | 0.787 | 12.14 | 1 |
| w/o spatial attention | 5.031 | 0.783 | 12.51 | 1 |
| w/o affective condition | 4.730 | 0.768 | 12.88 | 1 |
| MeanFlow retraining and sampling | 4.078 | 0.769 | 13.89 | 1 |
| Shortcut retraining and sampling | 4.241 | 0.728 | 13.75 | 8 |
| Full model | 3.929 | 0.770 | 12.36 | 1 |
| Module | Time (s) |
|---|---|
| Motion generation | 0.006 |
| Online retargeting | 1.248 |
| Whole-body controller computation | 0.008 |
| Hyperparameter | Value |
|---|---|
| History length | 4 latent tokens (16 motion frames) |
| Motion window length | 128 frames (32 latent tokens) |
| Streaming hop length | 112 frames (28 latent tokens) |
| Training clip stride | 20 motion frames |
| RVQ temporal downsampling | |
| Residual quantizer stages | 6 per body region |