ECHO-G: Embodied Co-speech Humanoid mOtion Generation
Organizations: Beihang University · Mondo Robotics · The Hong Kong University of Science and Technology · Nanjing University
Abstract
Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio-text-robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio-text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.
Figures & tables
| Method | Co-Speech Motion Characteristics | Robot Motion Quality | Efficiency | ||||||
| FGD | Div | MM | BA | Jerk | Foot Err. (m) | C-Slide (m/s) | E2E Time (ms/frame) | Peak RAM (MB) | |
| EMAGE+GMR | 4.976 | 0.749 | 0 | 0.172 | 26.951 | 0.013 | 0.163 | 21.3 | 2661 |
| GestureLSM+GMR | 5.008 | 0.561 | 1.016 | 0.158 | 24.444 | 0.010 | 0.169 | 20.5 | 3804 |
| Human-Retargeted | 4.725 | 0.408 | 1.498 | 0.161 | 8.001 | 0.003 | 0.050 | 6.56 | 18714 |
| Ours | 2.278 | 0.320 | 1.786 | 0.063 | 9.039 | 0.008 | 0.052 | 5.96 | 18542 |
| Method | FGD | Div | MM | BA | Jerk | Foot Err. (m) | C-Slide (m/s) |
| Audio-only | 2.360 | 0.360 | 1.702 | 0.081 | 4.497 | 0.008 | 0.038 |
| Text-only | 2.436 | 0.429 | 1.681 | 0.113 | 5.940 | 0.009 | 0.049 |
| Ours | 2.278 | 0.320 | 1.786 | 0.063 | 9.039 | 0.008 | 0.052 |
| Method | Overall | Human- likeness | Rhythm Matching | Motion Quality |
| EMAGE+GMR | 1.80 | 1.87 | 1.95 | 1.59 |
| Human-Retargeted | 2.65 | 2.67 | 2.38 | 2.90 |
| Unimodal (pooled) | 3.05 | 3.01 | 3.13 | 3.01 |
| Ours | 3.49 | 3.45 | 3.31 | 3.70 |