cs.ROSep 30, 2026

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

Authors: Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu

Organizations: Beihang University · Mondo Robotics · The Hong Kong University of Science and Technology · Nanjing University

Abstract

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio-text-robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio-text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.

Figures & tables

Explore similar work

CardsList
  1. SocialHumanoid: Towards Expressive Humanoid Behavior via One-Step Co-Speech Motion Generation

    Sep 27, 2026Chengqun Yang, Tengjie Zhu, Liang Xu +10Human Motion GenerationHumanoid

  2. PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation

    Jun 18, 2026Zhangzhao Liang, Xiaofen Xing, Mingyue Yang +2Human Motion GenerationHuman Motion

  3. WaveSync: Constrained Wavefront Optimization for Synchronized Co-Speech Gestures in Humanoid Robots

    Jun 15, 2026Thang Tran Viet, Thanh Nguyen Canh, Gia Huy Uong +4Co-Speech Gesture GenerationImportance