cs.ROOct 5, 2026

Talk, Render, Act: Integrating Social Gesture and Digital Face with Synchronized Speech for Conversational Humanoid Robot

Authors: Jin Jiang, Kun Li, Jiancong Ma, Shengcai Liao

Organizations: College of Computing and Artificial Intelligence (CCAI), United Arab Emirates University (UAEU), Al Ain, Abu Dhabi, United Arab Emirates

Abstract

Expressive humanoid interaction requires speech, facial animation, and body gestures to form a coherent response. However, many full-body humanoid robots produce speech and gestures without a visually expressive face, while talking-face animation and robot gesture generation are typically developed separately. We present Talk, Render, Act (TRABot), an agent-based framework comprising specialized agents for motion-atom construction, dialogue generation, motion planning, and facial animation. First, to produce natural and semantically meaningful gestures, we construct Robot-Ready Semantic Motion Atoms by segmenting long-form, G1-retargeted BEAT2 motion into units with natural gesture boundaries, human-verified communicative functions, and feasible trajectories. Second, to preserve semantic order and coordinate body motion with the spoken response, we introduce a Semantic-Conditioned Compositional Planner. Given an ordered semantic function sequence and an estimated response duration, the planner selects approved atoms to realize the longest feasible action sequence while accounting for transitions and neutral recovery. Finally, we deploy a Streaming Face-Speech-Body Integration system on a physical G1 humanoid, combining streaming dialogue audio, audio-driven facial animation, and semantically planned body motion in a unified real-time interaction loop. Quantitative and qualitative experiments demonstrate that TRAbot achieves the best overall performance among all compared conditions in terms of naturalness, expressiveness, and multimodal coherence.

Figures & tables

Explore similar work

CardsList
  1. SocialHumanoid: Towards Expressive Humanoid Behavior via One-Step Co-Speech Motion Generation

    Sep 27, 2026Chengqun Yang, Tengjie Zhu, Liang Xu +10Human Motion GenerationHumanoid

  2. WaveSync: Constrained Wavefront Optimization for Synchronized Co-Speech Gestures in Humanoid Robots

    Jun 15, 2026Thang Tran Viet, Thanh Nguyen Canh, Gia Huy Uong +4Co-Speech Gesture GenerationImportance

  3. ECHO-G: Embodied Co-speech Humanoid mOtion Generation

    Sep 30, 2026Yizhao Li, Pusen Gao, Ming Wang +3Co-Speech Gesture GenerationHuman Motion Generation