Talk, Render, Act: Integrating Social Gesture and Digital Face with Synchronized Speech for Conversational Humanoid Robot
Organizations: College of Computing and Artificial Intelligence (CCAI), United Arab Emirates University (UAEU), Al Ain, Abu Dhabi, United Arab Emirates
Abstract
Expressive humanoid interaction requires speech, facial animation, and body gestures to form a coherent response. However, many full-body humanoid robots produce speech and gestures without a visually expressive face, while talking-face animation and robot gesture generation are typically developed separately. We present Talk, Render, Act (TRABot), an agent-based framework comprising specialized agents for motion-atom construction, dialogue generation, motion planning, and facial animation. First, to produce natural and semantically meaningful gestures, we construct Robot-Ready Semantic Motion Atoms by segmenting long-form, G1-retargeted BEAT2 motion into units with natural gesture boundaries, human-verified communicative functions, and feasible trajectories. Second, to preserve semantic order and coordinate body motion with the spoken response, we introduce a Semantic-Conditioned Compositional Planner. Given an ordered semantic function sequence and an estimated response duration, the planner selects approved atoms to realize the longest feasible action sequence while accounting for transitions and neutral recovery. Finally, we deploy a Streaming Face-Speech-Body Integration system on a physical G1 humanoid, combining streaming dialogue audio, audio-driven facial animation, and semantically planned body motion in a unified real-time interaction loop. Quantitative and qualitative experiments demonstrate that TRAbot achieves the best overall performance among all compared conditions in terms of naturalness, expressiveness, and multimodal coherence.
Figures & tables
| Category | Communicative functions | Common family |
| emphasis | emphasize | beat |
| enumeration | enumerate | beat / metaphoric |
| contrast | contrast_compare | metaphoric |
| reference | refer_self , refer_other , refer_external | deictic |
| depiction | depict_size_shape , depict_direction_motion | iconic |
| turn management | turn_open , turn_yield | beat / metaphoric / emblem |
| Case | User utterance | Planned communicative-function sequence |
| 1 | Introduce artificial intelligence to me. | turn_open emphasize depict_size_shape neutral |
| 2 | Nice to meet you. Bye-bye. | greet_farewell neutral |
| 3 | What are three ways to relax? | enumerate enumerate enumerate neutral |
| 4 | Do you like swimming? | refer_self deny turn_yield neutral |
| 5 | What is the shape of a basketball? | depict_size_shape certainty neutral |
| 6 | Which way does the sun move across the sky? | depict_direction_motion neutral |
| Face offset (ms) | Body offset (ms) | Speech–body overlap (%) | Completed without failure (%) |
| 37.5 | 84.72 | 100 |
| Method | Binary BAS | Gaussian BAS |
| Simple Motion | 0.297 0.206 | 0.210 0.129 |
| Random Motion | 0.281 0.197 | 0.190 0.145 |
| Ours (Full model) | 0.308 0.141 | 0.225 0.130 |
| Condition | Natural | Express | Coherence | Gesture-speech | Preference (%) |
| Speech | 1.0958 | 1.1042 | 1.0958 | N/A | 0.83 |
| Speech+Face | 2.1042 | 2.0958 | 2.0833 | N/A | 0.42 |
| Simple Motion | 3.1458 | 3.1333 | 3.1583 | 3.1458 | 0.42 |
| Random Motion | 3.8667 | 3.9083 | 3.8708 | 3.8625 | 13.75 |
| Full model | 4.7875 | 4.7583 | 4.7917 | 4.7708 | 84.58 |
| Method | Planned Atoms per reply | Unused budget (s) |
| Random Match | 1.451 | 2.376 |
| Ours (Full model) | 1.705 | 0.993 |