cs.AISep 29, 2026

DualTrack: Synchronized speech-gesture generation via symmetric coupling of pretrained priors

Authors: Yuanzhuo Hu, Zehan Liu, Xiaoyi Qin, Ming Li

Organizations: The Chinese University of Hong Kong, Shenzhen, China · X Square Robot

Abstract

Joint speech-gesture synthesis must coordinate two modalities despite limited paired data. Existing approaches often lack bidirectional interaction, have limited language coverage, or simplify body and finger representations. We present DualTrack, which couples pretrained speech and motion priors on a shared 12.5 Hz timeline. Causal adapters exchange previous-packet information, while current-state fusion coordinates the streams before they separately complete sixteen-codebook packets. We evaluate 43 BEAT2 recordings in four languages, with speakers held out from joint training and validation. On the shared English/Spanish inputs, without speech or motion prefixes, DualTrack achieves lower word error rate and full-motion Fréchet Gesture Distance, higher beat consistency and speech naturalness than the evaluated GELINA baseline.

Figures & tables

Explore similar work

CardsList
  1. DuoGesture: Motion-Grounded Semantic Conditioning and Biomechanical Beat Priors for Co-Speech Gesture Generation

    May 25, 2026Ferdinand Paar, Lanmiao Liu, Aslı Özyürek +2Co-Speech Gesture GenerationBiomechanics

  2. Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation

    Aug 31, 2026Vida Adeli, Soroush Mehraban, Jacob Rommann +3Co-Speech Gesture GenerationHuman Motion Generation

  3. StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring

    Aug 3, 2026Xiangyue Zhang, Jianfang Li, Jiaxu Zhang +2Co-Speech Gesture GenerationAudio-Video Generation