Joint speech-gesture synthesis must coordinate two modalities despite limited paired data. Existing approaches often lack bidirectional interaction, have limited language coverage, or simplify body and finger representations. We present DualTrack, which couples pretrained speech and motion priors on a shared 12.5 Hz timeline. Causal adapters exchange previous-packet information, while current-state fusion coordinates the streams before they separately complete sixteen-codebook packets. We evaluate 43 BEAT2 recordings in four languages, with speakers held out from joint training and validation. On the shared English/Spanish inputs, without speech or motion prefixes, DualTrack achieves lower word error rate and full-motion Fréchet Gesture Distance, higher beat consistency and speech naturalness than the evaluated GELINA baseline.
Figures & tables
Figure 1: The overview of DualTrack (a) Adapt the SeMoCo-based tokenizer with one semantic and fifteen kinematic codebooks. (b) Freeze the tokenizer and pretrain MotionPrior’s temporal and depth predictors. (c) Couple MotionPrior and Qwen3-TTS through causal history exchange and Tick Fusion. Each stream predicts sixteen codes per 80 ms packet.
Gesture
Speech
Language
Model
N
FGD Full ↓
FGD Body ↓
BC ↑
Div-Body
Div-Full
WER/CER(%) ↓
NMOS ↑
English
Ground truth
22
0.000
0.000
0.442
4.671
16.085
39.34
2.972
Gelina
22
7.513
1.952
0.341
2.651
2.651
18.09
2.709
DualTrack
22
7.346 (5.436)
1.586 (0.882)
0.606
3.497
10.277
2.89
3.847
DualTrack †
22
6.889 (5.436)
1.573 (0.882)
0.567
3.570
11.499
3.45
3.830
Chinese
Ground truth
9
0.000
0.000
0.526
4.668
16.781
5.91
2.771
Table 1: Speech and gesture generation of Gelina and DualTrack on BEAT2 test subsets.
Model
Full FGD ↓
BC ↑
Div.
NMOS ↑
Full DualTrack
5.7175
0.5653
10.582
3.885
w/o residual+fusion
5.9637
0.5394
10.015
3.619
Table 2: Ablation results on the complete test split (136 recordings). Div. denotes full-body motion diversity. BC denotes beat consistency
Co-speech gesture generation requires both semantic expressivity and biomechanically plausible rhythmic motion. Existing holistic gesture models mix lexically grounded semantic gestures with frequent prosody-aligned beat gestures. This limits semantic grounding, speech-motion alignment, and kinematic smoothness. We propose \emph{DuoGesture}, a neuro-inspired and biomechanically informed approach that decomposes co-speech gesture synthesis into semantic and beat streams. The two streams are coordinated by a \emph{Semantic Variational Information Bottleneck}, a stochastic frame-level gate that learns when semantic gestures should override rhythmic beat motion. The semantic stream is controlled by \emph{Motion-Grounded Semantic Conditioning}, which replaces purely linguistic word embeddings with motion-language representations to provide motion-aligned semantic priors for long-tailed lexical triggers of gestures. The beat stream is further regularised by an \emph{Inertial Beat Prior}, an anthropometry-weighted arm-chain module that reduces jitter and improves rhythmic consistency without constraining semantic frames. Objective evaluations and subjective experiments show that DuoGesture outperforms strong baselines, while component ablations confirm the complementary roles of semantic grounding, stochastic stream selection, and biomechanical regularisation.
Ferdinand Paar, Lanmiao Liu, Aslı Özyürek +2
Radboud University, Nijmegen · Max Planck Institute for Psycholinguistics · Utrecht University
Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.
Vida Adeli, Soroush Mehraban, Jacob Rommann +3
Pickford AI · University of Toronto · Vector Institute
Real-time co-speech gesture generation must produce 3D motion clip by clip as speech arrives. Existing streaming methods are open-loop: each clip depends on past context, but the model cannot check or correct its trajectory. Small errors therefore accumulate and cause drift over long sequences. We observe that this failure is mainly caused by the lack of a forward constraint rather than poor short-clip quality. A plausible key pose at the end of each clip provides a destination anchor that limits drift. Based on this observation, we propose StreamTalk, a closed-loop framework with a periodic generate-retrieve-refine cycle. Streaming Pose-Guided Generation first predicts a coarse clip, retrieves a plausible tail pose from a speaker-specific motion database, and refines the clip using this pose before continuing to the next window. During training, Stochastic Anchor Masking randomly masks pose and translation frames, teaching the model to recover complete motion from sparse boundary conditions. A part-aware DiT separates hand, body, and translation streams to reduce interference between global displacement and local articulation. On BEAT2, StreamTalk achieves state-of-the-art FGD, reduces long-horizon drift relative to open-loop baselines, and runs in real time at 76 FPS. Project page: https://xiangyue-zhang.github.io/StreamTalk/.
Xiangyue Zhang, Jianfang Li, Jiaxu Zhang +2
The University of Tokyo · Alibaba Group · Nanyang Technological University +1