Joint speech-gesture synthesis must coordinate two modalities despite limited paired data. Existing approaches often lack bidirectional interaction, have limited language coverage, or simplify body and finger representations. We present DualTrack, which couples pretrained speech and motion priors on a shared 12.5 Hz timeline. Causal adapters exchange previous-packet information, while current-state fusion coordinates the streams before they separately complete sixteen-codebook packets. We evaluate 43 BEAT2 recordings in four languages, with speakers held out from joint training and validation. On the shared English/Spanish inputs, without speech or motion prefixes, DualTrack achieves lower word error rate and full-motion Fréchet Gesture Distance, higher beat consistency and speech naturalness than the evaluated GELINA baseline.
Figures & tables
Figure 1: The overview of DualTrack (a) Adapt the SeMoCo-based tokenizer with one semantic and fifteen kinematic codebooks. (b) Freeze the tokenizer and pretrain MotionPrior’s temporal and depth predictors. (c) Couple MotionPrior and Qwen3-TTS through causal history exchange and Tick Fusion. Each stream predicts sixteen codes per 80 ms packet.
Gesture
Speech
Language
Model
N
FGD Full ↓
FGD Body ↓
BC ↑
Div-Body
Div-Full
WER/CER(%) ↓
NMOS ↑
English
Ground truth
22
0.000
0.000
0.442
4.671
16.085
39.34
2.972
Gelina
22
7.513
1.952
0.341
2.651
2.651
18.09
2.709
DualTrack
22
7.346 (5.436)
1.586 (0.882)
0.606
3.497
10.277
2.89
3.847
DualTrack †
22
6.889 (5.436)
1.573 (0.882)
0.567
3.570
11.499
3.45
3.830
Chinese
Ground truth
9
0.000
0.000
0.526
4.668
16.781
5.91
2.771
Table 1: Speech and gesture generation of Gelina and DualTrack on BEAT2 test subsets.
Model
Full FGD ↓
BC ↑
Div.
NMOS ↑
Full DualTrack
5.7175
0.5653
10.582
3.885
w/o residual+fusion
5.9637
0.5394
10.015
3.619
Table 2: Ablation results on the complete test split (136 recordings). Div. denotes full-body motion diversity. BC denotes beat consistency