When Semantics Matter: Reliability-Aware Semantic-Rhythm Control for Co-Speech Gesture Generation
Authors: Zhirui Xing, Long Ye, Kaige Li, Ziyi Xu, Ming Meng
Organizations: Hainan International College, Communication University of China, Lingshui, China · School of Data Science and Intelligent Media, Communication University of China, Beijing, China · School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University, Shenzhen, China
Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible motions, they often rely heavily on acoustic prosody while underutilizing textual semantics, especially when semantic annotations are incomplete, noisy, or unavailable. Consequently, the generated gestures may follow speech rhythm while failing to express the intended semantics. To address this problem, we propose a reliability-aware semantic-rhythm control framework for co-speech gesture generation. We first learn a discrete motion prior that represents continuous gestures in a compact and structured motion-code space. We then introduce a dual-branch semantic contribution estimation mechanism consisting of a full multimodal branch and an audio-only branch. Their distributional discrepancy is formulated as conditional information gain to quantify how much textual semantics changes the predicted motion. Based on this estimate, a controllable semantic-rhythm objective selectively strengthens semantic guidance in content-relevant segments while limiting unnecessary semantic intervention in rhythm-dominant segments. Furthermore, we treat background noise as an acoustic reliability condition and introduce noise-conditioned feature modulation together with beneficial latent perturbation to improve generation robustness under realistic acoustic environments. Experiments on benchmark datasets demonstrate that the proposed framework achieves a favorable balance among semantic expressiveness, rhythmic synchronization, motion diversity, and robustness, enabling reliable and controllable co-speech gesture generation.
Figures & tables
Figure 2: Overview of the proposed reliability-aware semantic contribution modeling framework.
Figure 3: Dual-branch semantic contribution estimation. The full multimodal and audio-only branches predict distributions over the same motion codebook; their KL divergence quantifies the contribution of textual semantics.
Method
BEAT
TED-Expressive
FGD ↓
BC ↑
Diversity ↑
SRGR ↑
FGD ↓
BC ↑
Diversity ↑
CaMN ( Liu et al. 2022a )
8.510
0.797
206.789
0.231
9.284
0.681
117.362
SemGes ( Liu et al. 2025a )
4.762
0.453
305.706
0.256
5.106
0.713
139.824
CoCoGesture ( Qi et al. 2026 )
4.685
0.743
317.030
0.252
5.347
0.721
143.517
EMAGE ( Liu et al. 2024a )
4.969
0.729
269.520
0.256
5.582
0.734
132.406
GestureLSM ( Liu et al. 2025b )
4.521
0.789
320.846
0.259
4.526
0.762
150.684
Table 1: Comparison with representative co-speech gesture generation methods on BEAT ( Liu et al. 2022a ) and TED-Expressive ( Liu et al. 2022b ) .
Figure 4: Qualitative comparison with GlobalDiff and MIBURI under four representative speaking dynamics: gradually calming, increasingly impassioned, consistently calm, and continuously intense. Three temporally ordered poses are overlaid for each example, with dashed arrows indicating the dominant hand trajectories. Compared with the baselines, our method better adapts gesture amplitude, direction, and temporal evolution to the changing vocal intensity.
Configuration
FGD ↓
BC ↑
Diversity ↑
SRGR ↑
Direct Fusion
5.084
0.789
314.672
0.241
Dual Branch
4.931
0.793
317.486
0.248
Dual Branch + CIG
4.846
0.799
320.318
0.254
Full Model
4.362
0.805
322.455
0.261
Table 2: Ablation study of semantic contribution modeling on BEAT.
Figure 5: Qualitative comparison under different acoustic conditions.
Clean
10 dB
0 dB
Configuration
FGD ↓
BC ↑
SRGR ↑
FGD ↓
BC ↑
SRGR ↑
FGD ↓
BC ↑
SRGR ↑
w/o Reliability
4.879
0.800
0.254
5.421
0.748
0.229
6.736
0.672
0.194
AdaLN Only
4.824
0.802
0.256
5.183
0.769
0.241
6.081
0.706
0.211
Perturbation Only
4.837
0.801
0.255
5.246
0.763
0.238
6.204
0.699
0.207
Full Model
4.362
0.805
0.261
5.012
0.782
0.249
5.742
0.731
0.224
Table 3: Ablation study of acoustic reliability modeling on BEAT.