When Semantics Matter: Reliability-Aware Semantic-Rhythm Control for Co-Speech Gesture Generation
Authors: Zhirui Xing, Long Ye, Kaige Li, Ziyi Xu, Ming Meng
Organizations: Hainan International College, Communication University of China, Lingshui, China · School of Data Science and Intelligent Media, Communication University of China, Beijing, China · School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University, Shenzhen, China
Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible motions, they often rely heavily on acoustic prosody while underutilizing textual semantics, especially when semantic annotations are incomplete, noisy, or unavailable. Consequently, the generated gestures may follow speech rhythm while failing to express the intended semantics. To address this problem, we propose a reliability-aware semantic-rhythm control framework for co-speech gesture generation. We first learn a discrete motion prior that represents continuous gestures in a compact and structured motion-code space. We then introduce a dual-branch semantic contribution estimation mechanism consisting of a full multimodal branch and an audio-only branch. Their distributional discrepancy is formulated as conditional information gain to quantify how much textual semantics changes the predicted motion. Based on this estimate, a controllable semantic-rhythm objective selectively strengthens semantic guidance in content-relevant segments while limiting unnecessary semantic intervention in rhythm-dominant segments. Furthermore, we treat background noise as an acoustic reliability condition and introduce noise-conditioned feature modulation together with beneficial latent perturbation to improve generation robustness under realistic acoustic environments. Experiments on benchmark datasets demonstrate that the proposed framework achieves a favorable balance among semantic expressiveness, rhythmic synchronization, motion diversity, and robustness, enabling reliable and controllable co-speech gesture generation.
Figures & tables
Figure 2: Overview of the proposed reliability-aware semantic contribution modeling framework.
Figure 3: Dual-branch semantic contribution estimation. The full multimodal and audio-only branches predict distributions over the same motion codebook; their KL divergence quantifies the contribution of textual semantics.
Method
BEAT
TED-Expressive
FGD ↓
BC ↑
Diversity ↑
SRGR ↑
FGD ↓
BC ↑
Diversity ↑
CaMN ( Liu et al. 2022a )
8.510
0.797
206.789
0.231
9.284
0.681
117.362
SemGes ( Liu et al. 2025a )
4.762
0.453
305.706
0.256
5.106
0.713
139.824
CoCoGesture ( Qi et al. 2026 )
4.685
0.743
317.030
0.252
5.347
0.721
143.517
EMAGE ( Liu et al. 2024a )
4.969
0.729
269.520
0.256
5.582
0.734
132.406
GestureLSM ( Liu et al. 2025b )
4.521
0.789
320.846
0.259
4.526
0.762
150.684
Table 1: Comparison with representative co-speech gesture generation methods on BEAT ( Liu et al. 2022a ) and TED-Expressive ( Liu et al. 2022b ) .
Figure 4: Qualitative comparison with GlobalDiff and MIBURI under four representative speaking dynamics: gradually calming, increasingly impassioned, consistently calm, and continuously intense. Three temporally ordered poses are overlaid for each example, with dashed arrows indicating the dominant hand trajectories. Compared with the baselines, our method better adapts gesture amplitude, direction, and temporal evolution to the changing vocal intensity.
Configuration
FGD ↓
BC ↑
Diversity ↑
SRGR ↑
Direct Fusion
5.084
0.789
314.672
0.241
Dual Branch
4.931
0.793
317.486
0.248
Dual Branch + CIG
4.846
0.799
320.318
0.254
Full Model
4.362
0.805
322.455
0.261
Table 2: Ablation study of semantic contribution modeling on BEAT.
Figure 5: Qualitative comparison under different acoustic conditions.
Clean
10 dB
0 dB
Configuration
FGD ↓
BC ↑
SRGR ↑
FGD ↓
BC ↑
SRGR ↑
FGD ↓
BC ↑
SRGR ↑
w/o Reliability
4.879
0.800
0.254
5.421
0.748
0.229
6.736
0.672
0.194
AdaLN Only
4.824
0.802
0.256
5.183
0.769
0.241
6.081
0.706
0.211
Perturbation Only
4.837
0.801
0.255
5.246
0.763
0.238
6.204
0.699
0.207
Full Model
4.362
0.805
0.261
5.012
0.782
0.249
5.742
0.731
0.224
Table 3: Ablation study of acoustic reliability modeling on BEAT.
While the field of co-speech gesture generation has seen significant advances, producing holistic, semantically grounded gestures remains a challenge. Existing approaches rely on external semantic retrieval methods, which limit their generalisation capability due to dependency on predefined linguistic rules. Flow-matching-based methods produce promising results; however, the network is optimised using only semantically congruent samples without exposure to negative examples, leading to learning rhythmic gestures rather than sparse motion, such as iconic and metaphoric gestures. Furthermore, by modelling body parts in isolation, the majority of methods fail to maintain crossmodal consistency. We introduce a Contrastive Flow Matching-based co-speech gesture generation model that uses mismatched audio-text conditions as negatives, training the velocity field to follow the correct motion trajectory while repelling semantically incongruent trajectories. Our model ensures cross-modal coherence by embedding text, audio, and holistic motion into a composite latent space via cosine and contrastive objectives. Extensive experiments and a user study demonstrate that our proposed approach outperforms state-of-the-art methods on two datasets, BEAT2 and SHOW.
Lanmiao Liu, Esam Ghaleb, Aslı Özyürek +1
Max Planck Institute for Psycholinguistics · Donders Institute for Brain Cognition and Behaviour · Utrecht University
Co-speech gesture generation requires both semantic expressivity and biomechanically plausible rhythmic motion. Existing holistic gesture models mix lexically grounded semantic gestures with frequent prosody-aligned beat gestures. This limits semantic grounding, speech-motion alignment, and kinematic smoothness. We propose \emph{DuoGesture}, a neuro-inspired and biomechanically informed approach that decomposes co-speech gesture synthesis into semantic and beat streams. The two streams are coordinated by a \emph{Semantic Variational Information Bottleneck}, a stochastic frame-level gate that learns when semantic gestures should override rhythmic beat motion. The semantic stream is controlled by \emph{Motion-Grounded Semantic Conditioning}, which replaces purely linguistic word embeddings with motion-language representations to provide motion-aligned semantic priors for long-tailed lexical triggers of gestures. The beat stream is further regularised by an \emph{Inertial Beat Prior}, an anthropometry-weighted arm-chain module that reduces jitter and improves rhythmic consistency without constraining semantic frames. Objective evaluations and subjective experiments show that DuoGesture outperforms strong baselines, while component ablations confirm the complementary roles of semantic grounding, stochastic stream selection, and biomechanical regularisation.
Ferdinand Paar, Lanmiao Liu, Aslı Özyürek +2
Radboud University, Nijmegen · Max Planck Institute for Psycholinguistics · Utrecht University
While recent advances in co-speech gesture generation have achieved impressive rhythmic synchronization, synthesizing gestures that are both semantically meaningful and faithful to a speaker's unique non-verbal style remains an open challenge. Semantic gestures, such as iconic shapes or deictic pointing, are statistically sparse, making them difficult to learn effectively within standard generative models. We present SiGnature, a framework for Stylized and Semantic Gesture generation that reconciles precise semantic control with high-fidelity style preservation. Unlike prevalent methods that rely on entangled latent representations, SiGnature operates in an explicit joint-rotation space. This design enables our core contribution, Joint Motion Integration (JMI), a training-free inference mechanism capable of injecting any external motion sequence, particularly in-the-wild semantic gestures, directly into the diffusion process. JMI automatically identifies the specific active joints'' conveying a semantic action and injects them into the generation, while relying on the diffusion backbone to synthesize the remaining body dynamics, including posture and flow, in accordance with the pre-learned style of the target speaker. This allows for the plug-and-play integration of arbitrary motions, including complex semantic gestures, without retraining or introducing the Frankenstein'' artifacts typical of cut-and-paste methods. Extensive experiments and perceptual studies demonstrate that SiGnature offers superior semantic motion control while maintaining smooth and natural co-speech gesture generation and preserving the distinct characteristics of the speaker, thereby outperforming state-of-the-art baselines.