Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation
Authors: Nick Milkin, Lanmiao Liu, Esam Ghaleb, Asli Ozyurek, Zerrin Yumak
Organizations: Max Planck Institute for Psycholinguistics, Nijmegen, The Netherlands · Utrecht University, Utrecht, The Netherlands · Donders Institute for Brain, Cognition and Behaviour, Nijmegen, The Netherlands
Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify. We present a perceptually grounded and semantics-aware benchmark that combines standardized model comparison, human-centered metric validation, and fine-grained semantic evaluation. We first curate a list of 13 objective metrics covering different aspects, including distributional similarity, geometric fidelity, kinematic quality, cross-modal synchrony, and semantic appropriateness. For the semantic-appropriateness category, we propose a new metric, Semantic Gesture Preservation (SGP), which measures how far semantic gestures in the ground truth are preserved in the generated gestures. For this, we augment the BEAT2 dataset's annotations using a multi-modal LLM. We then conduct a perceptual study where 101 participants score generated gestures among five dimensions, including human-likeness, motion diversity, absence of animation errors, speech timing and content match. We systematically analyze objective metric--subjective score correlations. Unlike Semantic Score (SC), which shows no significant association with the evaluated perceptual dimensions, SGP is selectively aligned with speech-aware human judgments. We construct five target-specific composite metrics aligned with the subjective dimensions. These composites improve perceptual alignment across all five dimensions, with the largest gains for absence of animation errors and content match, indicating that complementary objective signals can better approximate human judgments than individual metrics alone. Overall, our results show that objective metrics require validation against subjective evaluations.
Figures & tables
Figure 1: Perceptual study interfaces for the audio (left) and muted (right) conditions, evaluating speech-aware and visual-quality dimensions, respectively. Participants complete one condition using the same seven-point Likert scale and video interface.
Method
FGD ↓
BC ↑
Div. ↑
Dens. ↑
Cov. ↑
Dice ↑
LDLJ REL→0
SRGR ↑
PCK ↑
MJD ↓
Chamfer ↓
Hausdorff ↓
Foot Contact ↓
SC ↑
SGP ↑
EMAGE
5.523
0.692
88
0.8972
0.07275
0.6407
0.0530
0.05494
0.3309
0.2410
0.01375
0.2908
0.5048
0.248
0.4372
SemConFlow
2.245
0.780
120
0.01559
0.06647
0.6793
-0.4294
0.05868
0.3552
0.2007
0.01112
0.2530
0.5932
0.314
0.4205
SemTalk
4.266
0.727
116
0.1123
0.06673
0.6241
-0.6020
0.05059
0.3037
0.2303
0.13976
0.4663
0.4040
0.268
0.4014
GestureLSM
4.268
0.525
112
0.009239
0.009817
0.6573
1.5386
0.04536
0.2715
0.2391
0.01797
0.2909
0.8883
0.248
0.2622
Table 1: Objective comparison of the benchmarked holistic co-speech gesture generation methods on BEAT2. All metrics are computed over the 265 shared test sequences, except SGP, which is evaluated on the 15 sequences used in the perceptual study with Gemini-based semantic annotations. Arrows indicate the preferred direction, and bold denotes the best result for each metric.
Figure 2: Subjective ratings across five perceptual dimensions. Points show mean ratings with standard-deviation error bars; BEAT2 serves as the reference, and the remaining methods are generated results.
Perceptual dimension pair
Generated ( n=60 )
Including GT ( n=75 )
Human-likeness – Motion diversity
0.728
0.626
Human-likeness – Absence of animation errors
0.928
0.960
Human-likeness – Speech timing
0.650
0.813
Human-likeness – Content match
0.624
0.801
Motion diversity – Absence of animation errors
0.770
0.647
Motion diversity – Speech timing
0.665
0.564
Table 2: Pairwise Spearman correlations among subjective perceptual dimensions.
Figure 3: Spearman correlations between thirteen objective metrics and five subjective evaluation dimensions ( n=60 ), with significance indicated by ∗p<0.05 , ∗∗p<0.01 , and ∗∗∗p<0.001 .
Target Composite Metric
MSE ↓
MAE ↓
Composite ρ↑
Best individual
Δ∣ρ∣↑
Human-likeness
0.2846
0.4140
0.6568
LDLJREL
+0.0468
Motion diversity
0.2035
0.3610
0.5375
Coverage
+0.0117
Absence of animation errors
0.1560
0.2984
0.7065
LDLJREL
+0.0819
Speech timing
0.2671
0.4349
0.7439
Coverage
+0.0246
Content match
0.3665
0.4844
0.7824
Coverage
+0.0805
Table 3: LOOCV performance of target-specific composite predictors over 60 generated model–sequence observations using 13 predictors. MSE and MAE report out-of-fold error on the seven-point subjective scale, while composite ρ measures Spearman correlation with subjective ratings. Best individual is the strongest single objective metric by absolute Spearman correlation, and Δ∣ρ∣ denotes the composite gain.
Figure 4: Participant interface for the muted condition. Ratings become available only after the complete stimulus has been viewed.
Figure 5: Participant interface for the audio condition. After viewing the audiovisual stimulus, participants rate speech timing and content match.
Figure 6: Attention-check interfaces used in the two perceptual studies. The muted condition uses an instructed-response check, while the audio condition uses a content-based question about the preceding stimulus.
Figure 7: Prompt used for LLM-based semantic gesture annotation. The model receives a short video segment together with its predefined semantic word and returns a structured description of gesture form and contextual meaning.
Figure 8: Representative event-level motion example used for semantic gesture annotation. The figure visualizes a short rendered motion clip of approximately 2–3 s centered on the annotated semantic interval. This event clip, together with the corresponding target word and structured annotation prompt, is provided to Gemini 2.5 Pro as input for fine-grained semantic annotation.
Figure 9: Example of an LLM-generated semantic gesture annotation. The target semantic expression is paired with a fine-grained description of the visible motion and an interpretation of its communicative meaning.
Hainan International College, Communication University of China, Lingshui, China · School of Data Science and Intelligent Media, Communication University of China, Beijing, China · School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University, Shenzhen, China