cs.CLAug 9, 2026

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

Authors: Esam GhalebHugh Mee WongKristina Kobrock

Organizations: Max Planck Institute for Psycholinguistics · Utrecht University · Osnabrück University

Abstract

Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet dialogue models typically rely on transcripts alone. We study how much referential information gestures and their combination with speech carry in multimodal dialogue under different partner visibility conditions. % We build models that identify the intended referent in a video-mediated referential communication game based on either the speech transcript, the skeletal representation of gesture, or both modalities. Our results show that gesture alone is predictive of the intended referent and that multimodal fusion is most beneficial when the transcript-based model is uncertain. Training-only alignment of learned representations with the referent image further improves the fusion model performance. % In a comparison with human interaction data, we further see pragmatic effects of interlocutor visibility on gesture production and informativeness as well as an entrainment effect in speech and multimodal, but not gesture, performance across rounds of repeated interaction. We thus make contributions to the technical modelling of multimodal information in human dialogue and the analysis of human interaction data via trained model representations.

Explore similar work

CardsList
  1. Negation Beyond the Verbal Channel: Temporal Multimodal Correlates in Dialogue

    Sep 14, 2026Leon Hammerla, Patrick Schrottenbacher, Alexander MehlerNonverbal CuesNegation

  2. Recognizing Co-Speech Gestures in-the-Wild

    May 29, 2026Sindhu B Hegde, K R Prajwal, Andrew ZissermanCo-Speech Gesture GenerationMultimodal Dataset