cs.CLSep 22, 2026

Enriching Speech Emotion Representations with Conversational Context

Authors: Arthur PeuvotRomaric BesançonGaël de ChalendarBianca VieruIoana Vasilescu

Abstract

Detecting emotions is necessary for building systems that can accurately and adaptively interact with humans. Speech Emotion Recognition (SER) has become an important research focus to develop intelligent spoken interfaces. However, most studies predict emotions at the utterance level, ignoring the conversational context, along with the emotional flow and speaker interactions it carries. In this paper, we introduce ACERT (Averaged Contextual Emotion Representation through Time), a module that integrates a flexible-length window of conversational context to better capture emotional evolution in spoken interactions. To evaluate the robustness of this method, we conducted experiments on datasets spanning diverse emotionally expressive styles and contexts. ACERT outperforms current state-of-the-art (SOTA) approaches on IEMOCAP, establishes the first context-aware benchmark on SAFE, and obtains strong results on MELD for unweighted, class-balanced metrics. Ablation studies show that ACERT's gains come from emotional and conversational continuity, rather than from speaker identity or acoustic conditions.

Explore similar work

CardsList
  1. TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

    Jun 29, 2026Sathvik Manikantan Napa Ugandhar, Hao Zhang, Alison Gunzler +5Emotion Recognition