cs.CVJul 16, 2026

Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach

Authors: Elena RyuminaMaxim MarkitantovAlexandr AxyonovFedor ShchetininTimur AbdulkadirovDmitry RyuminAlexey Karpov

Organizations: St. Petersburg Federal Research Center of the Russian Academy of Sciences · St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), St. Petersburg, Russia · HSE University, St. Petersburg, Russia · ITMO University, St. Petersburg, Russia

Abstract

Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, acoustic, facial, and contextual patterns, while top-performing systems often rely on computationally expensive ensembles. We present a single text-centered multimodal approach for video-level ambivalence and hesitancy recognition for the 11th Affective & Behavior Analysis in-the-Wild (ABAW) Challenge. The proposed approach combines linguistic, acoustic, facial, and scene features using text-centered multimodal fusion model. Text Residual Fusion treats text as the anchor modality and applies gated residual adjustments based on the other modalities. Experiments on the Behavioural Ambivalence/Hesitancy (BAH) corpus confirm that text is the strongest unimodal modality. The Text Residual Fusion model achieves an average Macro F1-score (MF1) of 75.14% across the Development and Public Test subsets. On the Private Test subset, it reaches an MF1 of 78.24%, outperforming the text model by 4.03%. These results demonstrate that complementary multimodal information can improve recognition performance without requiring a large model ensemble.

Explore similar work

CardsList